Video analysis system
The video analysis system addresses the inefficiencies of existing technologies by objectively evaluating user reactions in online sessions, enhancing communication efficiency and user comfort by analyzing reactions without showing user images.
Patent Information
- Application Number
- JP2023529315
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-06-23
- Publication Date
- 2025-10-20
- Estimated Expiration
- 2041-06-23
AI Technical Summary
Existing video analysis technologies are not designed for the digital transformation of work and online communication, lacking efficiency and objectivity in evaluating user reactions during online meetings and lectures.
A video analysis system that analyzes user reactions based on video images obtained during online sessions, processing data to exclude user images and providing objective evaluation, allowing users to communicate with peace of mind.
Enables efficient and objective evaluation of online communication by analyzing user reactions without displaying their images, promoting better communication and understanding of emotional responses.
Smart Images

Figure 0007756446000001 
Figure 0007756446000002 
Figure 0007756446000003
Abstract
Description
[Technical Field]
[0001] The present invention relates to a video image analysis system that analyzes the biological responses of participants based on video images obtained in an online session conducted with multiple participants. [Background technology]
[0002] There is known a technique for analyzing the emotions felt by others in response to a speaker's comments (see, for example, Patent Document 1). There is also known a technique for analyzing changes in a subject's facial expression over a long period of time and estimating the emotions felt during that time (see, for example, Patent Document 2). There is also known a technique for identifying the factors that most influenced changes in emotions (see, for example, Patent Documents 3 to 5). There is also known a technique for comparing a subject's usual facial expression with their current facial expression and issuing an alert if the facial expression is gloomy (see, for example, Patent Document 6). There is also known a technique for comparing a subject's normal (expressionless) facial expression with their current facial expression to determine the degree of the subject's emotions (see, for example, Patent Documents 7 to 9). There are also known techniques for analyzing organizational emotions and the atmosphere felt by individuals within a group (see, for example, Patent Documents 10 and 11). [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Application Publication No. 2019-58625 [Patent Document 2] Japanese Patent Application Laid-Open No. 2016-149063 [Patent Document 3] Japanese Patent Application Publication No. 2020-86559 [Patent Document 4] Japanese Patent Application Laid-Open No. 2000-76421 [Patent Document 5] Japanese Patent Application Publication No. 2017-201499 [Patent Document 6] Japanese Patent Application Publication No. 2018-112831 [Patent Document 7] Japanese Patent Application Laid-Open No. 2011-154665 [Patent Document 8] Japanese Patent Application Laid-Open No. 2012-8949 [Patent Document 9] JP 2013-300 A [Patent Document 10] Japanese Patent Application Laid-Open No. 2011-186521 [Patent Document 11] WO15 / 174426 publication Summary of the Invention [Problem to be solved by the invention]
[0004] All of the above technologies are merely secondary functions in situations where communication in the physical world is the primary focus. In other words, they were not created in response to the recent digital transformation of work and the global pandemic, in which communication for work, classes, etc. is primarily conducted online.
[0005] The present invention aims to enable more efficient communication in situations where online communication is the norm, such as meetings and lectures, by enabling users to communicate with peace of mind and objectively evaluating such communication. [Means for solving the problem]
[0006] According to the present invention, there is provided a video analysis system that analyzes the reactions of users based on video images obtained by photographing the users in an environment where online sessions are held with multiple users, regardless of whether the users are displayed on the screen during the online session, and that includes a video acquisition unit that acquires video images for each of the multiple users by photographing the users during the online session, an analysis unit that analyzes changes in the user's biological reactions based on the video images acquired by the video acquisition unit, and a video processing unit that processes data related to the video images based on predetermined conditions so that the images of the users are not included in the video images. [Effects of the Invention]
[0007] According to the present disclosure, by analyzing and evaluating the video of the video session while displaying the video on the terminal, which does not include the user's image, evaluation can be made objectively, particularly regarding the content.
[0008] In particular, according to the present invention, in situations where online communication is the norm, users can communicate with peace of mind and objectively evaluate the communication exchanged, in order to communicate more efficiently. [Brief explanation of the drawings]
[0009] [Figure 1] 1 is a diagram showing an overall system according to an embodiment of the present invention; [Figure 2] FIG. 2 is a functional block diagram of an evaluation terminal according to an embodiment of the present invention; [Figure 3] FIG. 2 is a diagram illustrating a first example of a functional configuration of an evaluation terminal according to an embodiment of the present invention. [Figure 4] FIG. 2 is a diagram illustrating a second example of a functional configuration of an evaluation terminal according to an embodiment of the present invention. [Figure 5] FIG. 10 is a diagram illustrating a third example of a functional configuration of an evaluation terminal according to an embodiment of the present invention. [Figure 6]7 is a screen display example according to the functional configuration example 3 of FIG. 6. [Figure 7] 7 is another example of a screen display according to the functional configuration example 3 of FIG. 6. [Figure 8] FIG. 10 is a diagram illustrating another configuration of functional configuration example 3 of the evaluation terminal according to the embodiment of the present invention. [Figure 9] FIG. 10 is a diagram illustrating another configuration of functional configuration example 3 of the evaluation terminal according to the embodiment of the present invention. [Figure 10] FIG. 1 is a diagram illustrating an example of a functional configuration of a system according to an embodiment of the present invention. [Figure 11] 10A and 10B are diagrams showing an example of a moving image processed and output by the moving image processing unit according to the embodiment. [Figure 12] 10A and 10B are diagrams illustrating an example of an output form by an evaluation output unit according to the embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0010] The present disclosure will be described below with reference to the following embodiments. (Item 1) A video analysis system that analyzes reactions of a user based on video images obtained by photographing the user in an environment where an online session is held by a plurality of users, regardless of whether the user is displayed on a screen during the online session, comprising: a video acquisition unit that acquires, for each of the plurality of users, a video obtained by photographing the user during the online session; an analysis unit that analyzes changes in biological reactions of the user based on the moving images acquired by the moving image acquisition unit; a video processing unit that processes data relating to the video based on a predetermined condition so that the video does not include an image of the user; A video analysis system comprising: (Item 2) Item 1. The video analysis system according to item 1, The video processing unit performs the processing based on input information acquired by input to a terminal by the user. (Item 3) Item 2. The video analysis system according to item 2, a presentation unit that presents information about the execution of processing by the video processing unit to the user terminal; A video analysis system in which the video processing unit performs the processing based on input information obtained by the user inputting information into a terminal in response to the information presented by the presentation unit. (Item 4) The video analysis system according to any one of items 1 to 3, The video processing unit extracts only the audio contained in the video and generates it as a new video. (Item 5) The video analysis system according to any one of items 1 to 3, The video processing unit processes the video to display an object corresponding to an image of a user included in the video in place of the image of the user. (Item 6) The video analysis system according to any one of items 1 to 5, an output control unit that outputs the video to a terminal of each user participating in the online session; the moving image processing unit processes a moving image obtained by photographing one user; A video analysis system in which the output control unit outputs video data processed by the video processing unit to a terminal of a user other than the one user, and outputs video data that has not been processed by the video processing unit to the terminal of the one user.
[0011] Preferred embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings. In this specification and drawings, components having substantially the same functional configurations are designated by the same reference numerals, and redundant description will be omitted.
[0012] <Basic function> The video session evaluation system of this embodiment is a system that analyzes and evaluates the unique emotions (e.g., feelings of pleasure or discomfort in response to one's own or another's words or actions, or the degree of such feelings) of a subject of analysis among multiple people in an environment where the multiple people are engaged in a video session (hereinafter, both one-way and two-way sessions are referred to as online sessions) that are different from those of the other people. An online session may be, for example, an online conference, online class, or online chat. Terminals installed in multiple locations are connected to a server via a communication network such as the Internet, allowing video images to be exchanged between the multiple terminals through the server. Video images handled in an online session include facial images and audio of users using the terminals. Video images also include images of documents shared and viewed by multiple users. It is possible to switch between facial images and document images on the screen of each terminal, displaying only one of them, or to divide the display area and display both facial images and document images simultaneously. It is also possible to display an image of one of the multiple people in full screen, or to split the images of some or all of the users into small screens and display them. Of multiple users participating in an online session using terminals, it is possible to designate one or more users as analysis targets. For example, the leader, facilitator, or administrator of the online session (hereinafter collectively referred to as the organizer) designates one of the users as the analysis target. The organizer of an online session is, for example, a lecturer of an online class, a chairperson or facilitator of an online conference, or a coach of a session aimed at coaching. The organizer of an online session is usually one of the multiple users participating in the online session, but may also be a different person who does not participate in the online session. Note that it is also possible to not designate an analysis target, and to treat all participants as analysis targets. Also, the leader, facilitator, or administrator of the online session (hereinafter collectively referred to as the organizer) can designate one of the users as the analysis target. The organizer of an online session is, for example, a lecturer of an online class, a chairperson or facilitator of an online conference, or a coach of a session aimed at coaching.The host of an online session is typically one of multiple users participating in the online session, but may also be a different person who is not participating in the online session.
[0013] In the video session evaluation system according to this embodiment, when a video session is established between multiple terminals, at least a video image acquired from the video session is displayed. The displayed video image is acquired by the terminal, and at least a facial image contained in the video image is identified for each predetermined frame. An evaluation value for the identified facial image is then calculated. The evaluation value is shared as needed. In particular, in this embodiment, the acquired video image is stored in the terminal, analyzed and evaluated on the terminal, and the results are provided to the user of the terminal. Therefore, even if a video session contains personal information or confidential information, for example, the video itself can be analyzed and evaluated without providing the video itself to an external evaluation agency, etc. Furthermore, if necessary, the evaluation results (evaluation values) can be provided to an external terminal, allowing the results to be visualized and cross-analysis, etc. to be performed.
[0014] As shown in Figure 1, the video session evaluation system according to this embodiment includes user terminals 10 and 20 each having at least an input unit such as a camera unit and a microphone unit, a display unit such as a display, and an output unit such as a speaker, a video session service terminal 30 that provides a two-way video session to the user terminals 10 and 20, and an evaluation terminal 40 that performs part of the evaluation of the video session.
[0015] <Hardware configuration example> Each functional block, functional unit, and functional module described below can be configured, for example, using hardware, a DSP (Digital Signal Processor), or software provided in a computer. For example, when configured using software, the software is actually configured with a computer's CPU, RAM, ROM, etc., and is implemented by running a program stored in RAM, ROM, a hard disk, a semiconductor memory, or other storage medium. The series of processes performed by the system and terminal described herein can be implemented using software, hardware, or a combination of software and hardware. A computer program for implementing each function of the information sharing support device 10 according to this embodiment can be created and installed on a PC or the like. A computer-readable storage medium storing such a computer program can also be provided. Examples of storage media include a magnetic disk, an optical disk, a magneto-optical disk, and a flash memory. The computer program may also be distributed, for example, via a network, without using a storage medium.
[0016] The evaluation terminal according to this embodiment acquires moving images from a video session service terminal, identifies at least facial images contained in the moving images for each predetermined frame, and calculates an evaluation value for the facial images (details will be described later).
[0017] <How to get the video> As shown in FIG. 2, the video session service (hereinafter simply referred to as "this service") provided by the video session service terminal enables two-way image and audio communication with user terminals 10 and 20. This service displays video captured by the camera of the other user terminal on the display of the user terminal, and can output audio captured by the microphone of the other user terminal from the speaker. This service is also configured to enable both or either user terminals to record video and audio (collectively referred to as "video, etc.") in the memory of at least one of the user terminals. The recorded video information Vs (hereinafter referred to as "recorded information") is cached in the user terminal that initiated the recording and is recorded only locally on one of the user terminals. If necessary, users can view the recorded information themselves or share it with others within the scope of their use of this service.
[0018] <Functional configuration example 1> Fig. 4 is a block diagram showing an example of the configuration according to this embodiment. As shown in Fig. 4, the video session evaluation system of this embodiment is realized as a functional configuration possessed by a user terminal 10. That is, the user terminal 10 has, as its functions, a video image acquisition unit 11, a biological reaction analysis unit 12, a peculiar determination unit 13, a related event identification unit 14, a clustering unit 15, and an analysis result notification unit 16.
[0019] The video acquisition unit 11 acquires from each terminal video images obtained by capturing images of multiple people (multiple users) using a camera provided in each terminal during an online session. The video images acquired from each terminal may or may not be set to be displayed on the screen of each terminal. In other words, the video acquisition unit 11 acquires video images from each terminal, including video images currently being displayed and video images currently not being displayed on each terminal.
[0020] The biological response analysis unit 12 analyzes changes in biological responses for each of multiple people based on the moving images (whether or not they are being displayed on the screen) acquired by the moving image acquisition unit 11. In this embodiment, the biological response analysis unit 12 separates the moving images acquired by the moving image acquisition unit 11 into a set of images (a collection of frame images) and audio, and analyzes changes in biological responses from each.
[0021] For example, the biological reaction analysis unit 12 analyzes changes in biological reactions related to at least one of facial expression, eye movement, pulse rate, and facial movement by analyzing the user's facial image using frame images separated from the video acquired by the video acquisition unit 11. Furthermore, the biological reaction analysis unit 12 analyzes changes in biological reactions related to at least one of the user's speech content and voice quality by analyzing the audio separated from the video acquired by the video acquisition unit 11.
[0022] When a person's emotions change, this is reflected in changes in biological reactions such as facial expressions, eye movements, pulse rate, facial movements, speech content, and voice quality. In this embodiment, changes in the user's emotions are analyzed by analyzing changes in the user's biological reactions. One example of the emotion analyzed in this embodiment is the degree of comfort / discomfort. In this embodiment, the biological reaction analysis unit 12 quantifies changes in biological reactions according to a predetermined standard, thereby calculating a biological reaction index value that reflects the details of the changes in biological reactions.
[0023] The analysis of facial expression changes is performed, for example, as follows: For each frame image, a facial region is identified within the frame image, and the identified facial expressions are classified into multiple categories according to an image analysis model that has been trained in advance by machine learning. Based on the classification results, the system analyzes whether a positive or negative facial expression change has occurred between consecutive frame images, and the magnitude of the change, and outputs a facial expression change index value according to the analysis results.
[0024] The analysis of changes in gaze is performed, for example, as follows. That is, for each frame image, the eye area is identified within the frame image, and the direction of both eyes is analyzed to analyze where the user is looking. For example, it is analyzed whether the user is looking at the face of the speaker currently being displayed, at the shared material currently being displayed, or looking outside the screen. It may also be possible to analyze whether the gaze movement is large or small, or whether the movement is frequent or infrequent. The gaze change is also related to the user's concentration level. The biological response analysis unit 12 outputs a gaze change index value according to the analysis result of the gaze change.
[0025] The analysis of pulse rate changes is performed, for example, as follows. That is, for each frame image, the facial area is identified within the frame image. Then, using a trained image analysis model that captures the numerical value of facial color information (G in RGB), changes in the G color of the facial surface are analyzed. The results are arranged along the time axis to form a waveform representing changes in color information, and the pulse is identified from this waveform. When a person is nervous, their pulse rate increases, and when they feel calm, their pulse rate slows down. The biological response analysis unit 12 outputs a pulse rate change index value according to the analysis results of the pulse rate changes.
[0026] The analysis of changes in facial movement is performed, for example, as follows. That is, for each frame image, a facial area is identified within the frame image, and the facial direction is analyzed to analyze where the user is looking. For example, it is analyzed whether the user is looking at the face of the currently displayed speaker, the currently displayed shared material, or looking off-screen. It may also be analyzed whether the facial movement is large or small, or whether the movement is frequent or infrequent. It may also be analyzed by combining facial movement and eye movement. For example, it may be analyzed whether the user is looking directly at the currently displayed speaker's face, looking up or down, or looking at an angle. The biological response analysis unit 12 outputs a facial direction change index value according to the analysis result of the change in facial direction.
[0027] The analysis of speech content is performed, for example, as follows. That is, the biological response analysis unit 12 converts speech for a specified period of time (for example, approximately 30 to 150 seconds) into a string of characters by performing known speech recognition processing, and then performs morphological analysis on the string of characters to remove words unnecessary for expressing the conversation, such as particles and articles. The remaining words are then vectorized, and an analysis is performed to determine whether a positive or negative emotional change has occurred, and the extent of the emotional change, and a speech content index value corresponding to the analysis result is output.
[0028] Voice quality analysis is performed, for example, as follows: The biological response analysis unit 12 identifies the acoustic features of the voice by performing known voice analysis processing on the voice for a specified period of time (for example, approximately 30 to 150 seconds). Then, based on the acoustic features, it analyzes whether a positive or negative voice quality change has occurred and to what extent the voice quality change has occurred, and outputs a voice quality change index value according to the analysis result.
[0029] The biological reaction analysis unit 12 calculates a biological reaction index using at least one of the facial expression change index, eye direction change index, pulse rate change index, facial direction change index, speech content index, and voice quality change index calculated as described above. For example, the biological reaction index is calculated by weighting the facial expression change index, eye direction change index, pulse rate change index, facial direction change index, speech content index, and voice quality change index.
[0030] The unique determination unit 13 determines whether or not the change in biological reaction analyzed for the subject of analysis is unique compared to the change in biological reaction analyzed for other people other than the subject of analysis. In this embodiment, the unique determination unit 13 determines whether or not the change in biological reaction analyzed for the subject of analysis is unique compared to other people based on the biological reaction index values calculated for each of the multiple users by the biological reaction analysis unit 12.
[0031] For example, the unique determination unit 13 calculates the variance of the biological reaction index values calculated for each of multiple people by the biological reaction analysis unit 12, and by comparing the biological reaction index value calculated for the person being analyzed with the variance, determines whether the changes in the biological reactions analyzed for the person being analyzed are unique compared to others.
[0032] The following three patterns can be considered when changes in the analyzed biological reactions of the subject are unique compared to others. The first is when no particularly large changes in biological reactions occur in others, but a relatively large change in biological reactions occurs in the subject. The second is when no particularly large changes in biological reactions occur in the subject, but a relatively large change in biological reactions occurs in others. The third is when relatively large changes in biological reactions occur in both the subject and others, but the nature of the change differs between the subject and others.
[0033] The related event identification unit 14 identifies an event occurring with respect to at least one of the subject, other people, and the environment when a change in a biological reaction determined to be unique by the unique determination unit 13 occurs. For example, the related event identification unit 14 identifies, from video images, the behavior and words of the subject when a unique change in a biological reaction occurs in the subject. The related event identification unit 14 also identifies, from video images, the behavior and words of other people when a unique change in a biological reaction occurs in the subject. The related event identification unit 14 also identifies, from video images, the environment when a unique change in a biological reaction occurs in the subject. The environment may be, for example, shared documents displayed on the screen or something that appears in the background of the subject.
[0034] The clustering unit 15 analyzes the degree of correlation between a change in biological reaction determined to be unique by the unique determination unit 13 (for example, one or more combinations of eye contact, pulse rate, facial movement, speech content, and voice quality) and an event occurring when the unique change in biological reaction occurs (an event identified by the related event identification unit 14), and if the correlation is determined to be at a certain level or above, clusters the person or event being analyzed based on the analysis results of the correlation.
[0035] For example, if a specific change in biological reaction corresponds to a negative emotional change and the event occurring when the specific change in biological reaction occurs is also a negative event, a correlation of a certain level or higher is detected. The clustering unit 15 clusters the analysis subject or event into one of multiple pre-segmented classifications according to the content of the event, the degree of negativity, the magnitude of correlation, etc.
[0036] Similarly, if a change in a specific biological reaction corresponds to a positive emotional change and the event occurring when the change in the specific biological reaction occurs is also a positive event, a correlation of a certain level or higher is detected. The clustering unit 15 clusters the analysis subject or event into one of multiple pre-segmented classifications according to the content of the event, the degree of positivity, the magnitude of correlation, etc.
[0037] The analysis result notification unit 16 notifies the person designating the subject of analysis (the subject of analysis or the organizer of the online session) of at least one of the changes in biological reactions determined to be specific by the specific determination unit 13, the events identified by the related event identification unit 14, and the classifications clustered by the clustering unit 15.
[0038] For example, the analysis result notification unit 16 notifies the analysis subject of his / her own words and actions as an event occurring when a unique change in biological reaction occurs in the analysis subject that is different from that of others (one of the three patterns described above; the same applies below). This allows the analysis subject to understand that when he / she behaves in a certain way, he / she feels differently from others. At this time, the analysis subject may also be notified of the unique changes in biological reaction identified for the analysis subject. Furthermore, the analysis subject may also be notified of changes in biological reaction of others to be compared.
[0039] For example, if the emotions felt by others in response to words or actions made by the subject without any particular awareness and with normal emotions, or words or actions made by the subject with a particular awareness and with a certain emotion differ from the emotions felt by the subject himself at the time of the words or actions, the subject will be notified of his or her own words or actions at that time. This makes it possible to discover words or actions that are well-received by others or that are not well-received by others, despite the subject's own awareness.
[0040] Furthermore, the analysis result notification unit 16 notifies the organizer of the online session of events occurring when a unique change in biological reaction occurs in the analysis subject that is different from that of others, along with the unique change in biological reaction. This allows the organizer of the online session to know what events are influencing what emotional changes as phenomena unique to the designated analysis subject. Then, it becomes possible to take appropriate measures for the analysis subject based on the content of the information obtained.
[0041] Furthermore, the analysis result notification unit 16 notifies the organizer of the online session of events occurring when a unique change in the biological reaction of the analysis subject occurs that is different from that of others, or of the clustering results of the analysis subject. This allows the organizer of the online session to understand the behavioral tendencies unique to the analysis subject and predict possible future behaviors and conditions, etc., depending on which classification the specified analysis subject is clustered into. This then makes it possible to take appropriate measures for the analysis subject.
[0042] In the above embodiment, an example has been described in which a biological reaction index value is calculated by quantifying changes in biological reactions according to a predetermined standard, and whether or not the changes in biological reactions analyzed for the subject of analysis are unique compared to others is determined based on the biological reaction index values calculated for each of a plurality of people, but the present invention is not limited to this example. For example, the following may be used.
[0043] That is, the biological reaction analysis unit 12 analyzes the eye movement of each of the multiple people and generates a heat map showing the eye direction. The peculiar determination unit 13 compares the heat map generated by the biological reaction analysis unit 12 for the analysis subject with the heat map generated for other people, and determines whether the change in the biological reaction analyzed for the analysis subject is more peculiar than the change in the biological reaction analyzed for other people.
[0044] As described above, in this embodiment, the video of the video session is stored in the local storage of the user terminal 10, and the above-described analysis is performed on the user terminal 10. Although it may depend on the machine specifications of the user terminal 10, it is possible to analyze the video information without providing it to an external party.
[0045] <Functional configuration example 2> As shown in FIG. 5, the video session evaluation system of this embodiment may include, as functional components, a moving image acquisition unit 11, a biological reaction analysis unit 12, and a reaction information presentation unit 13a.
[0046] The reaction information presenting unit 13a presents information indicating changes in biological reactions analyzed by the biological reaction analyzing unit 12a, including participants not displayed on the screen. For example, the reaction information presenting unit 13a presents information indicating changes in biological reactions to a leader, facilitator, or manager of the online session (hereinafter collectively referred to as the organizer). The organizer of the online session may be, for example, a lecturer of an online class, a chairperson or facilitator of an online conference, or a coach of a session for coaching purposes. The organizer of the online session is usually one of multiple users participating in the online session, but may also be a different person who does not participate in the online session.
[0047] In this way, the host of an online session can grasp the status of participants who are not displayed on the screen in an environment where an online session is being held with multiple people.
[0048] <Functional configuration example 3> Fig. 6 is a block diagram showing an example of the configuration according to this embodiment. As shown in Fig. 6, in the video session evaluation system of this embodiment, functions similar to those in the first embodiment described above are given the same reference numerals and descriptions thereof may be omitted.
[0049] The system according to this embodiment includes a camera unit that captures video footage of the video session, a microphone unit that captures audio, an analysis unit that analyzes and evaluates the video, an object generation unit that generates a display object (described later) based on information obtained by evaluating the acquired video, and a display unit that displays both the video footage of the video session and the display object while the video session is being executed.
[0050] As explained above, the analysis unit includes a video image acquisition unit 11, a biological reaction analysis unit 12, a peculiar determination unit 13, a related event identification unit 14, a clustering unit 15, and an analysis result notification unit 16. The functions of each element are as described above.
[0051] 7, the object generation unit, based on the analysis result of the video acquired from the video session by the analysis unit, displays an object 50 indicating the recognized face portion and information 100 indicating the analyzed and evaluated content as necessary, superimposed on the video. When multiple faces appear in the video, the object 50 may identify and display the faces of all of the multiple people.
[0052] Furthermore, even if the camera function of the video session is disabled on the other party's device (i.e., the camera is disabled by software within the video session application, rather than by physically covering the camera, etc.), if the other party's face is recognized by the other party's camera, the object 50 or the object 100 may be displayed in the area where the other party's face is located. This allows both parties to confirm that the other party is in front of the device even if the camera function is turned off. In this case, for example, the video session application may hide information acquired from the camera, while displaying only the object 50 or the object 100 corresponding to the face recognized by the analysis unit. Furthermore, the video information acquired from the video session and the information recognized and obtained by the analysis unit may be separated into different display layers, and the layer related to the former information may be hidden.
[0053] When there is an area for displaying multiple moving images, the object 50 or the object 100 may be displayed in all areas or only in a part of the areas. For example, as shown in Fig. 8, the object 50 or the object 100 may be displayed only in the moving image on the guest side.
[0054] The embodiments of the invention described above in Basic Configuration Example 1 to Basic Configuration Example 3 may be realized as a single device, or may be realized by a plurality of devices (e.g., cloud servers) partially or entirely connected via a network. For example, the control unit 110 and storage 130 of each terminal 10 may be realized by different servers connected to each other via a network. That is, this system includes user terminals 10 and 20, a video session service terminal 30 that provides two-way video sessions to the user terminals 10 and 20, and an evaluation terminal 40 that evaluates the video sessions. The following variations and combinations of configurations are possible: (1) All processing is done on the user's device As shown in Figure 8, by performing processing by the analysis unit on the terminal where the video session is being held, the analysis and evaluation results can be obtained simultaneously (in real time) with the time the video session is being held (although a certain amount of processing power is required). (2) Processing on the user terminal and the evaluation terminal 9, an analysis unit may be provided in an evaluation terminal connected via a network, etc. In this case, the video captured by the user terminal is shared with the evaluation terminal simultaneously with or after the video session, and after being analyzed and evaluated by the analysis unit in the evaluation terminal, information on objects 50 and 100 is shared with the user terminal together with or separately from the video data (i.e., information including at least the analysis data) and displayed on the display unit.
[0055] The following system is realized using each of the configurations of the above-described functional configuration examples 1 to 3 or a combination thereof.
[0056] <Embodiment> A video analysis system (hereinafter simply referred to as the "system") according to an embodiment of the present disclosure analyzes the reactions of participants in an online session based on video images obtained by filming all or specific participants in the session. The analysis may be performed regardless of whether the participants are displayed on the screen during the online session. For example, the system (analysis unit) according to this embodiment analyzes video images to statistically analyze and output information such as the amount and frequency of communication between users and their emotions at the time. Furthermore, the analysis unit analyzes not only the user's emotions but also the content of their comments based on the video images. The analysis of the content of such comments may be performed, for example, using known voice analysis techniques or natural language processing techniques for video images.
[0057] In order to analyze and analyse user responses, users need to record videos of themselves during actual online sessions, but they may not want to make the videos public to other companies. Even if video data is recorded and saved, users may not want to disclose the videos to third parties.
[0058] Therefore, in this embodiment, a system is realized that enables a user to analyze his or her own video data while preventing the video data in which the user appears from being disclosed to a third party.
[0059] Fig. 10 is a diagram showing an example of the functional configuration of the system according to this embodiment. The system shown in Fig. 10 includes a video processing unit 21, a presentation unit 22, and an output control unit 23. The video processing unit 21, the presentation unit 22, and the output control unit 23 can be realized by loading a program stored in a storage medium or the like provided in, for example, the user terminals 10 and 20 or the evaluation terminal 40 into a memory or the like and executing the program with a processor such as a CPU.
[0060] Video processing unit 21 has a function of processing data related to the video acquired from video acquisition unit 11 based on predetermined conditions so as not to include an image of the user included in the video. The video data referred to here may be video data generated at any time during an online session, or video data accumulated after the online session.
[0061] Fig. 11 is a diagram showing an example of a video that has been processed and output by the video processing unit 21 according to this embodiment. As shown in Fig. 11, a screen 1100 is a screen for displaying a review of the video of the online session, and a video 1101 showing each user who participated in the online session is displayed on the screen 1100. Note that video 1101 displays the video of all users who participated in the online session, but it may also display, for example, only at least one user who participated in the online session.
[0062] For example, the video processing unit 21 may perform processing to change the image of the user included in the video to an object corresponding to the image of the user and display it. Specifically, as in the display of User2 and User4 in the video 1101 shown in Fig. 11, the video processing unit 21 may process the video data so that the image of the user is replaced with an object such as a silhouette or character of the user. The type of processing performed by the video processing unit 21 is not particularly limited as long as it is processing to prevent the image of the user from being included.
[0063] Furthermore, the video processing unit 21 may extract only the audio included in the video and generate a new video. Specifically, the video processing unit 21 may process a video containing only the user's audio, as shown in the display of User3 in video 1101 in FIG. 11 .
[0064] Furthermore, the video processing unit 21 may process the video data based on input information acquired by a user's input to the user terminal. For example, if a user wants to prevent a video data image of the user from being displayed to a third party, the user inputs information to the user terminal 10, 20 to not allow the display of the user's image in the video data image of the user. The video processing unit 21 can acquire input information based on such input and process the video data.
[0065] The presentation unit 22 has a function of presenting to the user terminals 10, 20 information regarding the execution of processing related to the processing by the video processing unit 21. Such information may include, for example, information regarding whether or not the video data is to be processed, and information for executing the processing of the video data, such as information regarding which users the processed video data should be displayed to. The video processing unit 21 may perform processing based on input information acquired by the user inputting information into the user terminal in response to the information presented by the presentation unit 22. This allows the user to choose whether to make video data in which they appear public or private.
[0066] The output control unit 23 may have a function of outputting moving images to each of the user terminals 10 and 20. For example, when the moving image processing unit 21 processes moving image data to be displayed during an online session, the output control unit 23 outputs the processed moving images to each of the user terminals 10 and 20 participating in the online session.
[0067] Specifically, when the video processing unit 21 processes a video obtained by photographing a user, the output control unit 23 may output data of the processed video to user terminals of users other than the user. At this time, the processed video data may be output only to the user terminals of some of the other users. Alternatively, unprocessed video data may be output to the user terminal of the user. This allows users to check their own facial expressions and other information during an online session while preventing their image from being disclosed to third parties to whom they do not want it to be made public.
[0068] When analyzing video data stored after an online session, the video data may be video data that has not been processed by the video processing unit 21. That is, the video data that has been processed by the video processing unit 21 may be video data that is displayed to each user on the user terminals 10 and 20. The unprocessed video data may be stored separately in the evaluation terminal 40, a server, etc.
[0069] 12 is a flowchart showing an example of the processing flow of the system according to this embodiment. First, the video processing unit 21 acquires a video showing an image of one user in an online session (step S101). Then, the video processing unit 21 processes the acquired video so that the image of the user is not included in the video (step S103).
[0070] Next, the output control unit 23 can output the processed video data to each of the user terminals 10 and 20 (step S105).
[0071] As described above, according to one embodiment of the present disclosure, video data including an image of one user can be processed to remove the image, and the processed video data can be output to a device of another user. This allows for analysis of the user's facial expressions and other biometric responses, while video data that does not include the user's image can be delivered to the user's device. Therefore, for example, privacy can be protected by preventing others from seeing the user's image, and the results of the user's biometric response analysis can be obtained. Therefore, the user can participate in the online session with peace of mind while undergoing biometric response analysis.
[0072] The processes described herein using flowchart diagrams do not necessarily have to be performed in the order shown, some process steps may be performed in parallel, additional process steps may be employed, and some process steps may be omitted.
[0073] The above-described embodiments may be combined as appropriate. Furthermore, the effects described in this specification are merely descriptive or exemplary and are not limiting. In other words, the technology according to the present disclosure may achieve other effects that are apparent to those skilled in the art from the description of this specification, in addition to or in place of the above-described effects. [Explanation of symbols]
[0074] 10, 20 user terminals 21 Video Processing Department 22 Presentation section 23 Output control section 30 Video Session Service Terminals 40 Evaluation Devices
Claims
1. A video analysis system that analyzes reactions of a user based on video images obtained by photographing the user in an environment where an online session is held by a plurality of users, regardless of whether the user is displayed on a screen during the online session, comprising: a video acquisition unit that acquires, for each of the plurality of users, a video obtained by photographing the user during the online session; an analysis unit that analyzes changes in biological reactions of the user based on the moving images acquired by the moving image acquisition unit; a video processing unit that processes data relating to the video based on a predetermined condition so that the video does not include an image of the user; an output control unit that outputs the video to a terminal of each user participating in the online session; Equipped with the moving image processing unit processes a moving image obtained by photographing one user based on input information acquired by the user's input to the terminal, and the input information specifies to which user the processed moving image is to be displayed; A video analysis system in which the output control unit outputs video data processed by the video processing unit to the terminal of a user specified in the input information, outputs video data that has not been processed by the video processing unit to the terminal of a user not specified in the input information, and outputs video data that has not been processed by the video processing unit to the terminal of one of the users.
2. The video analysis system according to claim 1, a presentation unit that presents information about the execution of processing by the video processing unit to the user terminal; A video analysis system in which the video processing unit performs the processing based on input information obtained by the user inputting information into a terminal in response to the information presented by the presentation unit.
3. The video analysis system according to claim 1, The video processing unit extracts only the audio contained in the video and generates it as a new video.
4. The video analysis system according to claim 1, The video processing unit processes the video to display an object corresponding to an image of a user included in the video in place of the image of the user.
5. 1. A video analysis method for analyzing a user's reaction based on a video obtained by photographing the user in an environment where an online session is held by a plurality of users, regardless of whether the user is displayed on a screen during the online session, comprising: acquiring, for each of the plurality of users, a video image obtained by photographing the user during the online session; analyzing a change in a biological reaction of the user based on the moving image acquired in the step of acquiring the moving image; processing the data relating to the moving image based on a predetermined condition so as not to include an image of the user included in the moving image; outputting the video to a terminal of each user participating in the online session; Run In the processing step, processing is performed on a moving image obtained by shooting one user based on input information acquired by input to a terminal by the user, and the input information specifies to which user the processed moving image is to be displayed; In the outputting step, data of the moving image processed in the step of processing the moving image is output to a terminal of a user specified in the input information, data of the moving image that has not been processed in the step of processing the moving image is output to a terminal of a user not specified in the input information, and data of the moving image that has not been processed in the step of processing the moving image is output to a terminal of the one user. Video image analysis methods.
Citation Information
Patent Citations
Feeling analyzing system
JP2000076421A
Conference terminal device, display control method, and display control program
JP2010213133A
Expression change analysis system
JP2011154665A
Emotion estimation device and emotion estimation method
JP2011186521A
Face expression amplification device, expression recognition device, face expression amplification method, expression recognition method and program
JP2012008949A