Information processing system, information processing method, and information processing program

The system tracks user gaze and facial expressions to automatically identify and address emotional responses in electronic content, improving user engagement by providing targeted supplementary information.

JP2026023749APending Publication Date: 2026-02-13AKITA UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2024125916
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-08-01
Publication Date
2026-02-13

AI Technical Summary

Technical Problem

Existing systems fail to accurately identify specific objects within electronic content that evoke user emotions, requiring manual interaction which disrupts concentration.

Method used

An information processing system that uses a camera, eye tracker, and machine learning to track user gaze and facial expressions, embedding gaze features in frame images to classify emotional responses and display relevant content or supplementary explanations automatically.

Benefits of technology

Accurately identifies objects causing user emotions, allowing seamless display of supplementary information without disrupting reading flow, enhancing user engagement and understanding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026023749000001_ABST
    Figure 2026023749000001_ABST
Patent Text Reader

Abstract

To provide an information processing system, method and program for accurately specifying an object about which a user has some feeling from among displayed contents.SOLUTION: An information processing system 1 includes a camera that captures a moving image of a face of a user who views content, an eye tracker that tracks a line of sight of the user and outputs line-of-sight information, and an information processing apparatus including an image generation unit that generates a plurality of frame images constituting the moving image in which the face of the user is captured based on an image signal output from the camera, a coordinate point acquisition unit that acquires a coordinate point on the content to which the user directs the line of sight in a time-series order based on the line-of-sight information output from the eye tracker, an embedding processing unit that embeds the line-of-sight feature amount in each frame image, and an analysis unit that classifies a facial expression of the user by analyzing the plurality of frame images in which the line-of-sight feature amount is embedded using a machine learning model.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an information processing system, an information processing method, and an information processing program. [Background technology]

[0002] In recent years, the use of personal computers (PCs) and tablet devices has been promoted in classes at schools and universities, and students (hereinafter referred to as users) have many opportunities to view content displayed on screens, such as electronic texts and videos. Users have various feelings about text and image information, not just electronic texts, such as finding it difficult, interesting, boring, or fun.

[0003] As techniques relating to estimation of a user's feelings and emotions, for example, Patent Documents 1 to 4 are known. Patent Document 1 discloses a system that selects the most suitable book and appropriate advice information for each contracted user and individually delivers the book along with the advice information. This system determines whether the user is satisfied or not based on image information of the user's face photo.

[0004] Patent Document 2 discloses a viewer emotion assessment device that diagnoses the emotions (affect and feelings) of a viewer who directs his or her gaze toward a scene including a specific object toward the specific object. This viewer emotion assessment device analyzes the viewer's emotions toward the specific object based on the position of the viewpoint in the viewed image calculated from the viewed image and eye movement image, changes in the viewer's physiological response data and the accompanying acceleration, and changes in various parts of the viewer's face, and diagnoses both the emotions and feelings.

[0005] Patent Document 3 discloses a teaching support system that makes it easier for teachers to understand the emotions of students attending class via a camera. The system acquires a first captured image of the teacher's lesson and displays it to the students, while also acquiring a second captured image of the student and displaying it to the teacher, estimates the student's emotions, and associates an icon indicating the estimated student's emotions with the second captured image and displays it to the teacher.

[0006] Patent Document 4 discloses an information processing device for proposing to a user purchase conditions adjusted according to the type and level of the user's emotion without allocating human resources. This information processing device acquires first biometric information of the user acquired by a first device, and acquires second biometric information of the user acquired by a second device, which is different from the first biometric information, and recognizes the user's first emotion value based on the first biometric information and the second biometric information.

[0007] Non-Patent Document 1 discloses a technique for detecting unknown words using a webcam, in which the position of unknown words is identified by tracking the learner's gaze and applying a transducer-based machine learning model that encodes text information. [Prior art documents] [Patent documents]

[0008] [Patent Document 1] Japanese Patent Publication No. 2022-62732 [Patent Document 2] International Publication No. 2011 / 042989 [Patent Document 3] Japanese Patent Application Publication No. 2023-58477 [Patent Document 4] Japanese Patent Application Publication No. 2023-59435

[0009] [Non-Patent Document 1] Jiexin Ding et al., “Gaze Reader: Detecting Unknown Word Using Webcam for English as a Second Language (ESL) Learners, arXiv,<URL:https: / / doi.org / 10.48550 / arXiv.2303.10443> , arXiv:2303.10443v1 [cs.HC] ,Submitted on March 18, 2023 Summary of the Invention [Problem to be solved by the invention]

[0010] Some electronic textbooks have a function that allows users to embed supplementary explanations in terms and display them by selecting the term. However, to use such a function, users must manually move the mouse over the target term or tap on the screen. This can cause users to lose concentration while reading the text.

[0011] Therefore, if it were possible to automatically detect terms or sentences that users find difficult to read from within electronic text, it would be possible to display supplementary explanations without breaking the user's concentration.

[0012] However, Patent Document 1 only judges the satisfaction level based on the user's facial expression, and does not disclose a method for estimating the user's feelings or identifying the target in the content that evokes those feelings.

[0013] In Patent Document 2, the viewer's emotions toward an object are analyzed based on various data such as visual images, eye movement images, and physiological response data, but the specific processing of a discrimination method that integrates and uses these data is not disclosed.

[0014] Patent Document 3 allows teachers to understand students' emotions by displaying an icon indicating the student's emotion in association with the second captured image. Patent Document 3 estimates the student's emotion based on the student's eye movement, the student's biometric information, and the student's response to a question from the teacher, but does not disclose the process for making the estimation. Furthermore, Patent Document 3 does not identify the object (e.g., a term in a text) about which the student is feeling some emotion.

[0015] In Patent Document 4, a machine learning algorithm is used to determine the type and level of a user's emotion, but the object (e.g., a term in a text) about which the user is feeling emotion is not identified.

[0016] In Non-Patent Document 1, in view of the insufficient accuracy of eye tracking using a webcam, gaze information and text information are used to improve the accuracy of detecting unknown words. However, Non-Patent Document 1 does not estimate the user's feelings, and does not identify the target about which the user has some kind of feeling.

[0017] The present invention has been made in consideration of the above, and aims to provide an information processing system, an information processing method, and an information processing program that can accurately identify, from among displayed content, an object about which a user has some kind of emotional attachment. [Means for solving the problem]

[0018] In order to solve the above problem, an information processing system that is one aspect of the present invention comprises a camera configured to capture a video of a user's face as they view content displayed on a screen or in space and output an image signal; an eye tracker configured to track the user's gaze and output gaze information representing the direction of the user's gaze; an image generation unit configured to generate a plurality of frame images that constitute a video that shows the user's face based on the image signal output from the camera; a coordinate point acquisition unit configured to acquire, in chronological order, coordinate points on the content where the user was looking at at the time each frame image was captured based on the gaze information output from the eye tracker; a feature acquisition unit configured to acquire gaze features that represent the characteristics of the user's gaze; an embedding processing unit that embeds the gaze features in each of the plurality of frame images; and an analysis unit configured to classify the facial expressions of the user that appear in the plurality of frame images by analyzing the plurality of frame images with the embedded gaze features using a machine learning model.

[0019] The information processing system may further include a monitor that displays the content, and a display control unit that causes the monitor to display a display based on the classification result of the user's facial expression by the analysis unit.

[0020] In the above information processing system, the analysis unit may classify the user's facial expressions shown in each frame image into normal faces and expressions other than normal faces, and the display control unit may display a specific mark on the monitor in an area including the coordinate point of the gaze corresponding to the frame image classified as an expression other than the normal face.

[0021] In the above information processing system, the content may include text, the analysis unit may classify the user's facial expressions shown in each frame image into normal faces and expressions other than normal faces, and the display control unit may display on the monitor supplementary explanations regarding a series of text including gaze coordinate points corresponding to the frame images classified as expressions other than normal faces.

[0022] In the information processing system, the analysis unit may classify the facial expression of the user captured in each frame image into a normal face and an unreadable face. In the information processing system, the gaze feature amount may be an amount corresponding to a moving distance of the user's gaze between frames.

[0023] In the information processing system, the gaze feature amount may be a value based on a movement distance, a movement angle, or a movement speed of the gaze from a previous stop point to the current frame. In the information processing system, the gaze feature amount may be a value representing a gaze movement direction, or a value representing a change in the gaze movement direction between frames.

[0024] In the information processing system, the embedding processing unit may rewrite pixel values ​​in a partial region of each of the plurality of frame images with the gaze feature amount.

[0025] Another aspect of the present invention is an information processing method including: an image generation step of generating a plurality of frame images constituting a video showing a user's face based on an image signal output from a camera configured to capture a video of the user's face as they view content displayed on a screen or in a space and output an image signal; a coordinate point acquisition step of chronologically acquiring coordinate points on the content to which the user was looking at at the time each frame image was captured based on gaze information output from an eye tracker configured to track the user's gaze and output gaze information representing the direction of the user's gaze; a feature acquisition step of acquiring gaze features representing characteristics of the user's gaze; an embedding processing unit that embeds the gaze features in each of the plurality of frame images; and an analysis step of classifying the facial expressions of the user shown in the plurality of frame images by analyzing the plurality of frame images with the embedded gaze features using a machine learning model.

[0026] An information processing program that is yet another aspect of the present invention causes a computer to execute the following steps: an image generation step that generates a plurality of frame images that constitute a video that shows a user's face, based on an image signal output from a camera configured to capture a video of the user's face as they view content displayed on a screen or in a space and output an image signal; a coordinate point acquisition step that chronologically acquires coordinate points on the content that the user was looking at at the time each frame image was captured, based on gaze information output from an eye tracker configured to track the user's gaze and output gaze information that represents the direction of the user's gaze; a feature acquisition step that acquires gaze features that represent the characteristics of the user's gaze; an embedding processing unit that embeds the gaze features in each of the plurality of frame images; and an analysis step that classifies the facial expressions of the user that appear in the plurality of frame images by analyzing the plurality of frame images with the embedded gaze features using a machine learning model. [Effects of the Invention]

[0027] According to the present invention, it is possible to accurately identify, from among the displayed content, an object about which the user has some kind of emotional feeling. [Brief explanation of the drawings]

[0028] [Figure 1] 1 is a schematic diagram illustrating a schematic configuration of an information processing system according to an embodiment of the present invention. [Figure 2] 1 is a block diagram showing a schematic configuration of an information processing system according to an embodiment of the present invention. [Figure 3] 4 is a flowchart illustrating an operation of the information processing system according to the embodiment of the present invention. [Figure 4] FIG. 10 is a schematic diagram for explaining an example of processing in a coordinate point acquisition step. [Figure 5] 10 is a table illustrating information acquired in a coordinate point acquisition step. [Figure 6] FIG. 10 is a schematic diagram for explaining an example of processing in a feature amount obtaining step. [Figure 7] FIG. 10 is a schematic diagram for explaining an example of processing in an embedding step. [Figure 8] FIG. 10 is a schematic diagram for explaining an example of processing in a display step. [Figure 9] FIG. 1 is a schematic diagram illustrating an experimental environment in an example. [Figure 10] 1A and 1B are schematic diagrams for explaining image data used in Examples and Comparative Examples. [Figure 11] 1 is a table showing combinations of datasets in cross-validation. [Figure 12] 1 is a graph showing experimental results in Examples and Comparative Examples. DETAILED DESCRIPTION OF THE INVENTION

[0029] Hereinafter, an information processing system, an information processing method, and an information processing program according to embodiments of the present invention will be described with reference to the drawings. Note that the present invention is not limited to these embodiments. In addition, in the description of each drawing, the same parts are denoted by the same reference numerals.

[0030] The drawings referred to in the following description merely show the shapes, sizes, and positional relationships in a schematic manner to enable the understanding of the contents of the present invention. That is, the present invention is not limited to the shapes, sizes, and positional relationships exemplified in each drawing. Furthermore, there may be parts in which the dimensional relationships and ratios differ between the drawings.

[0031] (Outline of information processing system) The information processing system according to this embodiment is a system for identifying an object about which a user has some kind of emotional feeling from among displayed content. The information processing system according to this embodiment can be applied to the display of two-dimensional content displayed on a monitor screen, or three-dimensional content displayed in virtual reality (VR), augmented reality (AR), or mixed reality (MR), but the following description will focus on the case where two-dimensional content is displayed on a screen.

[0032] Examples of user emotions include being unable to read characters (such as kanji), being unable to understand terms, finding a sentence difficult to read, finding it interesting, boring, or enjoying it. When a user has some emotion toward content, there is some slight change in the user's facial expression or eye movement compared to when the user does not have that emotion. Specifically, examples include frowning or staring when the content is difficult to read, and smiling or moving their eyes quickly when the content is interesting. The information processing system according to this embodiment accurately identifies the target about which the user has some emotion, based on moving images of the user's face that capture the user's facial expression over time and gaze information that indicates the gaze movement.

[0033] (Configuration of information processing system) Fig. 1 is a schematic diagram showing a general configuration of an information processing system according to an embodiment of the present invention. Fig. 2 is a block diagram showing a general configuration of the information processing system. As shown in Figs. 1 and 2, the information processing system 1 according to this embodiment includes a camera 3 that captures an image of the face of a user viewing content, an eye tracker 4, and an information processing device 5. The information processing system 1 may further include a monitor 2 having a screen 2a on which content is displayed, and an operation input unit 6.

[0034] The monitor 2 is, for example, a liquid crystal display or an organic EL display, and displays content on a screen 2a under the control of a display control unit 536, which will be described later.

[0035] The camera 3 is an imaging means including a solid-state imaging element such as a CCD image sensor or a CMOS image sensor. The camera 3 is configured to capture video of the face of the user 10 and output an image signal. For example, a general-purpose webcam can be used as the camera 3. Alternatively, a camera built into a laptop or tablet terminal can also be used. In this embodiment, the camera 3 is set above the monitor 2, but is not limited to this position as long as it is positioned so as to capture an image of the user 10's entire face.

[0036] The eye tracker 4 is a non-contact eye tracking (gaze measurement) device that measures gaze movement using, for example, the pupil-central corneal reflex (PCCR), and is configured to track the gaze of the user 10 and output gaze information that indicates the gaze direction of the user 10. In this embodiment, the eye tracker 4 is installed below the screen 2a, but is not limited to this location as long as it is capable of measuring the gaze of the user 10.

[0037] The information processing device 5 can be configured by a general-purpose computer such as a personal computer (PC), a notebook PC, a tablet terminal, etc. As shown in FIG. 2, the information processing device 5 includes an external interface 51, a storage unit 52, and a processor 53.

[0038] The external interface 51 is an interface that connects the information processing device 5 to external devices (e.g., monitor 2, camera 3, eye tracker 4) and communication lines (e.g., Internet lines) and sends and receives information between the external devices and communication lines.

[0039] The storage unit 52 is a computer-readable storage medium, such as a semiconductor memory such as a ROM or a RAM, or a hard disk. The storage unit 52 includes a program storage unit 521, a learning model storage unit 522, a content storage unit 523, an image data storage unit 524, and a detection data storage unit 525.

[0040] The program storage unit 521 stores an operating system program, a driver program, application programs that execute various functions, various parameters used during the execution of these programs, etc. Specifically, the program storage unit 521 stores an information processing program for identifying an object about which the user 10 is feeling some emotion from among the content displayed on the screen 2a, based on the image signal output from the camera 3 and the gaze information acquired by the eye tracker 4.

[0041] The learning model storage unit 522 stores a machine learning model for classifying the facial expressions of the user captured in the frame images. As the machine learning model, for example, a learning model for video image classification called "Video Swin Transformer" developed by researchers at Microsoft Research Asia (MSRA) can be used (see Liu, et al., "Video Swin Transformer", arXiv, vol. 2106.13230 (2021)).<URL:https: / / arxiv.org / abs / 2106.13230> ).

[0042] The content storage unit 523 stores data of the content to be displayed on the screen 2a. Examples of the content include electronic texts such as books and textbooks, electronic versions of newspapers and magazines, electronic comics, etc. The content may also include supplementary explanations of terms in the text.

[0043] The image data storage unit 524 stores image data (frame data) representing a frame image generated based on the image signal output from the camera 3.

[0044] Based on the gaze information output from the eye tracker 4, the detection data storage unit 525 stores the coordinate point on the screen 2a at which the user is looking in each frame image, and the feature amount calculated based on the coordinate point.

[0045] The processor 53 is configured using, for example, a CPU (Central Processing Unit) and a GPU (Graphics Processing Unit), and by reading various programs stored in the program storage unit 521, controls each unit of the information processing system 1 in an integrated manner and executes various calculation processes for identifying, from among the content displayed on the screen 2a, an object about which the user 10 has some kind of emotion. Functional units realized by the processor 53 include an image generation unit 531, a coordinate point acquisition unit 532, a feature acquisition unit 533, an embedding processing unit 534, an analysis unit 535, and a display control unit 536.

[0046] The image generation unit 531 performs image processing such as demosaicing, white balance processing, and gamma correction on the image signal input from the camera 3 to generate multiple frame images that make up a video showing the user's face. The multiple frame images can be formed in a raster format having RGB channels. Furthermore, the image generation unit 531 may perform pixel value standardization processing or the like as preprocessing of the image data to be loaded into a machine learning model.

[0047] Based on the gaze information output from the eye tracker 4, the coordinate point acquisition unit 532 acquires, in chronological order, the coordinate points on the content of the screen 2a at which the user was looking at when each frame image was captured.

[0048] The feature acquisition unit 533 acquires gaze feature quantities that represent the characteristics of the user's gaze. More specifically, the gaze feature quantities are numerical representations of the characteristics of the user's gaze movement. Examples of gaze feature quantities that can be used include values ​​such as gaze fixation time, saccade (rapid eye movement), gaze direction, and coordinates acquired by the coordinate point acquisition unit 532, as well as values ​​calculated based on these values. Specific examples of gaze feature quantities include the distance the user's gaze moves between frames, in other words, the distance between adjacent coordinate points in the time series acquired by the coordinate point acquisition unit 532, and values ​​obtained by multiplying these distances by a predetermined constant. Furthermore, the gaze feature quantity to be acquired may be one type only, or multiple types. For example, two types of gaze feature quantities, the gaze movement distance and the saccade size, may be acquired.

[0049] The embedding processing unit 534 executes a process of embedding a gaze feature amount in each of the plurality of frame images. When the feature amount acquisition unit 533 acquires a plurality of types of gaze feature amounts, various types of gaze feature amounts may be embedded in each frame image. Furthermore, the embedding processing unit 534 may perform a standardization process on the gaze feature amounts during the embedding process. For example, the gaze feature amounts may be converted so that the average value or standard deviation of the gaze feature amounts falls within a predetermined range.

[0050] The analysis unit 535 classifies the facial expressions of the user captured in the frame images by analyzing the multiple frame images in which gaze features are embedded using a machine learning model stored in the learning model storage unit 522. The analysis unit 535, for example, classifies the facial expressions of the user captured in each of the series of frame images into normal faces and expressions other than normal faces. Here, normal facial expressions refer to expressions that do not reflect any particular emotion. On the other hand, expressions other than normal facial expressions refer to expressions that reflect some emotion, such as an expression of a face that is difficult to read (a face that is difficult to read) or an expression of a face that is amused.

[0051] The display control unit 536 controls various displays on the monitor 2. For example, the display control unit 536 executes processes such as displaying content on the monitor 2 in a predetermined format, and displaying on the monitor 2 a display based on the classification result of the user's facial expression by the analysis unit 535.

[0052] The operation input unit 6 is an input device such as a keyboard, a mouse, or a touch panel provided on the screen 2a, and inputs to the processor 53 a signal corresponding to an operation performed from the outside.

[0053] (Operation of information processing system) FIG. 3 is a flowchart showing the operation of the information processing system 1. First, the information processing system 1 reads out content from the content storage unit 523 in response to an operation on the operation input unit 6 (for example, an operation to select content), and displays the content on the screen 2a (step S101). Subsequently, the information processing system 1 starts video shooting with the camera 3 and gaze tracking with the eye tracker 4 (step S102).

[0054] The information processing device 5 generates frame images based on the image signals of the moving image output from the camera 3 (step S103: image generating step).

[0055] Next, the information processing device 5 synchronizes the gaze information output from the eye tracker 4 with the frame image, and acquires the coordinate point of the user's gaze at the timing when the frame image was captured (step S104: coordinate point acquisition step).

[0056] FIG. 4 is a schematic diagram for explaining an example of processing in the coordinate point acquisition step. FIG. 5 is a table illustrating information acquired in the coordinate point acquisition step. As shown in FIGS. 4 and 5, the information processing device 5 (coordinate point acquisition unit 532) acquires a coordinate point P(k)=(x k ,y k ),P(k+1)=(x k+1 ,y k+1),P(k+2)=(x k+2 ,y k+2 ), ... are obtained and linked to the frame images F(k), F(k+1), F(k+2) taken at the same time as each coordinate point was obtained.

[0057] If the imaging frame rate of the camera 3 and the frequency of acquisition of gaze information by the eye tracker 4 differ, data when the two are synchronized can be extracted appropriately. For example, if the acquisition frequency of gaze information is higher than the imaging frame rate, gaze information can be extracted in accordance with the imaging frame rate. Conversely, if the imaging frame rate is higher, image data can be extracted in accordance with the acquisition frequency of gaze information.

[0058] Next, the information processing device 5 acquires a gaze feature amount at each coordinate point (step S105: feature amount acquiring step). FIG. 6 is a schematic diagram for explaining an example of processing in the feature amount acquiring step. In the following, an example will be explained in which the movement distance of the user's gaze between frames is used as the gaze feature amount C(k). The gaze feature amount C(k) can be acquired based on the distance |P(k+1)-P(k)| between adjacent coordinate points in a time series. The gaze feature amount C(k) may be the distance (pixels) between the coordinate points on the screen 2a itself, or may be a value obtained by multiplying the distance between the coordinate points by a predetermined constant.

[0059] Next, the information processing device 5 embeds the corresponding gaze feature C(k) in each frame image (step S106: embedding step). As shown in FIG. 5, for example, for frame image F(k), a gaze feature C(k) representing the distance from a gaze coordinate point P(k) corresponding to frame image F(k) to a gaze coordinate point P(k+1) corresponding to the next frame image F(k+1) is embedded. The embedding process is performed by rewriting pixel values ​​in a partial region of frame image F(k) with the gaze feature. The region whose pixel values ​​are rewritten may be a specific column, row, or block, or may be any region. Furthermore, the region may be concentrated in one location of frame image F(k) (for example, the last column of the image) or may be distributed to multiple locations (for example, the four corners of the image). The region may be located on the periphery of frame image F(k) so as not to overlap with the user's face captured in frame image F(k). The embedding process may be performed on each RGB channel image or on the composite channel.

[0060] 7 is a schematic diagram for explaining an example of processing in the embedding step. As shown in FIG. 7, the information processing device 5 (embedding processing unit 534) divides the frame image F(k) into each channel of RGB, and generates an R channel image F(k). R , G channel image F(k) G , and the B channel image F(k) B For each of the frame images F(k), the pixel values ​​of the last row are rewritten (overwritten) with the gaze feature C(k) corresponding to the frame image F(k). For example, in FIG. 7, the image F(k) of each RGB channel constituting the frame image F(k) R ,F(k) G ,F(k) B The last line of the above is rewritten to the gaze feature C(k)=3. Also, for the frame image F(k+1), the image F(k+1) of each RGB channel is R ,F(k+1) G ,F(k+1) B The last line of the above is rewritten as the gaze feature C(k+1)=9.

[0061] Next, the information processing device 5 classifies the user's facial expression by analyzing a series of frame images (analysis target images) in which gaze feature amounts are embedded using a machine learning model for video classification (step S107: analysis step). In detail, the information processing device 5 (analysis unit 535) inputs chronologically consecutive analysis target images into the machine learning model as a group of data, and determines the user's facial expression captured in the data. At this time, the analysis unit 535 may classify the user's facial expression into a normal face and an expression other than a normal face (for example, an unreadable face).

[0062] Subsequently, the information processing device 5 displays the analysis results on the screen 2a (step S108: display step). FIG. 8 is a schematic diagram for explaining an example of processing in the display step. As an example, as shown in FIG. 8(a), a specific mark a1 may be displayed in an area including the coordinate point of the gaze on the content corresponding to (i.e., linked to) a frame image determined to be other than a normal face (e.g., a face that is difficult to read). This makes it possible to determine that the portion (e.g., a term) where the mark a1 is displayed is difficult to read for the user.

[0063] As another example, if the content displayed on screen 2a includes text and supplementary explanations about terms in the text are embedded in the content, a window a2 containing supplementary explanations about the series of text (terms or sentences) that includes the coordinate point may be displayed, as shown in (b) of Figure 8. This allows the user to continue reading the content by referring to the supplementary explanations that are displayed as needed.

[0064] (Verification experiment) An experiment was conducted to verify the accuracy of emotion estimation by the information processing system 1 according to the embodiment of the present invention.

[0065] 1. Experimental environment Fig. 9 is a schematic diagram for explaining the experimental environment. As shown in Fig. 9, in the experiment, a monitor 2 and a keyboard as an operation input unit 6 were placed on a desk, a camera 3 was attached to the top of the monitor 2, and an eye tracker 4 was attached to the bottom of the monitor 2. Then, a user 10 was seated in a chair so that the distance L between the monitor and the face of the user 10 was approximately 50 cm.

[0066] 2. Equipment used We used a Logitech C920 webcam (manufactured by Logitech Corporation) as camera 3. The specifications are as follows: Resolution: Full HD 1920 x 1080 pixels Frame rate: 30fps Viewing angle (diagonal): 78° In the experiment, only the central 896 x 896 pixel portion of the 1920 x 1080 pixel frame image, which contained the user's face, was extracted and used.

[0067] The Tobii Eye Tracker 4C (manufactured by Tobii, Inc., USA) was used as eye tracker 4. The gaze information output is 90 Hz. Therefore, gaze information was extracted from the gaze information output by eye tracker 4 at a frequency of 30 Hz to match the frame rate of camera 3, which is 30 fps.

[0068] 3. Data acquisition method Twelve subjects in their twenties whose native language is Japanese and who use PCs on a daily basis were asked to silently read the text displayed on Monitor 2.

[0069] Monitor 2 displayed one page of 10 different sentences, each of which was created by quoting reading questions from the Level 2 Japanese Kanji Aptitude Test. The difficulty level of the sentences was set to the most difficult level (high school graduate, university, general level) that could be read silently without skipping, based on a preliminary survey. The sentences were displayed in the following format: Number of characters: 42.7 (average number of characters per sentence) Font: MS Mincho, monospaced Font size: Actual measurement: approx. 2.2cm square (viewing angle: approx. 2.5°)

[0070] The subject was filmed silently with a camera to obtain moving image (frame image) data, and the subject's gaze was tracked with an eye tracker 4 to obtain coordinate data of the subject's gaze on the screen.

[0071] Additionally, subjects were asked to press the space bar on their keyboard when they perceived a character as difficult to read while reading silently. The frame image captured when the space bar was pressed was then classified as an unreadable face image, and the other frame images were classified as normal face images. In this way, frame images classified based on the subject's judgment while reading silently were obtained as correct answer data. The obtained correct answer data was used as training data and test data for the machine learning model.

[0072] 5. Verification method A machine learning model was trained using a portion of the correct answer data as training data for the facial expression classification methods according to the following Examples and Comparative Examples 1 and 2. Another portion of the correct answer data was used as test data to classify faces captured in frame images into difficult-to-read faces and steady faces, and the accuracy (accuracy rate) of identifying difficult-to-read faces was calculated. Furthermore, cross-validation was performed on the identification accuracy of Examples and Comparative Examples 1 and 2. In this verification experiment, a difficult-to-read face refers to an expression in which a user is silently reading difficult-to-read characters, and a steady face refers to an expression in which a user is silently reading characters other than difficult-to-read characters.

[0073] The learning conditions for the machine learning models are as follows: Number of epochs: The number of epochs that showed the highest classification accuracy (precision) for the test data. Learning rate: default value Batch size: Maximum value that can be set

[0074] (1) Example Using the method described in the above embodiment, gaze features were calculated based on the coordinate points of the gaze and embedded in the frame images. The facial expressions of the subjects captured in the frame images with the embedded gaze features were then classified into difficult-to-read faces and normal faces using the above-mentioned machine learning model "Video Swin Transformer." Note that when embedding the gaze features, standardization processing was performed so that the average value of the gaze features for each subject was 0 and the standard deviation was 1.

[0075] (2) Comparative Example 1 Video images (frame images) showing the subjects' faces were classified into difficult-to-read faces and steady faces using the above-mentioned machine learning model "Video Swin Transformer." Unlike the examples, gaze information was not used.

[0076] (3) Comparative Example 2 Still images of the subjects' faces were classified into difficult-to-read faces and normal faces using the machine learning model "POSTER (Pyramid Cross-Fusion Transformer Network)." Unlike the examples, gaze information was not used. Here, "POSTER" is a deep learning model that recognizes facial expressions from still images (see Zheng, et al., "POSTER: A Pyramid Cross-Fusion Transformer Network for Facial Expression Recognition," arXiv, Vol. 2204.04083 (2022)).<URL:https: / / arxiv.org / abs / 2204.04083> ).

[0077] FIG. 10 is a schematic diagram for explaining the image data used in the example and comparative examples 1 and 2. In FIG. In Example and Comparative Example 1, a predetermined number of frames (10 frames in this experiment) of consecutive frames of the same facial expression (difficult to read face or normal face) from the correct answer data were treated as a series of moving images. Specifically, as shown in Fig. 10, frame images No. 1 to No. 30 were acquired, and frames No. 1 to No. 13 and No. 26 to No. 29 were classified as normal faces and frames No. 14 to No. 25 and No. 30 as difficult to read faces based on the subject's key operation. In this example, frames No. 1 to No. 10 were used as normal face moving image 1, frames No. 2 to No. 11 as normal face moving image 2, frames No. 3 to No. 12 as normal face moving image 3, and frames No. 4 to No. 13 as normal face moving image 4. Meanwhile, frames No. 14 to No. 23 were used as difficult to read face moving image 1, frames No. 15 to No. 24 as difficult to read face moving image 2, and frames No. 16 to No. 25 as difficult to read face moving image 3. However, if the number of steady face videos differs from the number of hard-to-read face videos, videos with the larger number are randomly removed to make the two numbers equal. In the example shown in Figure 10, one video is removed from steady face videos 1 to 4.

[0078] On the other hand, in Comparative Example 2, each frame image of the correct answer data was treated as a still image. That is, in the example shown in Fig. 10, a total of 17 images, frame Nos. 1 to 13 and 26 to 29, are used as still images of normal faces, and a total of 13 images, frame Nos. 14 to 25 and 30, are used as still images of hard-to-read faces. However, if the number of normal face images differs from the number of hard-to-read face images, images are randomly excluded from the one with the larger number of images to make both numbers equal. In the example shown in Fig. 10, four images are excluded from frame Nos. 1 to 13 and 26 to 29, which are normal faces.

[0079] FIG. 11 is a table showing the combination of datasets in cross-validation. In cross-validation, as shown in FIG. 11, datasets were created by having subjects silently read sentences 1 to 10. Among the datasets obtained, data obtained when a specific sentence (e.g., sentence 1) was silently read was used as test data, and data obtained when other sentences (e.g., sentences 2 to 10) were silently read were used as training data. In this experiment, 10 types of sentences were prepared, and 10 datasets were created. For each of these datasets, a machine learning model was trained using the training data, and inference was performed using the trained model, and the accuracy for the test data was calculated. Thus, 10 accuracies were calculated for each of the Example and Comparative Examples 1 and 2. Furthermore, the average accuracy and standard deviation were calculated for each of the Example and Comparative Examples 1 and 2. Note that by using such cross-validation, it is possible to determine the average accuracy and its variance for unknown data not included in the training data, and to calculate accuracy that takes into account the effectiveness of each method (Example and Comparative Examples 1 and 2) when the sentence changes.

[0080] Furthermore, a t-test with Bonferroni correction was performed on the results of the cross-validation. Here, the t-test is a testing method for investigating whether there is a difference between two groups of values. In this experiment, a t-test was performed on three combinations: Example and Comparative Example 1, Example and Comparative Example 2, and Comparative Example 1 and Comparative Example 2 (number of tests: 3). In this case, the significance level of 5% (0.05) was adjusted by Bonferroni correction to 0.0167 (=significance level 0.05 / number of tests: 3).

[0081] 7. Experimental Results The accuracy rate and standard deviation of the facial expressions classified in the Example and Comparative Examples 1 and 2 were calculated. FIG. 12 is a graph showing the experimental results for the Example and Comparative Examples 1 and 2. In FIG. 12, the vertical axis represents the accuracy rate. Regarding the accuracy rate, significant differences were observed in a paired t-test at a 5% significance level (after Bonferroni correction) between the Example and Comparative Example 1, between the Example and Comparative Example 2, and between Comparative Examples 1 and 2 (i.e., p-value<0.0167). This indicates that the Example has significantly improved classification accuracy compared to Comparative Examples 1 and 2.

[0082] From the above verification experiments, it was confirmed that the facial expression recognition method using moving images (Example, Comparative Example 1) can improve the recognition accuracy more than the facial expression recognition method using still images (Comparative Example 2). Furthermore, even in the case of using moving images (Example, Comparative Example 1), it was confirmed that the recognition accuracy can be improved by combining the moving images with gaze information (Example).

[0083] As described above, according to this embodiment, it is possible to accurately identify, from among the content displayed on the screen 2a, an object about which the user is feeling some kind of emotion. Specifically, in this embodiment, by synchronizing gaze information acquired by an eye tracker with image data of a moving image acquired by a camera, it is possible to identify the object at which the user is looking at when each frame image is captured. Furthermore, in this embodiment, gaze features are embedded in the frame images analyzed by the machine learning model, so that facial expression classification results that reflect the characteristics of the user's gaze movement can be obtained. This makes it possible to improve the accuracy of facial expression classification.

[0084] Furthermore, according to this embodiment, when mark a1 is displayed in an area on the content that includes the coordinate point of the gaze corresponding to a frame image classified as a difficult-to-read face, it is possible to automatically collect information regarding characters that are difficult for the user to read, etc.

[0085] Furthermore, according to this embodiment, when a window a2 of supplementary explanation regarding a series of text including gaze coordinate points corresponding to a frame image classified as a difficult-to-read face is displayed, the user can smoothly proceed through the content by referring to the supplementary explanation.

[0086] Such an embodiment of the present invention can be applied to the following fields. For example, when the information processing system 1 is applied to online classes, which are expected to become increasingly popular in the future, supplementary explanations can be automatically displayed individually for any part of the electronic text used in the class that the user finds difficult. This allows the user to resolve their questions on the spot without having to go through the trouble or effort of searching for the explanation themselves. This makes it possible to improve the user's understanding of the class without breaking their concentration. It also reduces the burden on teachers, as they no longer need to provide supplementary explanations to each individual user.

[0087] Furthermore, by collecting information on parts of the content that users find difficult, it is possible to provide individualized instruction to users (for example, supplementary lessons for students) and to gauge their level of understanding of the lessons. Furthermore, the information collected in this way can be used to revise the electronic textbook.

[0088] In the above-described embodiment of the present invention, electronic texts have been exemplified as content, but the information processing system 1 can also be used in cases where the content is materials displayed at online seminars for companies, materials shared at web meetings, etc. The information processing system 1 can also be applied to the management of employees' work during remote work. As a specific example, by estimating the workload (difficulty) and physical condition of employees based on their facial expressions, this can be useful for determining the need for assistance, improving work balance, and managing physical condition.

[0089] Furthermore, the embodiment of the present invention is not limited to two-dimensional content displayed on the screen 2a of the monitor 2, but can also be used for three-dimensional content by combining it with virtual reality (VR), augmented reality (AR), and mixed reality (MR) technologies. For example, in a virtual reality space, it is possible to identify a three-dimensional object that the user feels some kind of emotion about, thereby evaluating the object, or to display a sign to guide the user when the user is confused.

[0090] (Variation 1) In the above embodiment, the movement distance of the line of sight is used as the line of sight feature amount. However, various other values ​​can also be used as the line of sight feature amount. As an example, the characteristics of gaze movement can be expressed by the duration of gaze fixation and the magnitude of saccades. Here, a fixation is a state in which the gaze remains fixed on an object, and a saccade is a state in which the gaze moves instantaneously and rapidly from the previous fixation point to the next fixation point. In other words, if the coordinate points of the gaze obtained from the eye tracker at a predetermined frequency (e.g., 90 Hz) continuously remain at roughly the same position, it can be considered a fixation, and if they move, it can be considered a saccade.

[0091] Indicators representing a saccade include the distance (pixels) of gaze movement per unit time, the angle of movement obtained by converting this distance of movement into an angle (in other words, the angle at which the eyeball moves, or degrees), or the speed of movement (pixels / sec, degrees / sec) calculated based on the distance or angle of movement and unit time. In other words, the distance of gaze movement, the angle of movement, or the speed of movement or a value based on the speed from the previous stop point to the current frame can be used as the gaze feature to be embedded in the image.

[0092] Here, when the gaze feature value is the distance the gaze moves per frame as described in the above embodiment, the gaze feature value rarely becomes zero even when the gaze is stationary because the gaze is always moving slightly. In contrast, in this modified example, the gaze feature value when the gaze is stationary can be expressed as zero. Also, the absolute distance (or angle) the gaze moves from the previous stationary point can be used as the gaze feature value.

[0093] As another example, the characteristics of gaze movement can be expressed by the direction of gaze. Specifically, the direction in which the gaze moves during a saccade can be expressed as an azimuth angle (0° to 360°) on a two-dimensional plane, and the value obtained by normalizing this angle to a numerical range of 0 to 1 can be used as the gaze feature. Such gaze feature can be an index indicating that when the gaze direction is moving from left to right, the Japanese text is being read smoothly, and when the gaze direction is moving from right to left, the Japanese text is being reread or the reader has progressed to the next line.

[0094] As yet another example, the gaze feature may be a value obtained by comparing the gaze movement direction in the immediately preceding frame with the gaze movement direction in the current frame and normalizing the angle changed between them to a value ranging from 0 to 1. For example, if the gaze direction suddenly changes to a direction completely different from the previous direction, it is considered that a sudden change has occurred in the user's mood, and therefore such a gaze feature may be an index representing such a user's mood.

[0095] Alternatively, the frequency detected by frequency analysis of eye movement may be used as an index representing the frequency of saccades or the like as the gaze feature amount.

[0096] (Variation 2) In the above example, the user's facial expressions were classified into normal faces and difficult-to-read faces, but this embodiment can be applied to cases where the user's facial expressions are classified into various expressions, such as an interested face, a bored face, a sad face, etc.

[0097] For example, if a user is interested in a particular piece of text, they tend to read the text carefully and deliberately, slowly and thoroughly. In this case, the saccade distance becomes shorter, the frequency of fixations increases, and the angle of gaze movement becomes constant. Therefore, by using these indices as gaze feature quantities, it is possible to distinguish facial expressions of interest and identify the subject of the content in which the user is interested.

[0098] Furthermore, if a user feels bored, they may skip over parts of the text or read the text out of order. In such cases, the saccade distance increases, the frequency of fixations decreases, and the angle of gaze movement becomes inconsistent. Therefore, by using these indices as gaze features, it becomes possible to distinguish bored facial expressions, identify the content that the user finds boring, or determine whether the user is not interested in the content in the first place.

[0099] When a user feels sad, they may reminisce deeply about something. In this case, the user may gaze for a long time at a peripheral area unrelated to the text. Therefore, by using the presence or absence of a gaze target and the coordinate values ​​of the gaze as gaze features, it is possible to estimate the user's feelings.

[0100] The present invention is not limited to the above-described embodiment, and can be embodied in various other forms without departing from the spirit of the present invention. For example, some components may be removed from all of the components shown in the above embodiment, or other components may be appropriately combined. [Explanation of symbols]

[0101] 1...information processing system, 2...monitor, 2a...screen, 3...camera, 4...eye tracker, 5...information processing device, 6...operation input unit, 10...user, 51...external interface, 52...memory unit, 53...processor, 521...program memory unit, 522...learning model memory unit, 523...content memory unit, 524...image data memory unit, 525...detection data memory unit, 531...image generation unit, 532...coordinate point acquisition unit, 533...feature acquisition unit, 534...embedding processing unit, 535...analysis unit, 536...display control unit

Claims

1. a camera configured to capture a video of a user's face viewing content displayed on a screen or in a space and output an image signal; an eye tracker configured to track the user's gaze and output gaze information indicative of the user's gaze direction; an image generating unit configured to generate a plurality of frame images constituting a moving image showing the user's face based on the image signal output from the camera; a coordinate point acquisition unit configured to acquire, in chronological order, coordinate points on the content to which the user is directing his or her gaze at the timing when each frame image is captured, based on the gaze information output from the eye tracker; and a feature amount acquisition unit configured to acquire a gaze feature amount representing a feature of a user's gaze; an embedding processing unit that embeds the gaze feature amount in each of the plurality of frame images; an analysis unit configured to classify facial expressions of a user captured in the plurality of frame images by analyzing the plurality of frame images into which gaze feature amounts are embedded using a machine learning model; An information processing system comprising:

2. a monitor for displaying the content; a display control unit that displays on the monitor a display based on a classification result of the user's facial expression by the analysis unit; The information processing system of claim 1 , further comprising:

3. The analysis unit classifies the facial expressions of the user captured in each frame image into a normal face and a non-normal face, The information processing system according to claim 2 , wherein the display control unit causes the monitor to display a specific mark in an area including a coordinate point of a gaze corresponding to a frame image classified as an expression other than the steady face.

4. the content includes text; The analysis unit classifies the facial expressions of the user captured in each frame image into a normal face and a non-normal face, The information processing system according to claim 2 , wherein the display control unit causes the monitor to display supplementary explanations regarding a series of text including coordinate points of gaze corresponding to frame images classified as expressions other than the steady face.

5. The information processing system according to claim 4 , wherein the analysis unit classifies the facial expression of the user captured in each frame image into a normal face and an unreadable face.

6. 6. The information processing system according to claim 1, wherein the gaze feature amount is an amount corresponding to a moving distance of the user's gaze between frames.

7. The information processing system according to any one of claims 1 to 5, wherein the gaze feature amount is a value based on a movement distance, a movement angle, or a movement speed of the gaze from a previous stopping point to a current frame.

8. 6. The information processing system according to claim 1, wherein the gaze feature amount is a value representing a gaze movement direction or a value representing a change in the gaze movement direction between frames.

9. 6. The information processing system according to claim 1, wherein the embedding processing unit rewrites pixel values ​​in a partial area of ​​each of the plurality of frame images with the gaze feature amount.

10. an image generation step of generating a plurality of frame images constituting a moving image showing the user's face based on an image signal output from a camera configured to capture a moving image of the user's face viewing the content displayed on the screen or in the space and output the image signal; a coordinate point acquisition step of chronologically acquiring, based on gaze information output from an eye tracker configured to track the gaze of the user and output gaze information indicating the direction of the gaze of the user, coordinate points on the content to which the user was directing his / her gaze at the timing when each frame image was captured; a feature amount acquiring step of acquiring a gaze feature amount representing a feature of a user's gaze; an embedding processing unit that embeds the gaze feature amount in each of the plurality of frame images; an analyzing step of classifying facial expressions of the user captured in the plurality of frame images by analyzing the plurality of frame images into which gaze feature amounts are embedded using a machine learning model; An information processing method including:

11. an image generation step of generating a plurality of frame images constituting a moving image showing the user's face based on an image signal output from a camera configured to capture a moving image of the user's face viewing the content displayed on the screen or in the space and output the image signal; a coordinate point acquisition step of chronologically acquiring, based on gaze information output from an eye tracker configured to track the gaze of the user and output gaze information indicating the direction of the gaze of the user, coordinate points on the content to which the user was directing his / her gaze at the timing when each frame image was captured; a feature amount acquiring step of acquiring a gaze feature amount representing a feature of a user's gaze; an embedding processing unit that embeds the gaze feature amount in each of the plurality of frame images; an analyzing step of classifying facial expressions of the user captured in the plurality of frame images by analyzing the plurality of frame images into which gaze feature amounts are embedded using a machine learning model; An information processing program that causes a computer to execute the above.

Citation Information

Patent Citations

  • System that individually select most suitable book and appropriate advice information for each contracted user on one-off or regular basis and individually distribute book and attached advice information on one-off or regular basis

    JP2022062732A

  • Classroom support system, classroom support method, classroom support program

    JP2023058477A

  • Information processing device, method, and program

    JP2023059435A

  • Viewer's feeling determination device for visually-recognized scene

    WO2011042989A1