Model generation method and inference program

By training an inference model with eye-related data and optionally behavioral data, the method addresses the cost and complexity issues of brain activity-based methods, enabling accurate and affordable perception inference.

JP2025119824APending Publication Date: 2025-08-15NAT INST OF INFORMATION & COMM TECH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024014867
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-02-02
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

Existing methods for inferring perceived content, such as those based on brain activity measurements, are costly and cumbersome due to the use of large-scale devices like fMRI, making it difficult to infer perceptions in a simple and low-cost manner.

Method used

A method and program utilizing eye-related data, including gaze, blink, and pupil diameter data, to train an inference model that can infer semantic representations from content, optionally combined with behavioral data, to accurately represent perceptions.

Benefits of technology

Enables easy and cost-effective inference of perceived content by using eye-related data, potentially improving accuracy through the inclusion of behavioral data, thus overcoming the limitations of expensive brain activity measurement techniques.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025119824000001_ABST
    Figure 2025119824000001_ABST
Patent Text Reader

Abstract

To provide technique for easily inferring details that an individual perceives while viewing the content at low cost.SOLUTION: A model generation device according to one aspect of the present invention acquires eyeball-related data obtained by measuring an examinee viewing a content, uses the acquired eyeball-related data to perform machine learning on an inference model, and outputs the result of the machine learning. The machine learning includes training the inference model so as to acquire the capability of inferring a semantic expression in an information space corresponding to the details included in the content from the eyeball-related data.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a model generation method and an inference program. [Background technology]

[0002] How to quantify the effects of diverse content on diverse individuals and how to optimize content for diverse individuals are important issues in a wide range of fields, including education, entertainment, advertising, and policy.

[0003] With the rapid development of artificial intelligence technology, including large-scale language models, methods are being developed to quantitatively represent the content contained in content expressed in various modalities as fixed-length vectors in any latent space, such as image feature space, semantic space, concept space, etc. For example, Non-Patent Document 1 proposes an architecture for calculating continuous vector representations of words from a large-scale dataset.

[0004] Furthermore, with the recent advances in quantitative modeling techniques for brain activity, techniques for interpreting the content of perceptual experiences as vectors in a latent space are also being realized. For example, Patent Document 1 proposes a method for measuring brain activity and analyzing the measured brain activity to estimate perceived semantic content. [Prior art documents] [Patent documents]

[0005] [Patent Document 1] Japanese Patent Application Laid-Open No. 2016-195716 [Non-patent literature]

[0006] [Non-Patent Document 1] Tomas Mikolov, Kai Chen, Greg Corrado, Jeffrey Dean, “Efficient estimation of word representations in vector space”, [online], [Retrieved January 29, 2024] Internet<URL:https: / / arxiv.org / abs / 1301.3781> Summary of the Invention [Problem to be solved by the invention]

[0007] In the information analysis method based on brain activity proposed in Patent Document 1 and the like, a large-scale measuring device such as fMRI (functional magnetic resonance imaging) is used to measure brain activity. Large-scale measurement devices are expensive and restrictive, making it difficult to infer what individuals perceive in a simple, low-cost manner.

[0008] In one aspect, the present invention has been made in consideration of the above circumstances, and its purpose is to provide a technology for easily inferring what is perceived by an individual while viewing content at low cost. [Means for solving the problem]

[0009] In order to solve the above-mentioned problems, the present invention employs the following configurations, which can be combined as appropriate.

[0010] A model generation method according to one aspect of the present invention is an information processing method in which a computer executes the steps of acquiring eye-related data measured from a subject viewing content, performing machine learning of an inference model using the acquired eye-related data, and outputting the results of the machine learning. and training the inference model to acquire the ability to infer, from the eye-related data, a semantic representation in an information space corresponding to the appearance.

[0011] The present inventors have found through experimental examples described below that it is possible to infer from eye activity what is perceived by an individual while viewing content (i.e., eye-related data can be used as a signal source to replace brain activity data). Based on this finding, the present configuration generates a trained inference model (trained model) that has acquired the ability to infer what is perceived from content from eye-related data. Compared to brain activity, eye activity can be measured easily and at low cost. Therefore, with the present configuration, it is possible to use the generated trained model to infer what is perceived by an individual while viewing content easily and at low cost.

[0012] In the model generation method according to the above aspect, the computer may further execute a step of acquiring behavioral data indicating the behavior of the subject while viewing the content. The acquired behavioral data may be further used in machine learning of the inference model. In the machine learning, the inference model may be trained to acquire the ability to infer the semantic representation from the eye-related data and the behavioral data.

[0013] For example, behaviors while viewing content, such as searching for words that appear in the content or cheering for a player that appears in the content, may be triggered as a result of perceiving the content. In other words, because behaviors while viewing content are correlated with perceptual results, behavioral data of individuals while viewing content can provide clues for inferring perceived content. This configuration allows for the generation of a trained model that further accepts input of behavioral data. The generated trained model is expected to improve the accuracy of inferring what individuals perceive while viewing content by using the behavioral data together with eye-related data as explanatory variables.

[0014] In the model generation method according to the above aspect, the acquired eye-related data may be composed of gaze data, blink data, pupil diameter data, or a combination thereof. Gaze, blink data, and pupil diameter accurately represent eye activity. Therefore, according to this configuration, by using at least one of the gaze data, blink data, and pupil diameter data as eye-related data, it is possible to generate a trained model that has acquired the ability to appropriately infer perceptual content from eye activity.

[0015] Furthermore, aspects of the present invention may not be limited to the model generation stage. The present invention may also be directed to an inference stage that uses a trained inference model generated by the model generation method. For example, one aspect of the present invention may be an inference program that uses a trained inference model generated by the model generation method.

[0016] According to one aspect of the present invention, there is provided an inference program for causing a computer to execute the steps of: acquiring eye-related data measured from a subject viewing target content; inferring a semantic representation in an information space corresponding to content included in the target content from the acquired eye-related data using a trained inference model; and outputting a result of inferring the semantic representation. With this configuration, it is possible to infer the content perceived by an individual while viewing content simply and at low cost.

[0017] The inference program according to the above aspect may further cause the computer to execute a step of acquiring behavioral data indicating the behavior of the subject while viewing the target content. In the inferring step, the computer may infer the semantic expression from the acquired eye-related data and behavioral data. According to this configuration, We can expect to see an improvement in the accuracy of inferring perceived content.

[0018] Note that the present invention is not limited to the above-described model generation method. As another aspect of the model generation method according to the above-described aspects, the present invention may be an information processing device (model generation device) that realizes all or part of the above-described configurations, a program, or a storage medium readable by a machine such as a computer on which such a program is stored. A storage medium readable by a machine such as a computer is a medium that stores information such as a program by electrical, magnetic, optical, mechanical, or chemical action. Furthermore, the present invention is not limited to the above-described inference program. As another aspect of the inference program according to the above-described aspects, the present invention may be an information processing device (inference device) that realizes all or part of the above-described configurations, an information processing method (inference method), or a storage medium readable by a machine such as a computer on which a program is stored.

[0019] For example, a model generation device according to an aspect of the present invention may include a control unit. The control unit may be configured to execute the steps of acquiring eye-related data measured from a subject viewing content, performing machine learning on an inference model using the acquired eye-related data, and outputting a result of the machine learning. The machine learning may include training the inference model to acquire the ability to infer, from the eye-related data, a semantic representation in an information space corresponding to content included in the content.

[0020] For example, an inference device according to an aspect of the present invention may include a control unit configured to execute the steps of acquiring eye-related data measured from a subject viewing target content, inferring a semantic representation in an information space corresponding to content included in the target content from the acquired eye-related data using a trained inference model, and outputting a result of inferring the semantic representation.

[0021] For example, an inference method according to one aspect of the present invention may be an information processing method in which a computer executes the steps of acquiring eye-related data measured from a subject viewing target content, inferring a semantic representation in an information space corresponding to the content contained in the target content from the acquired eye-related data using a trained inference model, and outputting the result of inferring the semantic representation. [Effects of the Invention]

[0022] According to the present invention, it is possible to provide a technology for easily and at low cost inferring what is perceived by an individual while viewing content. [Brief explanation of the drawings]

[0023] [Figure 1] FIG. 1 shows a schematic diagram of an example of a situation in which the present invention is applied. [Figure 2] FIG. 2 is a diagram illustrating an example of a hardware configuration of the model generating device. [Figure 3] FIG. 3 is a diagram illustrating an example of the hardware configuration of the inference device. [Figure 4] FIG. 4 is a diagram illustrating an example of the software configuration of the model generating device. [Figure 5] FIG. 5 shows a schematic diagram of an example of the software configuration of the inference device. [Figure 6] FIG. 6 is a flowchart illustrating an example of a processing procedure of the model generating device. [Figure 7] FIG. 7 is a flowchart showing an example of a processing procedure of the inference device. [Figure 8] FIG. 8 shows the relationship between the principal component vectors and the decoding model obtained in the experimental example. [Figure 9] FIG. 9 shows the results of calculating the average value of the inference accuracy of the inference model for each dimension of the principal component vector in the experimental example. [Figure 10] Figure 10 shows the results of t-SNE representation of perceptual content inferred from pupil diameter. [Figure 11]Figure 11 shows the results of calculating the inter-individual change in inference accuracy of the coding model in the experimental example. DETAILED DESCRIPTION OF THE INVENTION

[0024] An embodiment according to one aspect of the present invention (hereinafter also referred to as "the present embodiment") will be described below with reference to the drawings. However, the present embodiment described below is merely an example of the present invention in all respects. Needless to say, various improvements and modifications can be made without departing from the scope of the present invention. In other words, when implementing the present invention, specific configurations according to the embodiment may be appropriately adopted. Note that, although data appearing in the present embodiment are described in natural language, more specifically, they are specified using computer-recognizable pseudo-language, commands, parameters, machine language, etc.

[0025] §1 Application Examples FIG. 1 schematically illustrates an example of a scenario in which the present invention is applied. The system according to this embodiment includes a model generation device 1 and an inference device 2. The model generation device 1 is one or more computers configured to generate a trained inference model 5 by controlling the implementation of machine learning. The inference device 2 is one or more computers configured to infer the perceptual content of a subject TU using the trained inference model 5.

[0026] The model generation device 1 according to this embodiment acquires training eye-related data 31 measured from a subject LU while the subject LU is viewing training content LC. The training content LC may also be referred to as training content. The model generation device 1 performs machine learning on an inference model 5 using the acquired eye-related data 31. The machine learning includes training the inference model 5 to acquire the ability to infer, from the eye-related data 31, a semantic representation 39 in an information space corresponding to the content included in the content LC. The semantic representation in the information space corresponds to the perceived content of the content. Therefore, this machine learning can generate a trained inference model 5 that has acquired the ability to infer the perceived content from the eye-related data measured while the subject is viewing the content. The model generation device 1 outputs the results of the machine learning. The generated trained inference model 5 may be provided to the inference device 2 in any manner and at any timing.

[0027] Meanwhile, the inference device 2 according to this embodiment acquires eye-related data 41 of a target person TU while they are viewing target content TC. Using a trained inference model 5, the inference device 2 infers a semantic representation 49 in information space corresponding to the content contained in the target content TC from the acquired eye-related data 41. This semantic representation 49 corresponds to the content perceived by the target person TU with respect to the target content TC. The inference device 2 outputs the result of inferring the semantic representation 49.

[0028] Experimental examples described below have revealed that it is possible to infer from eye activity what is perceived by an individual while viewing content. Based on this finding, the model generation device 1 according to this embodiment generates a trained inference model 5 that has acquired the ability to infer from eye-related data what is perceived from content. By using this trained inference model 5, it is possible to infer simply and at low cost what is perceived by an individual while viewing content. The inference device 2 according to this embodiment can infer simply and at low cost what is perceived by a target person TU while viewing target content TC.

[0029] (content) In one example, the content (content LC, target content TC) may be any information that can be viewed by a person. Viewing may be at least one of watching and listening. As long as it can be viewed, the data format of the content is not particularly limited and may be selected appropriately depending on the embodiment. For example, the content (content LC, target content TC) may be a text The content may be composed of text data, image data (video, still images), sound data, other numerical data, or a combination of these. As a typical example, the content may be video content including sound (e.g., video content for education / medical support, etc.). The training content LC and the target content TC may be the same or different.

[0030] (Eye-related data) In one example, the eye-related data (eye-related data 31, eye-related data 41) may be configured to indicate eye activity of a person (subject LU, target TU) while the person is viewing content. The activity may include a state. The eye activity may be expressed, for example, by gaze, blinking, pupil diameter, etc. As long as the eye-related data is configured in this manner, the configuration of the eye-related data is not particularly limited and may be selected appropriately depending on the embodiment.

[0031] In one example, the eye-related data (eye-related data 31, eye-related data 41) may be composed of gaze data, blink data, pupil diameter data, or a combination thereof. Gaze data, blink data, and pupil diameter accurately represent eye activity. Therefore, according to one example of this embodiment, the model generation device 1 can be expected to generate a trained inference model 5 that has acquired the ability to properly infer perceptual content from eye activity. Furthermore, by using this trained inference model 5 in the inference device 2, proper inference of the subject TU's perceptual content of the target content TC can be expected.

[0032] The eye-related data may be generated by measuring the activity of a person's eyeballs with any sensor. In a training data collection situation, a sensor SA may be used to appropriately collect training samples of the eye-related data 31. In an inference situation, a sensor SC may be used to acquire samples of the eye-related data 41. As long as the activity of the eyeballs can be measured, the type of sensor (sensor SA, sensor SC) is not particularly limited and may be appropriately selected depending on the embodiment. In one example, the sensor (sensor SA, sensor SC) may include an imaging device such as an infrared camera. The sensor (sensor SA, sensor SC) may further include a device that assists measurement, such as an infrared projector.

[0033] The data format of the eyeball-related data is not particularly limited and may be appropriately selected depending on the type of the sensor and the embodiment. Furthermore, at least a portion of the eyeball-related data may be composed of sensing data obtained by a sensor, or may be composed of analysis data (analysis results) obtained by analyzing the sensing data. The analysis may include, for example, feature extraction, calculation of feature amounts, etc. The feature extraction may include, for example, identifying the gaze direction, measuring blinks, measuring pupil diameter, etc. The feature amounts may include, for example, statistics such as maximum value, minimum value, median, mean value, variance, standard deviation, and n-th percentile value.

[0034] The model generation device 1 may acquire data directly from the sensor SA, or may acquire the data indirectly via another computer. When acquiring data directly from the sensor SA, the sensor SA may be connected to the model generation device 1. When acquiring data indirectly via another computer, the model generation device 1 may acquire the target data from the other computer via a network, a storage medium, etc. The model generation device 1 may acquire data measured in the past by the sensor SA, or may acquire data measured in real time by the sensor SA.

[0035] Similarly, the inference device 2 may acquire data directly from the sensor SC, or indirectly via another computer. When acquiring data directly from the sensor SC, the sensor SC may be connected to the inference device 2. When acquiring data indirectly via another computer, the inference device 2 may be connected to the other computer via a network, a storage medium, etc. The inference device 2 may acquire data measured in the past by the sensor SC, or may acquire data measured in real time by the sensor SC.

[0036] (Other data) In one example, data other than the eye-related data may be further used to infer the semantic expression. Figure 1 illustrates a situation in which behavioral data indicating behavior while viewing content is further used to infer the semantic expression, as an example of the other data.

[0037] 1 , the model generation device 1 may be further configured to acquire behavioral data 33 indicating the behavior of the subject LU while viewing the content LC. The acquired behavioral data 33 may be further used in machine learning of the inference model 5. Accordingly, in the machine learning, the inference model 5 may be trained to acquire the ability to infer a semantic representation 39 from the eye-related data 31 and the behavioral data 33.

[0038] As described above, actions taken while viewing content, such as searching for words that appear in the content or cheering for a player that appears in the content, may be triggered as a result of perceiving the content. In other words, because actions taken while viewing content are correlated with perception results, behavioral data of individuals viewing content can provide clues for inferring perceived content. Therefore, according to one example of the present embodiment, by generating a trained inference model 5 that further accepts input of behavioral data, it is expected that the accuracy of inferring semantic representations in the generated trained inference model 5 can be improved.

[0039] In addition, in one example, the inference device 2 may be configured to further acquire behavioral data 43 indicating the behavior of the subject TU while viewing the target content TC. In the inferring step, the inference device 2 may infer a semantic representation 49 from the acquired eye-related data 41 and behavioral data 43. According to one example of the present embodiment, by further taking the behavioral data 43 into consideration when inferring the semantic representation 49, it is possible to expect an improvement in the accuracy of the inference.

[0040] In one example, the behavioral data (behavioral data 33, behavioral data 43) may be configured to include information about behavior that may be related to the perception of content. As long as the behavioral data is configured in this manner, the configuration of the behavioral data is not particularly limited and may be determined appropriately depending on the embodiment. The behavioral data may be generated by measuring human behavior with any sensor. In a training data collection situation, a sensor SB may be used to appropriately collect training samples of the behavioral data 33. In an inference situation, a sensor SD may be used to acquire samples of the behavioral data 43.

[0041] As long as the sensors (SB, SD) can measure human behavior, the types of sensors (SB, SD) are not particularly limited and may be selected appropriately depending on the embodiment. In one example, the behavior of a subject observed while viewing content may include, for example, terminal operations (key operations, mouse operations, etc.), body movements (gestures, etc.), vocalizations (including speech), vital signs (body temperature, heart rate, pulse, blood pressure, electrocardiogram, electromyogram, sweating, activity level, etc.). Accordingly, the sensors (SB, SD) may include a computer, an imaging device, an acceleration sensor, an angular acceleration sensor, a gyro sensor, a motion capture, a microphone, a vital sensor, etc. The number of modalities in the behavioral data may be selected appropriately depending on the embodiment.

[0042] The information on the terminal operation may be acquired from any computer (user terminal) as appropriate. When collecting training data, the model generation device 1 is used as a user terminal, and the behavioral data 33 is configured to include information on the terminal operation, the model generation device 1 may be used as a computer (sensor SB) for acquiring the information on the terminal operation. Similarly, the inference device 2 , and may be used as a computer (sensor SD) for acquiring information on terminal operation. The vital sensor may include, for example, a thermometer, a heart rate monitor, a pulse rate monitor, a blood pressure monitor, an electrocardiograph, an electromyograph, a skin electrodermal response monitor, an activity monitor, etc. The vital sensor may be configured as a wearable terminal such as a smart watch.

[0043] The data format of the behavioral data is not particularly limited and may be appropriately selected depending on the type of the sensor and the embodiment. Furthermore, at least a portion of the behavioral data may be composed of sensing data obtained by a sensor, or may be composed of analysis data (analysis results) obtained by analyzing the sensing data. As with the eye-related data, the analysis may include, for example, feature extraction, feature amount calculation, etc. Feature extraction may include, for example, identifying operation content, gesture estimation, vital sign estimation, etc.

[0044] The model generation device 1 may acquire data directly from the sensor SB, or may acquire the data indirectly via another computer. When acquiring data directly from the sensor SB, the sensor SB may be connected to the model generation device 1. When acquiring data via another computer, the model generation device 1 may acquire the target data from the other computer via a network, a storage medium, or the like. The model generation device 1 may acquire data measured in the past by the sensor SB, or may acquire data measured in real time by the sensor SB. As described above, the model generation device 1 may also function as at least a part of the sensor SB.

[0045] Similarly, the inference device 2 may acquire data directly from the sensor SD, or may acquire it indirectly via another computer. When acquiring data directly from the sensor SD, the sensor SD may be connected to the inference device 2. When acquiring data via another computer, the inference device 2 may acquire the target data from the other computer via a network, a storage medium, etc. The inference device 2 may acquire data measured in the past by the sensor SD, or may acquire data measured in real time by the sensor SD. The inference device 2 may also serve as at least a part of the sensor SD.

[0046] The sensors that measure eye-related data (sensors SA and SC) and the sensors that measure behavioral data (sensors SB and SD) may be different or may at least partially overlap. When the sensors that measure eye-related data and the sensors that measure behavioral data at least partially overlap, the eye-related data may include at least a portion of the behavioral data. For example, when eye activity and human behavior are measured using an imaging device, the image data obtained by the imaging device may at least partially serve as both the eye-related data and the behavioral data.

[0047] Note that the manner in which semantic expressions are inferred is not limited to this example. In another example, data other than behavioral data may be used in addition to or instead of behavioral data to infer semantic expressions. In another example, semantic expressions may be inferred using only eye-related data.

[0048] (inference) In one example, inferring may include identifying and / or regressing, and inferring may include predicting the future, since content up to a certain time may have correlation with content from that time to a future time.

[0049] (Information space / semantic expression) In one example, the information space is configured to represent the contents included in the content by vectors. A vector may be composed of one or more numerical values. The data format of a vector may be selected arbitrarily. In one example, a vector may simply be numerical data. The content expressed in the information space may correspond to the meaning of a word, phrase, context, etc.

[0050] A phrase may be, for example, a word, and may correspond to a meaning at a point in time. A phrase may be directly included in the content, or may be indirectly derived from the content. A phrase indirectly derived from the content may include, for example, a phrase associated with the content, or a phrase derived by interpreting the content. A phrase may be represented by a vector point in information space.

[0051] A context may correspond to a sequence of meanings at two or more points in time, such as a sentence, a video scene, etc. Similar to a phrase, a context may be directly included in the content or may be indirectly derived from the content. A context may be represented by one or more vectors of points in an information space. In one example, the information space may be divided into clusters. Each cluster may be configured to represent a context. Each point in each cluster may be configured to represent a phrase. This allows semantic representations to be differentiated by context.

[0052] In one example, the semantic representation (semantic representation 39, semantic representation 49) may be a vector in an information space and indicate the content included in the content. The relationship between the semantic representation and the content of the content may be determined appropriately depending on the embodiment. In one example, the semantic representation may directly indicate the content of the content. For example, a numerical representation (such as a binary representation) of a phrase such as a word may be used as the semantic representation. In this case, in a machine learning scenario, a correct answer to the semantic representation 39 may be obtained directly from the content LC. In another example, the semantic representation may indirectly indicate the content of the content. For example, the semantic representation may be obtained by using any projection method that converts the content into a vector representation, such as the method proposed in Non-Patent Document 1 above. In this case, the information space may also be referred to as a latent space. In a machine learning scenario, a correct answer to the semantic representation 39 may be obtained by projecting the content LC using a predetermined projection method.

[0053] The projection method is not particularly limited and may be selected appropriately depending on the embodiment. For example, a trained machine learning model may be used to project the content onto the semantic representation. The machine learning model may have any configuration, such as a Transformer or a Vision Transformer. The projection method may be a known method. In one example, the content may be converted into a vector representation by a single projection. In another example, the content may be converted into a word (text) representation after undergoing an image-to-text conversion operation (image2txt). ) into a vector representation by multiple projections, such as by performing an operation (word2vec) to convert the vector into a vector representation in a distributed manner.

[0054] In an inference scenario, the semantic representation 49 corresponds to what the subject TU perceives with respect to the target content TC. As described above, the perceived content may include, for example, content perceived directly from the content, content indirectly derived from the content, etc. The content indirectly derived from the content may include, for example, content associated with the content, or content derived by interpreting the content. The perceived content may also include content predicted to appear at a future time based on the content up to the current time. In other words, perceiving may include predicting.

[0055] (inference model) The inference model 5 is configured to infer a semantic expression from a given input (e.g., eye-related data). As long as such inference processing can be performed, the configuration of the inference model 5 is not particularly limited and may be determined appropriately depending on the embodiment. The inference model 5 may be configured using any machine learning model.

[0056] The machine learning model is configured to have one or more calculation parameters that can be adjusted by machine learning. The one or more calculation parameters are used to calculate the desired inference (inference of semantic expression in this embodiment). The machine learning model may be configured, for example, by a neural network, a regression model, a decision tree model, a support vector machine, or other functional formulas (calculation models). The machine learning method may be selected appropriately depending on the machine learning model to be adopted.

[0057] In one example, the inference model 5 may include a neural network. The structure of the neural network is not particularly limited and may be determined appropriately depending on the embodiment. The structure of the neural network may be specified, for example, by the number of layers from the input layer to the output layer, the type of each layer, the number of nodes (neurons) included in each layer, and the connection relationships between the nodes in each layer. In one example, the neural network may include any mechanism such as a recurrent structure, a self-attention mechanism, or an autoregressive model. Furthermore, the neural network may include any layer such as a fully connected layer, a convolutional layer, a pooling layer, a deconvolutional layer, an unpooling layer, a normalization layer, a dropout layer, or a long short-term memory (LSTM). The neural network may include any type of model such as a diffusion model, a transformer model, or a generative model. The neural network may include a model capable of in-context learning, such as a large-scale language model (LLM) or a large-scale vision-language model (LVLM). In one example, by including a self-attention mechanism and an autoregressive model, the neural network can acquire the ability to perform in-context learning. The weights of the connections between the nodes included in the neural network and the thresholds of the nodes are examples of calculation parameters. The data format of the input and output of the inference model 5 is not particularly limited and may be selected appropriately depending on the embodiment.

[0058] Machine learning (i.e., training the inference model 5) involves adjusting (optimizing) the values of computational parameters using training samples. Typically, the model generation device 1 may perform supervised learning as a machine learning process using multiple training datasets, each of which is composed of a combination of input samples (training samples) and output samples (teacher signals, labels). The input samples are samples of input data such as the eye-related data 31 and behavioral data 33. The output samples are correct semantic representations 39 corresponding to the input samples. As described above, the correct semantic representations 39 may be obtained directly from the input samples, or by transforming the input samples using a predetermined projection method. In supervised learning, the values of computational parameters of the machine learning model may be adjusted so that the output obtained from the machine learning model when an input sample is given matches the corresponding output sample. However, the method of generating a trained model is not limited to this example and may be changed as appropriate depending on the embodiment. The training dataset is not limited to the above example and may be selected as appropriate depending on the embodiment. For example, data other than the above may be used as the training dataset when acquiring in-context learning ability, etc. Furthermore, the learning method does not have to be limited to supervised learning, and other methods such as unsupervised learning (including self-supervised learning) and reinforcement learning may also be used.

[0059] Furthermore, the input / output format of the inference model 5 is not particularly limited and may be determined appropriately depending on the embodiment. In one example, the input data may be provided to the inference model 5 as is, or may be provided after being preprocessed. In another example, the output of the inference model 5 may be configured to directly or indirectly indicate a semantic representation (semantic representation 39, semantic representation 49). When the output of the inference model 5 is configured to indirectly indicate a semantic representation, the semantic representation may be obtained by performing any information processing (such as interpretation processing) on the output of the inference model 5.

[0060] It should be noted that the machine learning process does not necessarily have to be executed within the model generation device 1. The model generation device 1 performing machine learning of the inference model 5 may include executing machine learning processing within the model generation device 1, and giving instructions to a computer other than the model generation device 1 to cause the other computer to execute machine learning processing. In the latter case, in one example, the model generation device 1 may be appropriately connected to the other computer, and may cause the other computer to execute machine learning processing while performing data communication.

[0061] Furthermore, experimental examples described below suggest that, although there are individual differences, semantic representations can be inferred even using trained models generated from eye-related data of others. Therefore, the subject LU in the training stage and the subject TU in the inference stage do not necessarily need to be the same person. The subject LU and the subject TU may be the same person, or they may be different people. That is, in the inference process of inferring the semantic representation 49 for the subject TU, the inference device 2 may use a trained inference model 5 generated from data obtained from the subject TU (e.g., eye-related data 31), or may use a trained inference model 5 generated from data obtained from someone other than the subject TU. Furthermore, the trained inference model 5 may be generated from data obtained from multiple subjects LU (e.g., eye-related data 31), or may be generated from data obtained from a single subject LU. In one example, the subjects LU and the subjects TU may be classified into groups, and a trained inference model 5 may be shared within the same group.

[0062] (System Configuration) In one example, as shown in Figure 1, the model generation device 1 and the inference device 2 may be connected to each other via a network. The type of network may be selected as appropriate from, for example, the Internet, a wireless communication network, a mobile communication network, a telephone network, a dedicated network, etc. However, the method of exchanging data between the model generation device 1 and the inference device 2 is not limited to this example and may be selected as appropriate depending on the embodiment. In another example, data may be exchanged using a storage medium.

[0063] 1, the model generation device 1 and the inference device 2 are separate computers. However, the configuration of the system according to this embodiment is not limited to this example and may be determined appropriately depending on the embodiment. In another example, the model generation device 1 and the inference device 2 may be configured as a single computer. In yet another example, at least one of the model generation device 1 and the inference device 2 may be configured as multiple computers.

[0064] §2 Configuration example [Hardware configuration] (Model generation device) 2 schematically illustrates an example of the hardware configuration of the model generation device 1 according to this embodiment. In this example of this embodiment, the model generation device 1 is a computer to which a control unit 11, a storage unit 12, an external interface 13, an input device 14, an output device 15, and a drive 16 are electrically connected.

[0065] The control unit 11 includes a CPU (Central Processing Unit) which is a hardware processor, The control unit 11 includes a RAM (Random Access Memory), a ROM (Read Only Memory), etc., and is configured to execute information processing based on programs and various data. The control unit 11 (CPU) is an example of a processor resource.

[0066] The storage unit 12 may be configured, for example, with a hard disk drive, a solid state drive, a semiconductor memory, etc. The storage unit 12 (and RAM, ROM) are examples of memory resources. In this embodiment, the storage unit 12 stores various information such as a model generation program 81, eye-related data 31, behavioral data 33, and learning result data 50.

[0067] The model generation program 81 is a program for causing the model generation device 1 to execute information processing (see FIG. 6, described below) related to machine learning of the inference model 5. The model generation program 81 includes a series of instructions for the information processing. The learning result data 50 is configured to indicate information related to the generated trained inference model 5. In this embodiment, the learning result data 50 may be generated as a result of executing the model generation program 81. Note that the configuration of the learning result data 50 is not particularly limited as long as it can hold information for executing the calculation processing of the trained inference model 5, and may be determined appropriately depending on the embodiment. In one example, the learning result data 50 may be configured to include information indicating values of calculation parameters adjusted by machine learning. In some cases, the learning result data 50 may be configured to include information indicating the configuration of the inference model 5 (e.g., the structure of a neural network, etc.).

[0068] The external interface 13 is configured to connect to an external device via a wired or wireless connection. The external interface 13 may be, for example, a USB (Universal Serial Bus) port, a communication port, a dedicated port, or the like. The type and number of external interfaces 13 may be determined appropriately depending on the embodiment. If the external interface 13 includes a communication port, the model generation device 1 may perform data communication with another computer (e.g., the inference device 2, etc.) via a network. The communication standard of the communication port may be selected arbitrarily. Furthermore, in this embodiment, the model generation device 1 may be connected to at least one of the sensor SA and the sensor SB via the external interface 13.

[0069] The input device 14 is a device for inputting, for example, a mouse, a keyboard, an operator, etc. The output device 15 is a device for outputting, for example, a display, a speaker, etc. A user can operate the model generation device 1 by using the input device 14 and the output device 15. The input device 14 and the output device 15 may be connected via an external interface 13. The input device 14 and the output device 15 may be integrated into one device, for example, a touch panel display, etc.

[0070] The drive 16 is a device for reading various information such as programs stored in a storage medium 91. At least one of the model generation program 81, the eyeball-related data 31, the behavioral data 33, and the learning result data 50 may be stored in the storage medium 91 instead of or together with the storage unit 12. The storage medium 91 is configured to store various information (such as stored programs) by electrical, magnetic, optical, mechanical, or chemical action so that a machine such as a computer can read the information. The model generation device 1 may acquire at least one of the model generation program 81, the eyeball-related data 31, the behavioral data 33, and the learning result data 50 from the storage medium 91. The storage medium 91 may be a disk-type storage medium such as a CD or a DVD, or may be a non-disk-type storage medium such as a semiconductor memory (e.g., a flash memory). The type of the drive 16 may be selected appropriately depending on the type of the storage medium 91. The drive 16 may be connected via an external interface 13. The storage medium 91 may also include a memory resource provided in another computer, such as a network attached storage (NAS). When the storage medium 91 is a memory resource of another computer, the drive 16 may be omitted.

[0071] Regarding the specific hardware configuration of the model generating device 1, components can be omitted, replaced, or added as appropriate depending on the embodiment. For example, the control unit 11 may include multiple hardware processors. The hardware processors may be a microprocessor, a field-programmable gate array (FPGA), a digital signal processor (DSP), a GPU (Gateway Processor), a 3D processor, a 3D image ... The external interface 13, the input device 14, and the output device 15 may be configured by a PU (Graphics Processing Unit), an ASIC (Application Specific Integrated Circuit), etc. , and at least one of the drive 16 may be omitted. If the behavioral data 33 is not used, the behavioral data 33 may be omitted from the storage unit 12. The model generation device 1 may be composed of multiple computers. In this case, the hardware configurations of the computers may or may not be the same. Furthermore, the model generation device 1 may be an information processing device designed specifically for the service to be provided, as well as a general-purpose server device, a general-purpose PC (Personal Computer), a tablet PC, a terminal device, etc.

[0072] (Inference device) 3 shows a schematic diagram of an example of the hardware configuration of the inference device 2 according to this embodiment. In this example of this embodiment, the inference device 2 is a computer to which a control unit 21, a storage unit 22, an external interface 23, an input device 24, an output device 25, and a drive 26 are electrically connected.

[0073] The control unit 21 to the drive 26 and the storage medium 92 of the inference device 2 may be configured similarly to the control unit 11 to the drive 16 and the storage medium 91 of the model generation device 1. The control unit 21 (CPU) is an example of a processor resource of the inference device 2, and the storage unit 22 (and RAM, ROM) is an example of a memory resource of the inference device 2. In this embodiment, the storage unit 22 stores various information such as an inference program 82 and learning result data 50.

[0074] The inference program 82 is a program for causing the inference device 2 to execute information processing (see FIG. 7 described below) related to the inference of the content perceived from the content. The inference program 82 includes a series of instructions for the information processing. The learning result data 50 may be managed separately from the inference program 82, or may be incorporated into the inference program 82. At least one of the inference program 82 and the learning result data 50 may be stored in a storage medium 92 instead of or together with the storage unit 22. The inference device 2 may acquire at least one of the inference program 82 and the learning result data 50 from the storage medium 92.

[0075] If the external interface 23 includes a communication port, the inference device 2 may perform data communication with other computers (e.g., the model generation device 1, etc.) via a network. The inference device 2 may also be connected to at least one of the sensors SC and SD via the external interface 23. An operator can operate the inference device 2 by using the input device 24 and the output device 25.

[0076] Note that, with regard to the specific hardware configuration of the inference device 2, components can be omitted, replaced, or added as appropriate depending on the embodiment. For example, the control unit 21 may include multiple hardware processors. The hardware processor may be configured with a microprocessor, FPGA, DSP, GPU, ASIC, etc. At least one of the external interface 23, input device 24, output device 25, and drive 26 may be omitted. The inference device 2 may be configured with multiple computers. In this case, the hardware configurations of the computers may or may not be the same. The inference device 2 may be an information processing device designed specifically for the service to be provided, as well as a general-purpose server device, a general-purpose PC, a tablet PC, a terminal device, etc.

[0077] [Software configuration] (Model generation device) 4 schematically shows an example of the software configuration of the model generation device 1 according to this embodiment. The control unit 11 of the model generation device 1 loads a model generation program 81 stored in the storage unit 12 into RAM, and executes instructions included in the model generation program 81 using the CPU. As a result, the model generation device 1 operates as a computer that includes an acquisition unit 111, a training unit 112, and an output processing unit 113 as software modules.

[0078] The acquisition unit 111 is configured to acquire eye-related data 31 measured from a subject LU while viewing content LC. The training unit 112 is configured to perform machine learning of the inference model 5 using the acquired eye-related data 31. The machine learning includes training the inference model 5 to acquire the ability to infer, from the eye-related data 31, a semantic representation 39 in an information space corresponding to the content included in the content LC. The output processing unit 113 is configured to output the results of the machine learning.

[0079] In one example, the acquisition unit 111 may be configured to further acquire behavioral data 33 indicating the behavior of the subject LU while viewing the content LC. In response to this, the training unit 112 may be configured to perform machine learning of the inference model 5 using the acquired eye-related data 31 and behavioral data 33. The machine learning may include training the inference model 5 to acquire the ability to infer a semantic representation 39 in an information space corresponding to the content included in the content LC from the eye-related data 31 and the behavioral data 33. In one example, the acquisition unit 111 may be configured to acquire eye-related data 31 composed of gaze data, blink data, pupil diameter data, or a combination thereof.

[0080] (Inference device) 5 schematically shows an example of the software configuration of the inference device 2 according to this embodiment. The control unit 21 of the inference device 2 loads an inference program 82 stored in the storage unit 22 into RAM, and executes instructions included in the inference program 82 using the CPU. As a result, the inference device 2 operates as a computer that includes an acquisition unit 211, an inference unit 212, and an output processing unit 213 as software modules.

[0081] The acquisition unit 211 is configured to acquire eyeball-related data 41 measured from a subject TU viewing the target content TC. The inference unit 212 is provided with a trained inference model 5 by holding learning result data 50. The inference unit 212 is configured to infer a semantic representation 49 in an information space corresponding to the content included in the target content TC from the acquired eyeball-related data 41, using the trained inference model 5. The output processing unit 213 is configured to output a result of inferring the semantic representation 49.

[0082] In one example, the acquisition unit 211 may be configured to further acquire behavioral data 43 indicating the behavior of the subject TU while viewing the target content TC. In response, the inference unit 212 may be configured to infer a semantic representation 49 from the acquired eye-related data 41 and behavioral data 43 using a trained inference model 5. In one example, the acquisition unit 211 may be configured to acquire eye-related data 41 composed of gaze data, blink data, pupil diameter data, or a combination thereof.

[0083] (others) In this embodiment, an example is described in which each software module of the model generation device 1 and the inference device 2 is implemented by a general-purpose CPU. However, some or all of the above software modules may be implemented by one or more dedicated processors or chipsets. Each of the above modules may be implemented as a hardware module. With regard to the software configuration of the model generation device 1 and the inference device 2, modules may be omitted, replaced, or added as appropriate depending on the embodiment.

[0084] §3 Example of operation [Model generation device] FIG. 6 is a flowchart showing an example of the processing procedure of the model generation device 1 according to this embodiment. The following processing procedure is a model generation method (information processing method) executed by a computer. ) However, the processing procedure of the model generation device 1 is merely an example, and each step may be changed as much as possible. Furthermore, steps in the following processing procedure may be omitted, replaced, or added as appropriate depending on the embodiment.

[0085] (Step S101) In step S101, the control unit 11 operates as the acquisition unit 111 and acquires the training eye-related data 31 measured from the subject LU who is viewing the content LC.

[0086] In one example, the control unit 11 may further acquire behavioral data 33 indicating the behavior of the subject LU while viewing the content LC. The sensor SA may be used to acquire the eyeball-related data 31, and the sensor SB may be used to acquire the behavioral data 33. At least a portion of the eyeball-related data 31 and the behavioral data 33 may be generated by the model generation device 1 or by another computer. The control unit 11 may acquire at least a portion of the eyeball-related data 31 and the behavioral data 33 via a network, an external storage device, a storage medium 91, or the like. The order in which the eyeball-related data 31 and the behavioral data 33 are acquired is not particularly limited and may be selected appropriately depending on the embodiment. The eyeball-related data 31 and the behavioral data 33 may be acquired at least partially in parallel. In another example, the acquired eyeball-related data 31 may be composed of gaze data, blink data, pupil diameter data, or a combination thereof. After acquiring the eyeball-related data 31, the control unit 11 proceeds to the next step S102.

[0087] (Step S102) In step S102, the control unit 11 operates as a training unit 112 and performs machine learning of the inference model 5 using the acquired eyeball-related data 31. In the machine learning, the control unit 11 trains the inference model 5 so that it acquires the ability to infer, from the eyeball-related data 31, a semantic representation 39 in the information space corresponding to the content included in the content LC. This makes it possible to generate a trained inference model 5 that has acquired the ability to infer a semantic representation from the eyeball-related data within the category of the training sample.

[0088] In one example, when behavioral data 33 has been acquired, the control unit 11 may further use the acquired behavioral data 33 in machine learning of the inference model 5. That is, in the machine learning process, the control unit 11 may train the inference model 5 so that when the eye-related data 31 and behavioral data 33 are given, the inference model 5 outputs an inference result that matches the correct answer of the corresponding semantic representation 39. This makes it possible to generate a trained inference model 5 that has acquired the ability to infer a semantic representation from the eye-related data and behavioral data.

[0089] In one example, training the inference model 5 may be optimizing the values of the calculation parameters of the inference model 5 according to the training samples. The machine learning method may be determined appropriately depending on the embodiment, such as the type and structure of the machine learning model used in the inference model 5. Any method, such as backpropagation or solving an optimization problem, may be adopted as a method for adjusting the calculation parameters. In another example, the control unit 11 may execute the machine learning calculation process within the model generation device 1, or may instruct another computer to execute the machine learning calculation process. When the machine learning (training the inference model 5) is completed, the control unit 11 proceeds to the next step S103.

[0090] (Step S103) In step S103, the control unit 11 operates as the output processing unit 113 and outputs the results of the machine learning.

[0091] The output destination and the content of the information to be output may be selected appropriately depending on the embodiment. In one example, the control unit 11 may generate learning result data 50 indicating the results of the machine learning as an output process and store the generated learning result data 50 in a predetermined storage area. The predetermined storage area may be, for example, RAM within the control unit 11, the storage unit 12, an external storage device, a storage medium, or a combination thereof. The storage medium may be, for example, a CD, a DVD, a semiconductor memory, or the like. The external storage device may be, for example, a data server such as a NAS. The external storage device may be, for example, an external storage device. When the machine learning of the inference model 5 is executed on another computer, the learning result data 50 may be generated on the other computer. In another example, the control unit 11 may output information (e.g., loss, etc.) obtained during the machine learning calculation process as an output process. The output destination may be, for example, RAM within the control unit 11, the storage unit 12, the output device 15, an external computer, an external storage device, a storage medium, or a combination thereof.

[0092] When the output of the machine learning results is completed, the control unit 11 ends the processing procedure of the model generation device 1 according to this operation example.

[0093] The generated learning result data 50 may be provided to the inference device 2 at any timing and by any method. For example, the control unit 11 may transmit the learning result data 50 to the inference device 2 as part of the processing of step S103 or separately from the processing of step S103. The inference device 2 may acquire the learning result data 50 (trained inference model 5) by receiving it. Also, for example, the inference device 2 may acquire the learning result data 50 by accessing the model generation device 1 or a data server via a network. Also, for example, the inference device 2 may acquire the learning result data 50 via a storage medium 92. Also, for example, the learning result data 50 may be pre-installed in the inference device 2.

[0094] Furthermore, the control unit 11 may update or generate new learning result data 50 by periodically or irregularly repeatedly executing the series of processes from step S101 to step S103 described above. During this repetition, at least a portion of the training samples used for machine learning may be changed, modified, added, deleted, etc. as appropriate. Then, the control unit 11 may update the learning result data 50 held by the inference device 2 by providing the updated or newly generated learning result data 50 to the inference device 2 by any method.

[0095] [Inference device] 7 is a flowchart showing an example of the processing procedure of the inference device 2 according to this embodiment. The following processing procedure is an example of an inference method (information processing method) executed by a computer. However, the processing procedure of the inference device 2 is merely an example, and each step may be changed as much as possible. Furthermore, steps in the following processing procedure may be omitted, replaced, or added as appropriate depending on the embodiment.

[0096] (Step S201) In step S201, the control unit 21 operates as the acquisition unit 211 and acquires the eyeball-related data 41 measured from the subject TU who is viewing the target content TC.

[0097] In one example, the control unit 21 may further acquire behavioral data 43 indicating the behavior of the subject TU while viewing the target content TC. The sensor SC may be used to acquire the eyeball-related data 41, and the sensor SD may be used to acquire the behavioral data 43. At least a portion of the eyeball-related data 41 and the behavioral data 43 may be generated by the inference device 2 or by another computer. The control unit 21 may acquire at least a portion of the eyeball-related data 41 and the behavioral data 43 via a network, an external storage device, a storage medium 92, or the like. The order in which the eyeball-related data 41 and the behavioral data 43 are acquired is not particularly limited and may be selected appropriately depending on the embodiment. The eyeball-related data 41 and the behavioral data 43 The acquisition of the above data may be performed at least partially in parallel. In one example, the acquired eyeball-related data 41 may be composed of gaze data, blink data, pupil diameter data, or a combination thereof. After acquiring the eyeball-related data 41, the control unit 21 proceeds to the next step S202.

[0098] (Step S202) In step S202, the control unit 21 operates as an inference unit 212 and uses a trained inference model 5 to infer a semantic representation 49 in the information space corresponding to the content contained in the target content TC from the acquired eyeball-related data 41.

[0099] In one example, when behavioral data 43 is acquired, the control unit 21 may use a trained inference model 5 to infer a semantic expression 49 from the acquired eye-related data 41 and behavioral data 43.

[0100] The computational processing of the trained inference model 5 may be performed as appropriate depending on the type, configuration, and other aspects of the inference model 5. For example, if the inference model 5 is configured as a neural network, the control unit 21 may input the eyeball-related data 41 (and behavioral data 43) into the trained inference model 5 and perform forward computational processing of the trained inference model 5. This allows the control unit 21 to obtain an output corresponding to the result of inferring the semantic representation 49 from the trained inference model 5. When the inference of the semantic representation 49 is completed, the control unit 21 proceeds to the next step S203.

[0101] (Step S203) In step S203, the control unit 21 operates as the output processing unit 213 and outputs the result of inferring the semantic expression 49.

[0102] The content of the information to be output as the inference result is not particularly limited as long as it is related to the inference result of the semantic representation 49, and may be determined appropriately depending on the embodiment. In one example, the control unit 21 may output the inference result of the semantic representation 49 as is. Outputting the inference result of the semantic representation 49 as is may include outputting a vector obtained by inference as the semantic representation 49, and outputting content (phrases, context, etc.) corresponding to the obtained vector. When the semantic representation is obtained by converting content using a predetermined projection method, the control unit 21 can obtain the content corresponding to the vector by performing an inverse conversion of the predetermined projection method on the inference result of the semantic representation 49.

[0103] In another example, the control unit 21 may perform any information processing (for example, a determination process, an analysis process, etc.) on the obtained inference result. Then, the control unit 21 may output the result of the information processing as information related to the inference result.

[0104] For example, the inferred semantic representation 49 corresponds to the content perceived by the target person TU. Therefore, the control unit 21 may determine the degree of match between the inferred semantic representation 49 (the inference result of the perceived content) and the content actually included in the target content TC. The control unit 21 may evaluate whether the target person TU is perceiving the target content TC appropriately according to the determined degree of match. Perceiving appropriately may include, for example, concentrating on watching the content, accurately understanding the content of the content, etc. If the determined degree of match is low, the control unit 21 may evaluate that the target person TU is not perceiving the target content TC appropriately (not concentrating on watching the target content TC, not understanding the content of the target content TC, etc.). On the other hand, if the determined degree of match is high, the control unit 21 may evaluate that the target person TU is perceiving the target content TC appropriately (concentrating on watching the target content TC, understanding the content of the target content TC, etc.). Whether the degree of match is high or low can be determined, for example, by comparing with a threshold value. The control unit 21 may output the result of this evaluation as an output process of the inference result. As an example of an application scenario, this output format may be adopted in an educational setting to evaluate whether a student is concentrating on watching educational content, whether the student has understood the educational content, etc.

[0105] Furthermore, for example, the control unit 21 may categorize the subject TU according to the inference result of the obtained semantic representation 49. Categorization is to determine the group to which the subject TU belongs. The control unit 21 may output the categorization result (i.e., the result of determining the group to which the subject TU belongs) as an output process of the inference result. Groups are presumed to correspond to perceptual tendencies. In other words, belonging to the same group is presumed to indicate similar perceptual tendencies. Therefore, as an example of an application scenario, this output format may be adopted in medical settings to infer the degree of a neurological disorder that affects the perception of content. The categorization result may correspond to the degree of the neurological disorder. This enables medical assistance (support). As another example of an application scenario, it is possible to estimate content that is likely to be noticed in the target content TC from the semantic representation 49 perceived by each group. Therefore, the control unit 21 may output the inferred semantic representation 49 for each group. The output inference result information may be used as content analysis material in production scenarios such as optimizing the target content TC or creating new content. The output format of this analysis data may be adopted regardless of categorization. That is, the control unit 21 may output the inference result of the semantic expression 49 without categorizing it. The output inference result may be used as the analysis data.

[0106] Furthermore, for example, a computational model that projects from the information space of the semantic representation to another space (such as another latent space) may be further prepared. The computational model may be, for example, a trained machine learning model or the like. The control unit 21 may convert the inference result of the semantic representation 49 into another representation by using the computational model to project the inference result of the semantic representation 49 into another space. The control unit 21 may output the obtained another representation as an output process of the inference result. The another representation may be selected arbitrarily. The another representation may be, for example, an emotional representation, a future semantic representation (future prediction), etc. When an emotional representation is adopted as the another representation, the control unit 21 can obtain the result of inferring the emotion of the subject TU when viewing the target content TC via the inference result of the semantic representation 49. When a future prediction is adopted as the another representation, the control unit 21 can obtain the result of future prediction of the content perceived by the subject TU regarding the target content TC via the inference result of the semantic representation 49. When inferring the semantic representation 49 includes future prediction, the another representation may be a semantic representation further in the future than the semantic representation 49. In one example, a plurality of computational models may be prepared, each of which projects onto a plurality of different spaces, by machine learning the projection relationships between the information space of the semantic representation and each of the plurality of different spaces. By selectively using the plurality of computational models, the control unit 21 can obtain a representation in any other space starting from the information space in which direct projection from eyeball-related data has been learned.

[0107] The output destination is not particularly limited and may be selected appropriately depending on the embodiment. The output destination may be, for example, RAM in the control unit 21, the storage unit 22, the output device 25, an external computer, an external storage device, a storage medium, or a combination thereof.

[0108] When the output of the inference result is completed, the control unit 21 ends the processing procedure of the inference device 2 according to this operation example. In one example, the control unit 21 may execute a series of processes from step S201 to step S203 in real time. That is, the control unit 21 may execute a process of inferring a semantic expression 49 from the eyeball-related data 41 (and behavior data 43) obtained in real time in step S202 and a process of outputting the inference result in step S203. In another example, the control unit 21 may infer a semantic expression 49 retrospectively from the eyeball-related data 41 (and behavior data 43) measured in the past. That is, the control unit 21 In step S201, eye-related data 41 (and behavior data 43) measured in the past may be acquired, and the processes of steps S202 and S203 may be executed on the acquired eye-related data 41 (and behavior data 43).

[0109] [Features] In this embodiment, the model generation device 1 generates a trained inference model 5 that has acquired the ability to infer from eyeball-related data what is perceived from content through the processing of step S102. By using this trained inference model 5, it is possible to infer, at low cost and in a simple manner, what is perceived by an individual while viewing content. The inference device 2, through the processing of steps S201 and S202 above, can infer, at low cost and in a simple manner, what is perceived by the subject TU while viewing the target content TC.

[0110] §4 Variations Although the embodiments of the present disclosure have been described in detail above, the above description is merely an example of the present disclosure in every respect. It goes without saying that various improvements or modifications can be made without departing from the scope of the present disclosure. The processes and means described in the present disclosure can be freely combined and implemented as long as no technical contradiction occurs.

[0111] For example, in the above embodiment, the inference device 2 holds the trained inference model 5 and executes the process of inferring the semantic representation 49. However, the entity executing the inference process does not have to be limited to the inference device 2. The trained inference model 5 (learning result data 50) may be held in another computer, and the inference device 2 may request the other computer to execute the inference process. In this case, the inference device 2 does not need to hold the trained inference model 5, and may obtain the results of inferring the semantic representation 49 from the other computer.

[0112] §5 Experimental Examples The following experiment was conducted to verify whether it is possible to infer the semantic expression of the content from eye activity while viewing the content, although the present invention is not limited to the following experimental example.

[0113] (preparation) First, we prepared 10-minute videos (7 videos). The videos were general video works. We had the subjects watch the prepared videos, and by measuring the pupil diameter (Y) while watching the videos, we obtained pupil diameter data. The pupil diameter data consisted of one-dimensional numerical data for each unit of time. In addition, we had a person (annotator) separate from the subjects annotate each scene in the prepared videos. Using a computational model (word2vec), we converted the annotations into a 1000-dimensional vector (X w2v ) was converted into a vector of the pupil diameter (Y) and the semantic expression (X). A publicly known model was used as the calculation model. This resulted in the correct value of the semantic expression in the information space (1000-dimensional numerical data for each unit time). w2v ) were divided into training samples and evaluation samples. Specifically, the evaluation experiment, in which samples obtained from six of the seven videos were used as training samples and samples obtained from one video were used as evaluation samples, was repeated seven times so that each video served as an evaluation sample, and the evaluation results were obtained by averaging the seven evaluations. 58 subjects participated in this experiment.

[0114] Next, to take into account the delay in pupil response, four time points were set to be used to infer the pupil diameter at the target time. The target time was set to 0 seconds, and the four time points were -2 seconds, +2 seconds, +6 seconds, and +10 seconds. Then, by performing regression analysis using the training samples, each parameter (W ) of the regression model (encoding model) that infers the pupil diameter data (Y) at the target time from the semantic representation vector (4000 dimensions in total) of these four time points was calculated. w2v ) was identified. The following equation 1 was used for the regression model.

[0115]

number

[0116] Next, the obtained 1000-dimensional parameters (W w2v_avrg ) values and evaluated the contribution of each parameter to the inference of pupil diameter, generating a principal component vector (W). The dimensions of the principal component vector (W) were set to 2, 10, 12, 14, 16, 18, and 20. The principal component vector (W) for each dimension was used to generate a semantic representation vector (X w2v ) into the principal component space, we obtain a compressed vector (Xe w2v ) training samples were obtained.

[0117] Next, to take time delays into account, 21 time points were set to be used to infer the vector of the semantic expression of the target time. The target time was set to 0 seconds, and the 21 time points were taken in 2-second increments from 20 seconds before to 20 seconds after. Then, by performing regression analysis using the training samples, a vector (Xe w2v The values of each parameter (D) of the regression model (decoding model) that infers the above-mentioned feature were identified for each dimension of the principal component vector.

[0118] Figure 8 shows the relationship between the principal component vectors and the decoding model obtained in the experiment. As shown in Figure 8, the transposed matrix (W t ) to obtain a compressed vector (Xe w2v ) is projected (inversely transformed) to the original 1000-dimensional semantic representation vector (Xd w2v ) can be obtained. Therefore, the decryption model and the transpose matrix of the principal component vector (W t ) to convert the pupil diameter (Y) into a semantic representation vector (X w2v An inference model for inferring the above was constructed for each dimension of the principal component vector.

[0119] (verification) Next, using the evaluation sample, we evaluated the accuracy of the obtained inference model for each dimension of the principal component vector and for each subject. Specifically, we used the above decoding model from the obtained inference model to analyze the compressed semantic representation vector (Xe w2v ) is calculated from the pupil diameter (Y), and the vector (Xe w2v ) is the original vector (X w2v ) is a compressed semantic representation vector (Xe w2v ) and calculated the inference accuracy (correlation coefficient) of the inference model according to whether or not the principal component vectors matched the corresponding values. In this accuracy evaluation, each of the 58 subjects was designated as the first subject, and the remaining 57 subjects were designated as the second subjects. The principal component vectors obtained from the 57 second subjects were averaged to obtain the principal component vectors to be applied to the first subjects. A decoding model was generated from the training samples obtained from the first subjects, and the generated decoding model was applied to the evaluation samples. The above evaluation was then repeated for each dimension of the principal component vector for the 58 subjects, and the obtained accuracies were averaged to calculate the average inference accuracy for each dimension of the principal component vector (Figure 9). In addition, using the obtained inference model, a semantic expression vector (Xd w2v) was calculated. Then, using t-SNE (t-distributed Stochastic Neighbor Embedding), the calculated vector (Xd w2v ) into space By placing it on top, the vector (Xd w2v ) is expressed in space (Figure 10).

[0120] Furthermore, to verify individual differences, the coding model (parameters (W w2v )) was used for a second subject different from the first subject, and the second subject's semantic representation was extracted from the vector of the semantic representation. We evaluated whether the system could correctly infer the pupil diameter of each subject at a target time. Each of the 30 subjects was designated as the first subject, and the remaining 29 subjects were designated as second subjects. The pupil diameter of the second subject was then inferred from the encoding model of the first subject, and accuracy was evaluated from the inference results (29 times). For each of the 30 subjects, the difference between the accuracy when the first subject's own pupil diameter was inferred using the encoding model of the first subject and the average accuracy of the 29 times was calculated (30 times) to calculate the change in inference accuracy between individuals (Figure 11). Note that a commercially available PC was used for all of the above calculations.

[0121] (result) Figure 9 shows the results of calculating the average inference accuracy of the inference model for each dimension of the principal component vector in the experimental example. Figure 10 shows the results of expressing the perceptual content inferred from pupil diameter using t-SNE. Figure 11 also shows the results of calculating the inter-individual change in inference accuracy of the coding model in the experimental example.

[0122] As shown in Figure 9, for example, when the number of dimensions of the principal component vector is three, the average accuracy (correlation coefficient) of inferring a 1,000-dimensional semantic expression from one-dimensional pupil diameter was approximately 0.18. These results demonstrate that the resulting inference model can achieve statistically significantly higher inference power (correlation coefficient of approximately 0.18) than a model without inference power. Furthermore, as shown in Figure 10, the content of the content could be restored from one-dimensional pupil diameter. Pupil diameter is an example of eye activity, and other eye activities such as gaze and blinking are also related to pupil diameter. Therefore, these results suggest that it is possible to infer the content perceived by an individual while viewing content from eye activity.

[0123] In addition, as shown in Figure 11, the inference accuracy of the coding model deteriorated for some subjects but not for the rest. These results suggest that, although there are individual differences, eye activity is similar to at least some extent, and that a trained inference model generated from other subjects' eye-related data can infer semantic representations. Furthermore, it is possible to categorize the inference results of semantic representations, suggesting that it is possible to analyze the tendencies of subjects by group. In other words, it is possible to apply the above-mentioned diagnosis and content optimization by grouping the inference results of semantic representations. [Explanation of symbols]

[0124] 1...Model generation device, 11...control unit, 12...storage unit, 13...external interface, 14...input device, 15...output device, 16...drive, 81...model generation program, 91...storage medium, 111...acquisition unit, 112...training unit, 113...output processing unit, LU...subject, LC...content, SA·SB…sensor, 31...eye-related data, 33...behavioral data, 39...Semantic expression, 2... Reasoning device, 21...control unit, 22...storage unit, 23...external interface, 24...input device, 25...output device, 26...drive, 82...inference program, 92...storage medium, 211...acquisition unit, 212...inference unit, 213...output processing unit, TU...target audience, TC...target content, SC·SD…sensor, 41...eye-related data, 43...behavioral data, 49...Semantic expression, 5...Inference model

Claims

1. The computer acquiring eye-related data measured from a subject viewing content; performing machine learning of an inference model using the acquired eye-related data; and outputting the results of the machine learning; Run The machine learning includes training the inference model to acquire an ability to infer, from the eye-related data, a semantic representation in an information space corresponding to a content included in the content. Model generation method.

2. The computer further performs a step of acquiring behavioral data indicative of the subject's behavior while viewing the content; The acquired behavioral data is further used in the machine learning of the inference model, In the machine learning, the inference model is trained to acquire the ability to infer the semantic representation from the eye-related data and the behavioral data. The model generation method of claim 1 .

3. The acquired eye-related data is composed of gaze data, blink data, pupil diameter data, or a combination thereof. The model generation method of claim 1 .

4. On the computer, acquiring eye-related data measured from a subject viewing target content; inferring, from the acquired eye-related data, a semantic representation in an information space corresponding to the content included in the target content using a trained inference model; and outputting the result of inferring the semantic representation; In order to execute Inference program.

5. causing the computer to further execute a step of acquiring behavioral data indicating behavior of the subject while viewing the target content; the inferring step causes the computer to infer the semantic representation from the acquired eye-related data and behavioral data; The inference program according to claim 4.

Citation Information

Patent Citations

  • Estimation method of perceived semantic content by analysis of brain activity

    JP2016195716A