Content display system, content image selection method, and content selection program
The content display system addresses the mismatch between user preferences and advertiser objectives by mapping user emotions onto an emotion map to select content that aligns with user emotions, enhancing user satisfaction and relevance in metaverse spaces.
Patent Information
- Application Number
- PCT/JP2024/040796
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-02-22
- Filing Date
- 2024-11-18
- Publication Date
- 2025-08-28
AI Technical Summary
Existing technologies in metaverse spaces fail to provide content that aligns with both user preferences and advertiser objectives, leading to mismatched information suggestions.
A content display system that analyzes candidate content features, estimates user emotions using biometric data, and maps these emotions onto an emotion map to select content that aligns with user emotions based on predefined selection rules, utilizing a machine learning model to convert content features into emotional values.
The system effectively selects and displays content that suits user emotions, ensuring both user satisfaction and advertiser relevance, thereby optimizing content delivery in metaverse environments.
Smart Images

Figure JP2024040796_28082025_PF_FP_ABST
Abstract
Description
Content display system, content image selection method, and content selection program
[0001] The present invention relates to a content display system, a content selection method, and a content selection program.
[0002] In recent years, technologies have been developed that allow customers and users to participate as avatars in a metaverse space created on a computer, creating virtual societies such as shopping malls, shops, exhibitions, and offices.
[0003] For example, Patent Document 1 discloses a technology for providing information optimized for each user that can also be applied to the metaverse space. Specifically, Patent Document 1 describes a form in which a user purchases a product using an e-commerce service. Patent Document 1 detects the user's state to detect what is being viewed, the level of concentration, and emotions. Patent Document 1 also detects the user's state to determine what is being viewed, the level of concentration, and emotions. Patent Document 1 then generates user preference information by referring to the determination results. As a result, Patent Document 1 discloses executing a program, such as information suggestion, based on the preference information.
[0004] Japanese Patent Application Laid-Open No. 2022-030832
[0005] According to the above-mentioned prior art, it is possible to propose information corresponding to the user's preference information based on the detected user state, but there is a problem that it is not possible to propose information that is in line with the objectives of the advertiser, etc. An object of the present invention is to provide a content playback system, a content selection method, and a content selection program that select and play content that suits the user's emotions from candidate content.
[0006] The above problem is solved as follows.
[0007] (1) A content display system including: a content analysis unit that analyzes feature values of candidate content to be played back by a playback unit; an emotion information acquisition unit that determines an estimated emotion of a user who can view the playback unit based on biometric information of the user; and a determination unit that selects content to be played back by the playback unit based on an emotion estimated to be evoked by the candidate content converted from the feature values of the candidate content, the estimated emotion of the user, and a selection rule.
[0008] (2) In the content display system described in (1), the determination unit maps the emotion estimated to be evoked by the candidate content and the user's estimated emotion on an emotion map, and selects candidate content on the emotion map that satisfies the selection rule for the estimated emotion.
[0009] (3) In the content display system described in (2), the selection rule is a position or distance relationship between the emotion estimated to be evoked by the candidate content on the emotion map and the estimated emotion of the user.
[0010] (4) The content display system described in (1) further includes an emotional value conversion model, which is a machine learning model that inputs the feature values of the content and outputs an emotional value estimated to be evoked by the image, and the feature values are converted into an emotional value estimated to be evoked by the candidate image using the emotional value conversion model.
[0011] (5) In the content display system described in (1), the emotion information acquisition unit estimates the user's emotion based on biometric information detected by at least one of a pupil diameter sensor, a gaze measurement sensor, and a facial expression reading sensor.
[0012] (6) A display content selection method including: a step of analyzing features of candidate content to be played back by a playback unit; a step of calculating an estimated emotion of a user who can view the playback unit based on biometric information of the user; a step of converting the features of the candidate content into an emotion estimated to be evoked by the candidate content; a step of mapping the emotion estimated to be evoked by the candidate content and the estimated emotion of the user on an emotion map; and a step of selecting candidate content on the emotion map that satisfies a selection rule for the estimated emotion and displaying the candidate content on a display unit.
[0013] (7) A display content selection program for causing a computer to execute the following steps: analyzing the features of candidate content to be played back by a playback unit; determining an estimated emotion of a user who can view the playback unit based on the biometric information of the user; converting the features of the candidate content into an emotion estimated to be evoked by the candidate content; mapping the emotion estimated to be evoked by the candidate content and the estimated emotion of the user on an emotion map; and selecting candidate content on the emotion map that satisfies a selection rule for the estimated emotion and displaying the selected candidate content on a display unit.
[0014] According to the present invention, content that suits the user's emotions is selected from candidate content and output, so that content that is useful to both the user and the information provider can be output.
[0015] It is a diagram for explaining the process of determining a display image of the image display system. It is a continuation diagram for explaining the process of determining a display image of the image display system. It is a system configuration diagram of the image display system of the embodiment. It is a flow diagram for explaining the operation of the image display system.
[0016] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings. The following embodiments are assumed to be an image display system that selects and displays images in a metaverse space such as a shopping mall, shop, exhibition, or office. The images are displayed as product images, advertisement images, etc., depending on the emotions of an avatar (user). The present invention can also be similarly implemented for product images displayed on a web screen on an EC (Electronic Commerce) site. The present invention can also be similarly implemented when advertisement images are displayed depending on the emotions of site users (users) on an EC (Electronic Commerce) site.
[0017] 1 and 2 are diagrams illustrating the display image determination process in an image display system according to an embodiment. The image display system maps the estimated emotion of an avatar (user) and the converted emotions of candidate images displayed as products or advertisements onto an emotion map. Then, a candidate image having a converted emotion that satisfies a predetermined rule on the emotion map for the estimated emotion of the avatar (user) is selected as the selected image.
[0018] The emotion map in Figure 1 applies Russell's circumplex model, which classifies emotions as "pleasant-unpleasant" on the horizontal axis and "arousal-calm (drowsy)" or "high-low arousal level" on the vertical axis. In Russell's circumplex model, each emotion is represented by the direction and magnitude of a vector on a two-dimensional coordinate axis. There is a relationship between the angles from the origin (the intersection of the two axes) between each mapped emotion, and the difference in vector direction represents the correlation coefficient. The strength of the emotion is indicated by the magnitude (distance) of the vector from the origin.
[0019] The circle 90 in Figure 1 indicates the estimated emotion of the avatar (user). As will be described in detail later, the emotion classification and intensity are estimated based on the user's biometric information and mapped to Russell's emotional circumplex model. That is, the image display system estimates emotion map coordinate values for the "pleasant / unpleasant" axis and the "arousal / calm" axis based on the user's biometric information. The estimated emotion map coordinate values are mapped to Russell's emotional circumplex model.
[0020] Star marks 91A, 91B, 91C, and 91D in Fig. 1 indicate the emotions evoked by a user when the user views candidate images A, B, C, and D of products or advertisements selected by the image display system. As will be described in detail later, the image display system converts the emotions evoked by the user when viewing the candidate images from the image feature quantities of the candidate images. The image display system then maps the emotional values (converted values) evoked by the candidate images onto Russell's emotional circle model.
[0021] In this way, the image display system clarifies the relationship between the estimated emotion of the avatar (user) and the converted emotion of the candidate image by mapping it onto Russell's emotional circle model in Figure 1, and then selects an image to be displayed to the avatar (user) from the candidate images based on the selection rules described below.
[0022] FIG. 2 is a diagram illustrating how an image corresponding to the inferred emotion of an avatar (user) is selected based on a predetermined selection rule in the Russell emotional cycle model in which the inferred emotion of the avatar (user) and the converted emotion of a candidate image are mapped, as described in FIG. 1 .
[0023] In an image display system that uses images as advertisements, this allows for the display of advertisements that are more effective and suited to the user's condition, and also avoids displaying advertisements that offend the user.
[0024] An example of a selection rule is indicated by arrow 92 in Figure 2. In this case, the selection rule is the one that provides the shortest distance on the "arousal-calm" axis and the longest distance on the "pleasant-unpleasant" axis. The image display system applies this selection rule to the avatar (user) indicated by the dotted circle, and selects candidate image D that is estimated to evoke the emotion indicated by star mark 91D as the image to be displayed to the avatar (user). In this way, by setting the selection rule as the position or distance relationship between the emotion estimated to be evoked by candidate content on the emotion map and the user's estimated emotion, it is possible to impart a pleasant emotion to the user without changing the user's level of arousal.
[0025] Another selection rule may be the furthest distance on the "arousal-calm" axis and the closest distance on the "pleasant-unpleasant" axis. This makes it possible to calm or arouse the user without changing the degree of pleasant emotion imparted to the user. The above two selection rules are rules that change the emotion of the avatar (user) to the equivalent emotion evoked by the selected image by showing the selected image to the avatar (user).
[0026] Furthermore, if the selection rule is that the distance is close to the estimated emotion of the avatar (user), it is possible to select an image that maintains the estimated emotion of the avatar (user).
[0027] Next, the configuration of an image display system according to an embodiment will be described. Fig. 3 is a system configuration diagram of the image display system according to an embodiment. In the following description, a case will be described in which the image display system displays images of products and advertisements to an avatar (user) in a metaverse space such as a shopping mall, a shop, an exhibition, or an office, or in an e-commerce site.
[0028] Image display system 1 has a display image analysis section 2, a biometric information sensor 3, an emotion information acquisition section 4, a conversion section 5 using an emotion value conversion model 51, and a determination section 6. Image display system 1 is also configured to have an image display control section 7, a display device 8, and an emotion value conversion model update section 9. The configuration will be explained in detail below.
[0029] The display image analysis unit 2 has a candidate image acquisition unit 21 and an image feature analysis unit 22, and quantifies the candidate images by determining image feature amounts of candidate images of products or advertisements to be displayed to the avatar (user). Specifically, the candidate image acquisition unit 21 acquires image information of the candidate images from a metaverse space generation system or an EC site via a network.
[0030] The image feature analysis unit 22 analyzes each candidate image acquired by the candidate image acquisition unit 21 to determine image feature quantities. The image feature quantities determined by the candidate image acquisition unit 21 include, for example, a saliency ratio, a pixel average of complexity, a white space ratio, a power spectrum linearity of spatial frequency, a power spectrum gradient of spatial frequency, a skewness, a number of dominant colors, a brightness contrast, a saturation contrast, a brightness average, and a saturation average.
[0031] The above describes a case where the image feature analysis unit 22 analyzes and obtains image feature amounts for the image information obtained by the candidate image acquisition unit 21. However, the present invention is not limited to this, and the image feature amounts of the candidate images may be obtained in advance by a metaverse space generation system or an EC site that provides the images, and the candidate images and their image feature amounts may be notified to the image display system 1.
[0032] Conversion unit 5 using emotional value conversion model 51 converts the image feature quantities of each candidate image determined by displayed image analysis unit 2 into emotional values that are estimated to be evoked in the user by these candidate images. The conversion is performed by emotional value conversion model 51, which is a machine learning model. Emotional value conversion model 51 uses a random forest algorithm, for example, to carry out emotion prediction and estimation using the image feature quantities as input.
[0033] Emotions evoked when a user views a plurality of learning images are obtained in advance through a questionnaire. Furthermore, the display image analysis unit 2 obtains image feature quantities from the plurality of learning images. That is, in the display image analysis unit 2, the image feature analysis unit 22 obtains image feature quantities from the image information acquired by the candidate image acquisition unit 21. Furthermore, the display image analysis unit 2 uses the image feature quantities as input, and regards the combination of emotions evoked when a user views the images, which has been obtained in advance through the questionnaire, as correct answer data. This correct answer data is then used to perform supervised learning of the emotion value conversion model 51.
[0034] Furthermore, instead of using emotion value conversion model 51, which is a machine learning model, conversion unit 5 may convert image feature amounts into emotion values evoked by the image using an emotion conversion database that indicates the correspondence between image feature amounts and emotion values.
[0035] The biometric information sensor 3 acquires biometric information for determining the estimated emotion of the avatar (user), and is, for example, a pupil diameter sensor, a gaze measurement sensor, a facial expression reading sensor (camera), a near-infrared camera, or other sensor. The biometric information sensor 3 acquires, for example, pupil diameter / pupil position, heart rate, feature points (facial expressions) in the lower half of the face, and brain waves as biometric information. The biometric information sensor 3 may be provided independently or may be configured integrally with the image display control unit 7. Specifically, the biometric information sensor 3 is provided in a VR (Virtual Reality) headset / VR goggles / HMD (Head Mounted Display), which is a display device in the metaverse space. Also, a webcam is considered to be the biometric information sensor 3.
[0036] The emotion information acquisition unit 4 has a biometric data acquisition unit 41 and an emotion estimation unit 42 based on biometric data analysis, and estimates the emotion of the avatar (user) based on the biometric information of the avatar (user) acquired by the biometric information sensor 3. In detail, the biometric data acquisition unit 41 acquires the biometric information of the target avatar (user) from the biometric information sensor 3 via a communication path. Here, the target avatar (user) refers to a user who is in a position where the display device 8 can be seen.
[0037] The emotion deduction unit 42 analyzes the biometric data using the biometric information acquired by the biometric data acquisition unit 41. The emotion deduction unit 42 analyzes the biometric data using gaze position, pupil diameter change, heart rate interval, facial expression, electroencephalogram (α / β / γ waves) fluctuations, etc., to calculate analyzed biometric data. The emotion deduction unit 42 then determines the emotion of the avatar (user) as an estimated emotion based on the analyzed biometric data. For example, if the emotion deduction unit 42 detects a tendency for pupil dilation through its analysis of the biometric data, the arousal level is determined to be high.
[0038] The determination unit 6 has an emotion model mapping unit 61 and a display image selection unit 62. The determination unit 6 selects an image to be shown to the avatar (user) based on the emotion map and the estimated emotion of the avatar (user) determined by the emotion information acquisition unit 4, the emotion value of the candidate image determined by the display image analysis unit 2 and the conversion unit 5 using the emotion value conversion model 51, and predetermined selection rules.
[0039] In detail, the emotion model mapping unit 61 maps the inferred emotion of the avatar (user) and the converted emotion of the candidate image onto an emotion map such as Russell's emotion circle model, as explained in Fig. 1. Then, the display image selection unit 62, as a display content selection method, finds a converted emotion on the emotion map that satisfies a selection rule of a positional relationship or distance condition for the inferred emotion of the avatar (user) mapped onto the emotion map, and selects an image from the candidate images that evokes the emotion closest to the converted emotion.
[0040] The image display control unit 7 displays the image selected by the determination unit 6 on the display device 8, which is a display unit. Specifically, the image display control unit 7 notifies the metaverse space generation system or the EC site of the selected image. Then, the metaverse space generation system displays this image on the VR headset / VR goggles / HMD, which is the display device 8, or the EC site displays this image on the monitor of the user terminal, which is the display device 8.
[0041] An update unit 9 for emotional value conversion model 51 performs supervised learning of emotional value conversion model 51 using the image feature amounts of the image selected by decision unit 6 as input and the converted emotion values as correct answer data, and updates emotional value conversion model 51.
[0042] In the above description, the image display system 1 calculates the emotion value evoked by a candidate image and selects an image according to the estimated emotion of the avatar (user). Instead of the candidate image, the image display system 1 calculates the feature values of content such as video and audio streams, converts them into emotion values evoked in the user who views the content, and maps them on an emotion map. Then, content such as video and audio streams may be selected according to the estimated emotion of the avatar (user). Alternatively, a display image may be combined with multiple pieces of content such as video and audio streams for selection.
[0043] More specifically, image display system 1 of the embodiment is configured as a computer comprising a CPU (Central Processing Unit), a storage device, an I / O control unit, a communication device, and an input / output device. The CPU executes a program in the storage device to realize the functions of image feature analysis unit 22 and conversion unit 5 using emotion value conversion model 51. The CPU also executes the program in the storage device to realize the functions of emotion estimation unit 42 based on biometric data analysis, emotion model mapping unit 61, and display image selection unit 62. The CPU also controls the I / O control unit and communication device based on the program to realize the functions of candidate image acquisition unit 21 and image display control unit 7.
[0044] Furthermore, the image display system 1 may be configured as part of a system for generating a metaverse space on a network or as part of the functions of an EC site, and may include a biometric information sensor 3 and a display device 8 as part of a user terminal.
[0045] The operation of the image display system 1 according to the embodiment will be described below along the steps (procedures) of the display content selection program with reference to the flow chart of FIG.
[0046] In step S1, the emotion information acquisition unit 4 acquires biometric information of the avatar (user) from the biometric information sensor 3 via a communication path.
[0047] In step S2, emotion information acquisition unit 4 calculates analyzed biometric data from the biometric information of the avatar (user) acquired in step S1, and determines an estimated emotion of the avatar (user) based on the analyzed biometric data.
[0048] In step S3, display image analysis unit 2 determines image feature amounts of candidate images that are candidates for display to the avatar (user). Conversion unit 5 using emotion value conversion model 51 then converts the image feature amounts of each candidate image determined by display image analysis unit 2 into emotion values estimated to be evoked in the avatar (user), using emotion value conversion model 51, which is a machine learning model, to determine emotions estimated to be evoked by the candidate images.
[0049] In step S4, the determination unit 6 maps the estimated emotion of the avatar (user) obtained in step S2 and the converted emotion estimated to be evoked by the candidate image obtained in step S3 onto an emotion map.
[0050] In step S5, the determination unit 6 obtains a converted emotion of the emotion map that satisfies the selection rule for the estimated emotion of the avatar (user) on the emotion map, and selects an image to be shown to the avatar (user).
[0051] In step S6, the image display control unit 7 displays the image selected in step S5 on the display device 8, allowing the avatar (user) to view it.
[0052] After a predetermined time has elapsed, in step S7, emotion information acquisition unit 4 again acquires biometric information of the avatar (user) from biometric sensor 3 via the communication path.
[0053] In step S8, emotion information acquisition unit 4 calculates analyzed biometric data using the avatar (user) biometric information acquired in step S7, and again determines an inferred emotion of the avatar (user) based on the analyzed biometric data. In step S9, update unit 9 of emotion value conversion model 51 receives as input the image feature quantities of the display image selected in step S5 and performs supervised learning of emotion value conversion model 51 using the user's inferred emotion determined in step S8 as ground truth data, thereby updating emotion value conversion model 51. After updating, image display system 1 terminates the process of selecting and displaying in the display image according to the avatar's (user's) emotion.
[0054] The present invention is not limited to the above-described examples, and includes various modifications. The above-described embodiments have been described in detail to clearly explain the present invention, and the present invention is not necessarily limited to those having all of the described configurations.
[0055] REFERENCE SIGNS LIST 1 Image display system (content playback system) 2 Display image analysis unit (content analysis unit) 21 Candidate image acquisition unit (candidate content acquisition unit) 22 Image feature analysis unit (feature analysis unit) 3 Biometric information sensor 4 Emotion information acquisition unit 41 Biometric data acquisition unit 42 Emotion estimation unit 5 Conversion unit 51 Emotion value conversion model 6 Determination unit 61 Emotion model mapping unit 62 Display image selection unit (content selection unit) 7 Image display control unit (playback control unit) 8 Display device (playback unit)
Claims
1. A content display system comprising: a content analysis unit that analyzes feature quantities of candidate content to be played back by a playback unit; an emotion information acquisition unit that determines an estimated emotion of a user who can view the playback unit based on biometric information of the user; and a determination unit that selects content to be played back by the playback unit based on the emotion estimated to be evoked by the candidate content converted from the feature quantities of the candidate content, the estimated emotion of the user, and a selection rule.
2. A content display system according to claim 1, wherein the determination unit maps the emotion estimated to be evoked by the candidate content and the user's estimated emotion onto an emotion map, and selects candidate content on the emotion map that satisfies the selection rule for the estimated emotion.
3. A content display system according to claim 2, wherein the selection rule is a relationship of position or distance between the emotion estimated to be evoked by the candidate content on the emotion map and the estimated emotion of the user.
4. A content display system according to claim 1, further comprising an emotional value conversion model, which is a machine learning model that inputs feature quantities of content and outputs emotional values estimated to be evoked by the image, and which converts the feature quantities into emotions estimated to be evoked by the candidate image using the emotional value conversion model.
5. A content display system according to claim 1, wherein the emotion information acquisition unit estimates the user's emotion based on biometric information detected by at least one of a pupil diameter sensor, a gaze measurement sensor, and a facial expression reading sensor.
6. A display content selection method comprising: a step of analyzing features of candidate content to be played back by a playback unit; a step of determining an estimated emotion of a user who can view the playback unit based on biometric information of the user; a step of converting the features of the candidate content into an emotion estimated to be evoked by the candidate content; a step of mapping the emotion estimated to be evoked by the candidate content and the estimated emotion of the user on an emotion map; and a step of selecting candidate content on the emotion map that satisfies a selection rule for the estimated emotion, and displaying the selected candidate content on a display unit.
7. A display content selection program for causing a computer to execute the following steps: analyzing the features of candidate content to be played back by a playback unit; determining an estimated emotion of a user who can view the playback unit based on the user's biometric information; converting the features of the candidate content into an emotion estimated to be evoked by the candidate content; mapping the emotion estimated to be evoked by the candidate content and the user's estimated emotion on an emotion map; and selecting candidate content on the emotion map that satisfies a selection rule for the estimated emotion and displaying the selected candidate content on a display unit.
Citation Information
Patent Citations
Information processing device, information processing method, and program
JP2022030832A
Emotion estimation device, emotion estimation method and program
JP2012059107A
Information processing device and emotion induction method
JP2023027783A
Taste determination system, taste determination method, and program
JP2023175013A
System and method for enhanced training using a virtual reality environment and bio-signal data
US20160077547A1