Display control device, display device, television receiver, display control system, and program

The display control device and system address the challenge of presenting personalized notification information by acquiring and analyzing user data to generate tailored content and speech, improving user satisfaction.

JP2026031404APending Publication Date: 2026-02-24SHARP KK
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2025101414
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-06-17
Publication Date
2026-02-24

AI Technical Summary

Technical Problem

Existing television technologies fail to effectively present notification information in accordance with user preferences, leading to suboptimal user satisfaction.

Method used

A display control device and system that includes units for acquiring content and user data, analyzing preferences, and generating personalized notification modes based on user responses and preferences, using a trained model to generate tailored content and speech.

Benefits of technology

Enables the presentation of notification information that aligns with user preferences, enhancing user satisfaction and interaction through personalized content and speech.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026031404000001_ABST
    Figure 2026031404000001_ABST
Patent Text Reader

Abstract

To present notification information according to a user's preference.SOLUTION: A display control device (100) includes a first acquisition unit (11) that acquires content data related to content, a second acquisition unit (21) that acquires user identification information for identifying a user, a preference information acquisition unit (21) that acquires preference information corresponding to the user identification information, and a control unit (22) that determines a notification mode of notification information to be notified to the user on the basis of the preference information.SELECTED DRAWING: Figure 4
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a display control device, a display device, a television receiver, a display control system, and a program. [Background technology]

[0002] Conventionally, television devices have been proposed that automatically provide personal media preferences based on user identification information detected by a camera. For example, in the technology described in Patent Document 1, a television device identifies a user based on pre-registered user information and a facial image detected by a camera, and displays media content related to the identified user or automatically changes settings. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Special Publication 2011-504710 Summary of the Invention [Problem to be solved by the invention]

[0004] In order to improve user satisfaction in television viewing, it is important to consider what information to provide to users, but the technology described in Patent Document 1 has had problems in this regard.

[0005] An object of one aspect of the present invention is to realize a technology that can present notification information in accordance with a user's preferences. [Means for solving the problem]

[0006] In order to solve the above problem, a display control device according to one embodiment of the present invention includes a first acquisition unit that acquires content data related to content, a second acquisition unit that acquires user identification information that identifies a user, a preference information acquisition unit that acquires preference information corresponding to the user identification information, and a control unit that determines the notification mode of notification information to be notified to the user based on the preference information.

[0007] A display control device according to one embodiment of the present invention includes a first acquisition unit that acquires content data related to content, a third acquisition unit that acquires user response information related to user responses, and a second control unit that generates preference information related to the user's preferences based on the user response information.

[0008] A display control system according to one aspect of the present invention includes a receiving device that receives content data relating to content, a display control device that controls the display of the content indicated by the content data, an input device that accepts input of user identification information relating to a user, a notification information control device that changes the notification mode of notification information notified to the user in accordance with the user information, and an output device that outputs at least one of the content and the notification information.

[0009] A display control system according to one embodiment of the present invention includes an acquisition unit that acquires content data, a first control unit that controls the display of content indicated by the content data, and a second control unit that outputs speech content generated by a trained model to one or more users and generates preference information indicating the preferences of the users by referring to the users' responses to the speech content. [Effects of the Invention]

[0010] According to one aspect of the present invention, notification information can be presented in accordance with the user's preferences. [Brief explanation of the drawings]

[0011] [Figure 1]1 is a block diagram showing a configuration of a display device according to a first embodiment of the present invention. [Figure 2] FIG. 1 is a diagram illustrating a display device according to a first embodiment of the present invention. [Figure 3] 1 is a diagram showing an example of a display on the display device according to the first embodiment of the present invention. [Figure 4] 1 is a block diagram showing a configuration of a display control device according to a first embodiment of the present invention. [Figure 5] 4 is a flowchart showing the flow of a preference information accumulation process performed by the display control device according to the first embodiment of the present invention. FIG. [Figure 6] 4 is a flowchart showing the flow of a preference information utilization process performed by the display control device according to the first embodiment of the present invention. FIG. [Figure 7] FIG. 10 is a diagram showing an example in which only notification information is displayed on a display device. [Figure 8] FIG. 10 is a diagram showing an example of displaying notification information together with content. [Figure 9] FIG. 10 is a block diagram showing an example of the configuration of a display control device according to a second embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0012] [Embodiment 1] Hereinafter, one embodiment of the present invention will be described in detail. <Display device 1> FIG. 1 is a block diagram showing the configuration of a display device 1 according to this embodiment. As shown in FIG. 1, the display device 1 includes a receiving unit 10, a sensing unit 20, a display control device 100, a display unit 30, and a speaker 40. The components shown in FIG. 1 are merely examples of the configuration of the display device 1 and are not limited to these components. For example, the display device 1 may include various components such as a remote control device operated by a user, an operation signal receiving unit that receives an operation signal from the remote control device, and a storage unit that stores video data. Furthermore, the components included in the display device 1 may be distributed across multiple devices via a network, for example. In such a case, the display device 1 may be referred to as a display system, and the display control device 100 may be referred to as an information processing device or an information processing system. In the example shown in FIG. 1, the display device 1 includes the display control device 100. However, this embodiment is not limited to this configuration. The display control device 100 may be implemented as a set-top box connected to the display device 1.

[0013] (Receiver 10) The receiving unit 10 receives content data and supplies the received content data to the display control device 100. The content data received by the receiving unit 10 includes, for example, encoded video data, encoded audio data, and related information accompanying the video data, but this does not limit the present embodiment. Furthermore, examples of the encoded video data include data (e.g., TS (Transport Steam)) encoded by various video encoding technologies such as MPEG2, MPEG4, H.264, and H.265, but this does not limit the present embodiment.

[0014] Furthermore, the related information may include, for example, at least one of program information related to the content, program guide data including the program information, information on data broadcasting provided in association with the content, and explanatory information related to the content, but these examples do not limit the present embodiment. Furthermore, for example, the related information may be information that can be extracted from the content data without performing a decoding process using a video encoding technology such as MPEG2, as described above. However, this example does not limit the present embodiment.

[0015] The receiving unit 10 may also be configured to acquire the content data from the Internet via wireless or wired communication. The receiving unit 10 may also be configured to acquire the content data from broadcast waves and include a tuner for selecting one of a plurality of channels included in the broadcast waves. When the receiving unit 10 includes a tuner, the display device 1 is also referred to as a television receiver (or simply a television). In other words, a television receiver may include the display device 1. Although the receiving unit 10 is provided outside the display control device 100 in FIG. 1, it may also be built into the display control device 100.

[0016] (Sensing unit 20) The sensing unit 20 senses one or more users who use the display device 1. As an example, the sensing unit 20 includes one or more cameras and captures images of the one or more users using the cameras. The sensing unit 20 then supplies sensing data including image data captured by the cameras to the display control device 100. The sensing unit 20 may also include a laser scanner (LiDAR device) that detects users by reflecting laser light, and may be configured to include scan data acquired by the laser scanner in the sensing data and supply the data to the display control device 100. The sensing unit 20 may also include one or more microphones that collect speech from the one or more users, and may be configured to include audio data indicating the audio collected by the microphones in the sensing data and supply the data to the display control device 100.

[0017] (Display section 30) The display unit 30 displays a display image generated by the display control device 100. The display image includes at least one of a still image and a moving image (video). As will be described later, the display image includes, for example, at least one of a video of the content indicated by the content data (also referred to as a content video) and notification information for the user. The display unit 30 may also be configured to display at least one of a video of the content indicated by the content data (content video) and notification information for the user.

[0018] As an example, the display unit 30 may be configured to include a display panel and a driver that drives the display panel based on image data of the display image, but this does not limit the present embodiment. Furthermore, a liquid crystal panel or an organic EL panel may be used as the display panel, but this does not limit the present embodiment. Display examples by the display unit 30 will be described later.

[0019] (Speaker 40) The speaker 40 outputs the sound of the content indicated by the content data to the user. If the notification information includes sound (also called notification sound), the speaker 40 outputs the notification sound to the user. In this embodiment, the term "audio" may include, but is not limited to, a human voice, and refers to any sound that propagates through a medium such as air.

[0020] (Example of using display device 1) FIG. 2 is a diagram showing an example of how the display device 1 is used. In the example shown in FIG. 2, the display device 1 is realized as a stationary display device with legs. As shown in FIG. 2, the display device 1 is configured to include a sensing unit 20 that senses a user U, a display unit 30 that displays display data, and speakers 40 (two in the example of FIG. 2) that output audio of content and audio of notification information. The display device 1 may also be realized as a wall-mounted display device. The display device 1 may include the above-mentioned display unit 30 and a display control device 100 described below.

[0021] <Display control device 100> Returning to FIG. 1, the configuration of each unit of the display control device 100 included in the display device 1 will be described. As shown in FIG. 1, the display control device 100 includes a first acquisition unit 11, a first control unit 12, a synthesis unit 13, a second acquisition unit 21, and a second control unit 22. Note that the designations "first," "second," and the like do not limit this embodiment. For example, either or both of the "first acquisition unit 11" and the "second acquisition unit 21" may be simply referred to as an "acquisition unit," and either or both of the "first control unit 12" and the "second control unit 22" may be simply referred to as a "control unit." Furthermore, the first control unit 12 may be referred to as a content display control unit.

[0022] (First acquisition unit 11) The first acquisition unit 11 acquires content data related to content. As an example, the first acquisition unit 11 acquires content data received by the above-mentioned receiving unit 10. The first acquisition unit 11 supplies the acquired content data to the first control unit 12.

[0023] (First control unit 12) The first control unit 12 controls the display of the content indicated by the content data supplied from the first acquisition unit 11 (in other words, the display of the content video). As an example, the first control unit 12 decodes video data included in the content data, and supplies the content video obtained by the decoding process to the synthesis unit 13. The content video forms a part of the display image displayed by the display unit 30. The first control unit 12 also decodes audio data included in the content data, and supplies the content audio obtained by the decoding process to the synthesis unit 13. The content audio forms a part of the audio output by the speaker 40. The first control unit 12 also extracts the above-mentioned related information from the content data, and supplies the extracted related information to the second control unit 22, for example.

[0024] (Second acquisition unit 21) The second acquisition unit 21 acquires sensing data of one or more users who use the display device 1 from the sensing unit 20. Here, as described above, the sensing data includes at least one of imaging data that includes the users in the angle of view and audio data that includes speech by the users. The second acquisition unit 21 supplies the sensing data to the second control unit 22.

[0025] The sensing data is an example of user information related to the user of the display device 1. Therefore, the second acquisition unit 21 may be expressed as acquiring user information related to the user of the display device 1. In this embodiment, the user information includes imaging data and audio data acquired by the sensing unit 20 or another device.

[0026] Furthermore, the sensing data is an example of user identification information for identifying the user of the display device 1. Therefore, the second acquisition unit may be expressed as acquiring user identification information for identifying the user of the display device 1.

[0027] Furthermore, the second acquisition unit 21 may be configured to acquire preference information corresponding to the user identification information. As an example, the second acquisition unit 21 may be configured to function as a preference information acquisition unit that acquires preference information corresponding to the user identification information from a preference information storage unit (a preference information DB 23 described later) that stores preference information of a plurality of users. Alternatively, the preference information acquisition unit may have a configuration different from the second acquisition unit 21. Hereinafter, the second acquisition unit 21 will also be referred to as the preference information acquisition unit 21. The preference information storage unit (preference information DB 23) stores preference information for each user generated by a preference information generation unit 206 described later.

[0028] Preference information corresponding to user identification information is preference information of the user identified by the user identification information. Preference information is information about the content in which the user is interested or prefers. While the specific content of preference information is not limited, examples of preference information directly related to content viewing include the type of content preferred (TV programs or their genres, viewing times, etc.), characters (performers, celebrities, anime or game characters, animals, sports teams or affiliated athletes, etc.), topics, and themes. Preference information settings related to device usability include UI layout settings, the order of recording lists and program guides (genre, recording (and broadcast) date and time, broadcast station order), UI design settings (font, design, font size, color), assignment of functions to remote control buttons, and startup time settings.

[0029] (Second control unit 22) The second control unit 22 executes an analysis process of analyzing at least one of the content data, the content video, the content audio, the related information, and the sensing data, and performs various processes by referring to the analysis results. As an example, the second control unit 22 executes at least one of a process of identifying the content by analyzing the content video or the related information, and a process of identifying utterances by the one or more users or a state of the one or more users by analyzing the sensing data, and performs various processes by referring to the results of these processes.

[0030] As an example, the second control unit 22 determines (generates) the notification mode of the notification information to be notified to the user by referring to the sensing data. In other words, the second control unit 22 changes the notification mode of the notification information to be notified to the user in accordance with the user information. Here, the notification information includes at least one of a predetermined image and speech via the image. For example, the notification information includes an object whose display mode and / or speech mode changes in accordance with at least one of the content data and the sensing data. The speech mode includes the content of the speech and tone of voice, etc. Here, the object is an example of a predetermined image. Furthermore, as an example, the object may be a character generated by the character generation unit 204 (described later), an image obtained by applying CG processing to an image externally acquired by the sensing unit 20, or an image created by an image generation AI from text data related to speech. Therefore, the notification mode of the notification information may be expressed as including at least one of a display mode of a predetermined image in accordance with the user information and a speech mode of speech in accordance with the user information.

[0031] Further, as one example, the second control unit 22 may perform a process of determining a speech style suitable for personal information identified based on the user information. For example, the second control unit 22 may perform a process of determining a speech style (such as a speech style for elderly men / a speech style for children) according to the age group or gender, etc., of the individual identified based on the user information (such as an elderly man / a child).

[0032] Further, for example, the second control unit 22 may perform a process of identifying the relative position of the user by referring to the imaging data, and changing at least one of the display mode and speech mode of the alarm information according to the identified relative position. Here, the relative position indicates the position of the user relative to the display device 1, and includes, for example, the direction from the display device 1 and the distance from the display device 1.

[0033] The speech style may include speech content, which is the content of the speech, and the second control unit 22 may determine the speech content to the user by referring to at least one of content and user information. The second control unit 22 may also be configured to control the timing of speech to the user based on at least one of content data and user information.

[0034] As an example, the second control unit 22 may perform at least one of the following processes: by referring to the sensing data, a process of identifying the number of one or more users; and a process of determining whether at least one of the one or more users is a registered user; and, depending on the results of the executed processes, determine the content of the speech to the user (for example, from a character) via a specified image.

[0035] Furthermore, second control unit 22 may control the speech of the character by determining whether to speak, the timing of the speech, and / or the content of the speech. Second control unit 22 may also be configured to control the display mode of the character in response to the control of the speech from the character.

[0036] The second control unit 22 may also be configured to control speech via an object (for example, speech from a character) based on content or content data. Here, "control based on content or content data" includes, for example, control based on at least one of content video, content audio, and related information obtained from the content data. Also, for example, the second control unit 22 may be configured to control speech in accordance with analysis results related to the content obtained by analyzing the content or content data. Here, "analyzing content" refers to analyzing content video and content audio, but is not limited to this and may also include analyzing related information. The analysis results may be analyzed by the display control device 100 or the display device 1, or may be obtained from the cloud.

[0037] The second control unit 22 may be configured to determine the notification mode of the notification information to be notified to the user based on preference information. In other words, the second control unit 22 may be configured to determine the notification mode of the notification information in accordance with the preference information of the user identified with reference to the sensing data. As an example, the second control unit 22 may be configured to determine at least one of the display mode and speech mode of a character (image) included in the notification information with reference to preference information.

[0038] Furthermore, the second control unit 22 may be configured to perform at least one of the following processes: determining the display mode of an image based on the preference information; and generating speech by determining the speech mode based on the preference information. As an example, the second control unit 22 may be configured to perform at least one of the following processes: determining a character to be displayed together with content from a plurality of character candidates by referring to the preference information; and selecting a voice tone for the character's speech from a plurality of voice tone candidates by referring to the preference information.

[0039] (Synthesis section 13) The synthesis unit 13 synthesizes the notification information determined (generated) by the second control unit 22 with the content video supplied from the first control unit 12. More specifically, the synthesis unit 13 synthesizes the notification video included in the notification information determined (generated) by the second control unit 22 with the content video supplied from the first control unit 12 to generate a display image including the content video and the notification video, and supplies the generated display image to the display unit 30. In addition, the synthesis unit 13 generates output audio including the notification audio and the content audio by integrating the notification audio included in the above notification information with the content audio supplied from the first control unit 12, and supplies the generated output audio to the speaker 40. Note that the above example does not limit the present embodiment, and as an example, the display image may be configured to include one of a content video and a notification video but not the other. More specifically, as an example, the display image displayed by the display unit 30 may be configured to include a content video or a notification video. Similarly, the output audio may be configured to include one of a content audio and a notification audio but not the other. More specifically, as an example, the output display audio output by the speaker 40 may be configured to include a content audio or a notification audio.

[0040] FIG. 3 shows a display example of a display image generated by the synthesis unit 13. As shown in FIG. 3, the display image generated by the synthesis unit 13 and displayed by the display unit 30 includes a first region R1 that displays a content video and a second region R2 that displays notification information. The second region R2 also includes a character CR included in the notification information, text data UC indicating the content of the utterance by the character CR, and text data UU indicating the utterance by the user (corresponding to "you" in FIG. 3). At least one of the display manner and the utterance manner of the character changes depending on at least one of the content data and the sensing data. Note that the example shown in FIG. 3 does not limit this embodiment. As an example, the display image may be configured not to include the first region R1.

[0041] Although specific examples of the character CR do not limit the present embodiment, as an example, the character may be any of a character simulating a user, a character simulating a living body such as an animal, a character simulating a non-living body such as a robot, or a character other than the above. Note that the character CR may also be referred to as an avatar.

[0042] (Specific Configuration Example of Display Control Device 100) Next, a specific configuration example of the display control device 100 according to this embodiment will be described with reference to Fig. 4. Fig. 4 is a block diagram showing a specific configuration example of the display control device 100. Note that in the following description, overlapping descriptions of matters that have already been described about the display control device 100 may be omitted.

[0043] (First control unit 12) As shown in FIG. 4, the first control unit 12 includes, for example, a content playback unit 121. The content playback unit 121 decodes video data included in the content data and supplies the decoded content video to the display unit 30. The content playback unit 121 generates decoded content video using a decoding process compliant with various video encoding technologies such as MPEG2, MPEG4, H.264, and H.265, for example. The content playback unit 121 also extracts related information included in the content data from the content data and supplies the extracted related information to the second control unit 22. The content playback unit 121 also decodes audio data included in the content data and supplies the decoded content audio to the speaker 40. Note that, as described in FIG. 1, the display control device 100 may be configured to include a synthesis unit 13, and the first control unit 12 may be configured to supply the content video and the content audio to the synthesis unit 13.

[0044] (Second control unit 22) 4, the second control unit 22 includes an analysis unit 221 and a notification information generation unit 222. The analysis unit 221 analyzes the sensing data supplied from the second acquisition unit 21, the decoded content video supplied from the content playback unit 121, and the related information supplied from the content playback unit 121.

[0045] The analysis processing by the analysis unit 221 may include, for example, a process of identifying the relative position of each of one or more users by analyzing the imaging data included in the sensing data, a process of identifying the state (posture, facial expression, emotion, etc.) of each of one or more users by analyzing the imaging data included in the sensing data, a process of identifying the speech content of each of one or more users by analyzing the audio data included in the sensing data, a process of identifying the relative position of each of one or more users by analyzing the audio data included in the sensing data, a process of identifying the content of the content at each point in time (scene, characters, actions of characters, speech of characters, etc.) by analyzing the content video, and a process of identifying the title, characters, plot, etc. of the content shown in the content video by analyzing related information, but these examples do not limit this embodiment.

[0046] Note that, for example, the various analytical processes performed by the analysis unit 221 can use a machine-learned inference model (prediction model) that executes various algorithms such as an object detection algorithm or an utterance extraction algorithm, but this does not limit the present embodiment. Furthermore, for example, the inference model may be configured as part of a trained model LM (described later), or may be realized as a model separate from the trained model LM (described later). Furthermore, the inference model may be configured to be included in the server device 200 (described later), or may be configured to be included in the display control device 100. The results of the above analytical processes performed by the analysis unit 221 are supplied to the alarm information generation unit 222.

[0047] The notification information generation unit 222 refers to the analysis result by the analysis unit 221 and generates notification information according to the analysis result. The notification video included in the generated notification information is supplied to the display unit 30 and displayed. Note that the display control device 100 may be configured to include a synthesis unit 13, which synthesizes the notification video and the content video. Furthermore, audio data included in the generated notification information is output from the speaker 40.

[0048] As shown in FIG. 4, the notification information generation unit 222 includes, for example, a prompt generation unit 201, an utterance content acquisition unit 202, a voice generation unit 203, a character generation unit 204, an utterance control unit 205, and a preference information generation unit 206.

[0049] (Prompt generation unit 201) The prompt generation unit 201 generates input data (prompt) to be input to the trained model LM by referring to the analysis result by the analysis unit 221, and inputs the generated input data to the trained model LM. Here, the trained model LM may be configured, for example, to be provided in a server device 200 connected to the display control device 100 via a network N, as shown in FIG. 4, or may be configured to be provided in the display control device 100. Furthermore, the trained model LM may be, for example, a large-scale language model, but is not limited to this. Any machine-learned generative model can be used as the trained model LM.

[0050] The specific prompt generated by the prompt generation unit 201 is not limited to this embodiment, but as an example, the prompt can be configured to include reference information including the content of the dialogue between the user and the character up to the present time, the content of the current content, the current state of the user, etc., and instruction information to generate speech content for the user based on the above reference information.

[0051] (Utterance content acquisition unit 202) The utterance content acquisition unit 202 acquires the utterance content generated by the trained model LM based on the prompt generated by the prompt generation unit 201. As an example, the utterance content acquisition unit 202 acquires the utterance content in the form of text data, but this does not limit the present embodiment.

[0052] The processing by the prompt generation unit 201 and the utterance content acquisition unit 202 can also be expressed as processing to generate input data (prompts) including the results of the analysis processing by the analysis unit 221 (information obtained from user identification information, such as the user's utterances), and input the generated input data (prompts) into the trained model LM, thereby determining the utterance content from the character to the user via images. The input data may further include preference information.

[0053] (Speech generation unit 203) The voice generation unit 203 generates voice data indicating the utterance content acquired by the utterance content acquisition unit 202. Here, the voice generation unit 203 may determine the tone of voice and accent of the utterance of the voice data by referring to the analysis result by the analysis unit 221. Alternatively, the prompt generated by the prompt generation unit 201 may include instruction information for specifying the tone of voice, and the voice generation unit 203 may determine the tone of voice and accent of the utterance of the voice data by referring to the response of the trained model LM to the instruction. The voice generated by the voice generation unit 203 is supplied to the speaker 40.

[0054] (Character generation unit 204) The character generation unit 204 generates notification information with reference to the analysis result by the analysis unit 221. As an example, the character generation unit 204 generates (determines) video data of a character (character video data) to be included in the notification information with reference to the analysis result by the analysis unit 221. As an example, the character generation unit 204 determines the character's appearance, the character's costume (clothing), more specifically, the character's gender, appearance and look (hairstyle, facial features, physique), clothing and accessories, and the character's facial expression and posture (face and body direction) when speaking, and the character's actions such as behavior (gestures), and supplies a character video expressing the determined content to the display unit 30. Note that, as described in FIG. 1, the display control device 100 may be configured to include a synthesis unit 13, and the character generation unit 204 may be configured to supply the character video to the synthesis unit 13.

[0055] (Speech control unit 205) The speech control unit 205 supplies the speech content from the character to the user, determined by the speech content acquisition unit 202, to the voice generation unit 203 and / or the character generation unit 204. In the present disclosure, "A and / or B" means either A only, B only, or both A and B. As described above, the voice generation unit 203 generates voice data from the received speech content and supplies it to the speaker 40. As a result, the voice of the character speaking to the user is played from the speaker 40. The character generation unit 204 generates notification information by converting the speech content into text and supplies it to the synthesis unit 13. The synthesis unit 13 displays the text at a predetermined position on the display unit 30. Alternatively, when both voice and text are presented to the user, the synthesis unit 13 displays the text at a predetermined position on the display unit 30 at the same time that the character's voice is played from the speaker 40.

[0056] (Preference information generation unit 206) The preference information generation unit (second control unit) 206 generates user preference information based on the user reaction information. The preference information generation unit 206 generates user preference information using the result of the analysis of the user reaction information by the analysis unit 221. The user reaction information is information related to the user's reaction among the information obtained from the sensing data, and is acquired by the second acquisition unit 21 (also referred to as the "third acquisition unit"). More specifically, the user reaction information is information such as the user's manner of speaking, facial expressions, and actions. Furthermore, the preference information generation unit 206 updates (adds and / or changes) the user's preference information based on the user identification information. Methods for generating and updating preference information will be described later.

[0057] As described above, the display control device 100 may include the first acquisition unit 11 that acquires content data related to content, the second acquisition unit 21 (third acquisition unit) that acquires user reaction information related to user reactions, and the preference information generation unit (second control unit 22) that generates preference information related to user preferences based on the user reaction information. The display control device 100 having such a configuration can ask the user questions and acquire (generate) the user preference information from the answers.

[0058] (Processing flow by the display control device 100) 5 is a flow diagram showing part of the flow of preference information accumulation processing by the display control device 100. The preference information accumulation processing is processing for acquiring information about a user's preferences from sensing data and recording and accumulating it in a database.

[0059] (Step S11) 5, first, in step S11, content data is played back. Specifically, the content playback unit 121 plays back the content received from the first acquisition unit 11 and supplies the content to the synthesis unit 13. Specific examples of content data have been described above, so a description thereof will be omitted here.

[0060] (Step S12) Subsequently, in step S12, the content and the notification information are displayed. Specifically, the synthesis unit 13 generates a display image by synthesizing the content and the notification information, and supplies the image to the display unit 30 for display. The processing of the synthesis unit 13 has been described above, so a description thereof will be omitted here.

[0061] (Step S13) Subsequently, in step S13, a response to the alarm information is acquired. Specifically, the second acquiring unit 21 acquires sensing data including at least one of an image and a voice of the user.

[0062] (Step S14) Next, in step S14, the response is analyzed to generate preference information. Specifically, the analysis unit 221 analyzes the user's speech, facial expressions, and the like included in the sensing data, and the preference information generation unit 206 generates preference information indicating the user's preferences based on the analysis. For example, the analysis unit 221 selects content that the user views frequently, watches for long periods of time, or responds to by yelling or laughing loudly. The analysis unit 221 then analyzes the type, characters, topics, and themes of such content to extract commonalities. The preference information generation unit 206 then generates preference information indicating the user's preferences from the type, characters, topics, and themes of such content. Note that the preference information generation process in step S14 may be performed when sensing data is acquired. In this case, the preference information is updated each time sensing data is acquired, but the timing of this update may be set arbitrarily by the user.

[0063] The display control device 100 also includes an interface that allows direct or indirect inquiries to be made to users to extract user preference information. Specifically, suppose a question is posed to a broadcast program being watched by multiple users, such as a baseball game broadcast, asking, "Do you like baseball?" A first user responds, "I'm a fan of this team, so I always watch when it's on," a second user responds, "We watch TV together as a family at this time," and a third user responds, "My father likes baseball, not me." In this case, the preference information generator 206 generates information indicating that the three people are a family and stores the family habits in the preference information DB 23. In this way, a large amount of information can be collected indirectly, even without direct responses to the question.

[0064] That is, according to the above-described method, by eliciting user preference information through dialogue, it is possible to prevent errors in linking personal information that would otherwise occur if a television operator were directly linked to content viewing information, or if a viewer in front of the television were directly linked to content viewing information. Furthermore, while it is generally conceivable to accumulate data linking a family profile with content viewing information, there is no need to miss an opportunity to link the content with the user who most actively views it.

[0065] Furthermore, when multiple users are in front of the television, the following process may be executed when linking personal information. First, the analysis unit 221 links the voice information of the multiple users acquired by the microphone and the image data (still images or moving images) acquired by the camera to the facial image of the speaker based on the movement of the user's mouth. Next, the notification information generation unit 222 generates a voice for asking the speaker a question. A specific example of a question is asking each speaker their name, such as "Third person from the right, what's your name?" When each of the multiple users in front of the television answers (speaks) their name to the audio question, the facial image, voice information, and name of each user are stored in a linked state in the preference information DB.

[0066] (Step S15) Next, in step S15, the preference information is saved. Specifically, the preference information generating unit 206 saves the updated preference information in the preference information DB (database) 23.

[0067] Through this process, preference information is accumulated for each user. Once the preference information is accumulated, it becomes possible to generate an avatar that is likely to be preferred by the user. In other words, it is possible to identify the user and present notification information such as an avatar generated in accordance with the user's preferences. This allows for a more intimate two-way dialogue between the user and the display device 1.

[0068] The present disclosure also includes a display control system capable of executing the above information processing. For example, the display control system may include a first acquisition unit 11 that acquires content data, a first control unit 12 that controls the display of content indicated by the content data, and a second control unit 22 (particularly including a preference information generation unit 206) that outputs utterance content generated by a trained model to one or more users and generates preference information indicating the preferences of the users by referring to the users' responses to the utterance content.

[0069] (Example of using preference information) FIG. 6 is a flow chart showing a specific example of processing executed by the display control device 100 to utilize user preference information.

[0070] (Step S21) First, in step S21, the display control device 100 identifies a user using sensing data. Specifically, the analysis unit 221 refers to the sensing data obtained by the sensing unit 20 to detect a user positioned in front of the display device 1. Then, the analysis unit 221 analyzes the image or sound to identify the detected user.

[0071] The identification method may be performed, for example, as follows. First, the analysis unit 221 executes a process of identifying the number of users positioned in front of the display device 1. Then, if the number of users is one, a personal identification process for the user is performed. As an example, the analysis unit 221 performs the personal identification process for the user by analyzing a facial image of the user indicated by the imaging data included in the sensing data. If there are multiple users, the analysis unit 221 performs the personal identification process for each user as described above. Furthermore, the analysis unit 221 determines whether the identified user corresponds to any one of multiple users registered in advance. In other words, the analysis unit 221 determines whether the identified user is a registered user.

[0072] (Step S22) Next, in step S22, the display control device 100 acquires the preference information of the identified user from the database. Specifically, if the identified user is a registered user, the character generation unit 204 accesses the preference information DB23 to acquire the preference information of the user. If the identified user is not a registered user, there is no saved preference information, and the display control device 100 does not execute the subsequent processes. In this case, the preference information generation unit 206 may newly register the user and create a data folder for the user in the preference information DB23.

[0073] (Step S23) Next, in step S23, the display control device 100 generates an avatar that reflects the user's preferences. Specifically, the character generation unit 204 references the acquired preference information and generates an avatar that resembles a character that is predicted to be preferred by the user. The clothing worn by the avatar may also be generated by reference to the user's preference information. For example, if the user's preference information indicates that the user has a tendency to enjoy watching baseball broadcasts and is a fan of Team A, it is possible to generate an avatar wearing the uniform of Team A based on this information.

[0074] (Step S24) Next, in step S24, the display control device 100 determines whether or not to make any changes to the avatar. Specifically, for example, the speech control unit 205 generates a speech in which the character asks the user whether the displayed character (avatar) is OK, and supplies the generated speech to the voice generation unit 203 and the character generation unit 204. The speech content is, for example, "Should I change my appearance?" The speech is then displayed on the display unit 30 and played as voice to wait for the user's response. The second acquisition unit 21 acquires the user's response, and the analysis unit 221 analyzes it and determines whether or not the user accepts the response. If the user's response is "That's right (YES)" or the like, the process proceeds to step S25. If the user's response is "It's fine as it is (NO)" or the like, the process proceeds to step S28.

[0075] (Step S25) In step S25, the display control device 100 interacts with the user. Specifically, for example, the speech control unit 205 generates a speech asking the user, for example, "What kind of outfit do you like?" and supplies the speech to the voice generation unit 203 and the character generation unit 204.

[0076] (Step S26) Next, in step S26, the display control device 100 extracts the user's preferences from the dialogue results. Specifically, the analysis unit 221 analyzes the user's answers to extract the user's preferences. For example, if the user answers "I like XX clothing," the analysis unit 221 outputs an analysis result for this user that says "XX clothing is preferred for this avatar."

[0077] (Step S27) Next, in step S27, the display control device 100 changes the avatar by referring to the extracted preference information. Specifically, the character generation unit 204 generates an avatar that resembles the character desired by the user by referring to the user's answer (newly acquired preference information). Alternatively, the character generation unit 204 may change only the clothing worn by the avatar based on the user's answer. Then, the process returns to step S24.

[0078] If the determination in step S24 is NO, the process proceeds to step S28. In step S28, the display control device 100 determines whether or not the preferences have been reflected in the avatar. In other words, it determines whether or not the avatar has been changed at least once in response to a user request (whether or not steps S25 to S27 have been executed at least once). If the determination result is YES, the process proceeds to step S29; if NO, the process ends.

[0079] (Step S28) In step S28, the display control device 100 records the extracted additional preference information in the preference information database. Specifically, the preference information generation unit (second control unit) 206 generates preference information by linking the avatar's utterance with the user's reaction information, and newly adds or changes the user's preference information to the preference information DB 23.

[0080] Through the above processing, the display control device 100 can present information in accordance with the user's preferences. Furthermore, the display control device 100 can accumulate user preferences through two-way dialogue with the user, and can provide an avatar that matches the user's preferences based on more detailed and specific preference information.

[0081] According to the above information processing, data about a user is not input in advance, but can be continuously changed (improved) depending on the user's viewing experience, environment, etc. Furthermore, data about individuals held by the display control device 100 can be utilized in the field of user experience.

[0082] The present disclosure also includes a display control system capable of executing the above-described information processing. For example, the display control system may include a receiving device (receiving unit 10) that receives content data related to content, a display control device 100 that controls display of content indicated by the content data, an input device (second acquiring unit 21) that accepts input of user identification information related to a user, a notification information control device (second control unit 22) that changes a notification mode of notification information to be notified to a user according to the user identification information, and an output device (display unit 30) that outputs at least one of the content and the notification information.

[0083] (Example of notification information display) Fig. 7 shows an example of display by the display unit 30 of the display device 1 that performs the processing according to this example. In the upper part of Fig. 7, the character CR utters "I wonder if this is a good look for me" (UC in R2 in the upper part of Fig. 7), and in response to this utterance, the user utters "I wonder if I like cats better" (UU in R2 in the upper part of Fig. 7).

[0084] In this case, as shown in R2 in the lower part of FIG. 7, the character video data of the character CR is changed to a cat. Then, the character CR utters, "I've transformed into a cat" (UC in R2 in the lower part of FIG. 7) while facing the user after the movement. Note that the user's information, "I think I prefer cats," may be recorded as new preference information in the preference information DB 23. Also, when the character CR speaks, the volume of the content video data may be made lower than the volume at which the character CR speaks. Note that when the above-mentioned dialogue is carried out, displaying the content is not necessarily required. Therefore, nothing may be displayed in the area of ​​the display unit 30 where content is normally displayed. FIG. 7 shows an example in which no content is displayed.

[0085] The above-described processing example is an example of processing by the second control unit 22 to link the utterance of the avatar with the user reaction information to generate and display preference information.

[0086] FIG. 8 is a diagram showing an example in which notification information (a dialogue between a user and an avatar in the example shown in FIG. 8) is displayed together with content on the display unit 30. The conversation between the user and the character CR in the second area R2 in the upper and lower rows of FIG. 8 is the same as in FIG. 7, but in FIG. 8, content video is displayed in the first area R1 in the upper and lower rows. In this way, the notification information may be displayed alone, or may be displayed together with the content. Furthermore, when displayed together with the content, the dialogue between the user and the character CR does not necessarily have to be about the content.

[0087] (Other specific processing examples) The specific processing by the display control device 100 is not limited to the above-described example. For example, when there are multiple users, the notification information generation unit 222 may generate a different avatar for each user. Furthermore, when there are multiple users, an avatar representing all of the multiple users may be generated. For example, when watching with family, an avatar that is close to the child or an avatar for the family may be generated. Furthermore, the mode of the avatar may be generated so that not only the appearance but also the voice, attitude, etc. are appropriate to the appearance. Furthermore, the avatar may be changed depending on the time of day or the event. For example, when watching the Olympics, an avatar that is reminiscent of the Olympics may be generated.

[0088] (Effects of the display control device 100) As described above, the display control device 100 according to this embodiment is configured to include a first acquisition unit that acquires content data related to content, a second acquisition unit that acquires user identification information that identifies a user, a preference information acquisition unit 21 that acquires preference information corresponding to the user identification information, and a control unit that determines the notification mode of notification information to be notified to the user based on the preference information. The display device 1 configured as described above can present notification information in accordance with the user's preferences. As described above, the display control device 100 can either display notification information together with content, or display only the notification information without the content.

[0089] [Embodiment 2] Another embodiment of the present invention will be described below. Fig. 9 is a block diagram showing a specific configuration example of a display control device 100A according to this embodiment. The display control device 100A shown in Fig. 9 has the same configuration as the display control device 100 shown in Fig. 4. Furthermore, the display control device 100A has a generation unit 200A, and the generation unit 200A has a language model LM. The other configurations are the same as those of the display control device 100 shown in Fig. 4, so duplicated explanations will be omitted.

[0090] [Software implementation example] The functions of the display control device 100 (hereinafter referred to as the "device") can be realized by a program that causes a computer to function as the device, and a program that causes a computer to function as each control block of the device (particularly each part included in the first control unit 12 and the second control unit 22).

[0091] In this case, the device includes a computer having at least one control device (e.g., a processor) and at least one storage device (e.g., a memory) as hardware for executing the program. The control device and storage device execute the program, thereby realizing the functions described in each of the above embodiments.

[0092] The program may be non-transitory and may be recorded on one or more computer-readable recording media. The recording media may or may not be included in the device. In the latter case, the program may be supplied to the device via any wired or wireless transmission medium.

[0093] Furthermore, some or all of the functions of the control blocks can be realized by logic circuits. For example, an integrated circuit in which a logic circuit that functions as each of the control blocks is formed is also included in the scope of the present invention. In addition, the functions of the control blocks can also be realized by, for example, a quantum computer.

[0094] Furthermore, each process described in each of the above embodiments may be executed by AI (Artificial Intelligence). In this case, the AI ​​may run on the control device or on another device (for example, an edge computer or a cloud server).

[0095] 〔summary〕 The present specification describes at least the following aspects.

[0096] (Aspect 1) A display control device comprising: a first acquisition unit that acquires content data related to content; a second acquisition unit that acquires user identification information that identifies a user; a preference information acquisition unit that acquires preference information corresponding to the user identification information; and a control unit that determines a notification mode of notification information to be notified to the user based on the preference information. According to the above configuration, the notification mode of the notification information to be notified to the user is determined based on the preference information, so that the notification information can be presented in accordance with the user's preferences.

[0097] (Aspect 2) A display control device as described in aspect 1, wherein the notification information includes at least one of a predetermined image and speech via the image, the notification manner includes at least one of a display manner of the image and a speech manner of the speech, and the control unit determines at least one of the display manner of the image and the speech manner via the image based on the preference information. According to the above configuration, an image is generated in accordance with the user's preferences, and the image is notified to the user with a sound appropriate for the image, so that notification information can be presented in accordance with the user's preferences.

[0098] (Aspect 3) 3. The display control device according to aspect 2, wherein the user identification information includes at least one of imaging data and audio data including speech by the user. According to the above configuration, the user can be identified from the image and sound acquired by the display control device.

[0099] (Aspect 4) The control unit The display control device according to aspect 3 executes at least one of a process of determining a display mode of the image based on the preference information and a process of generating speech by determining the speech mode based on the preference information. According to the above configuration, it is possible to determine the display mode of an image in accordance with the user's preferences and generate the tone of voice for the speech.

[0100] (Aspect 5) The control unit 5. The display control device according to any one of aspects 1 to 4, wherein preference information indicating the preferences of the user is updated based on the user identification information. According to the above configuration, preference information can be added and / or changed for each user.

[0101] (Aspect 6) The display control device according to any one of aspects 1 to 5, wherein the control unit generates input data including information obtained from the user identification information and the preference information, and determines the content of the speech to the user via the image by inputting the generated input data into a trained model. According to the above configuration, it is possible to determine the content of an utterance with high accuracy using a known model such as a large-scale language model.

[0102] (Aspect 7) The display control device according to any one of aspects 1 to 6, wherein the second acquisition unit further includes a preference information generation unit that acquires user response information regarding the user's response and generates the preference information based on the user response information. According to the above configuration, it is possible to generate estimated user preference information by referring to user response information.

[0103] (Aspect 8) 8. The display control device according to aspect 7, wherein the preference information generation unit generates preference information by linking the content of an utterance to the user with the user reaction information. According to the above configuration, since the user's reaction information during the interaction between the user and the avatar is used, more specific preference information can be generated.

[0104] (Aspect 9) A display control device comprising: a first acquisition unit that acquires content data related to content; a third acquisition unit that acquires user response information related to user responses; and a second control unit that generates preference information related to the user's preferences based on the user response information. According to the above configuration, since the user's preference information is generated based on the user's reaction information, the preference information can be updated and accumulated.

[0105] (Aspect 10) A display device comprising a display control device according to any one of aspects 1 to 9 and a display unit that displays at least one of the content and the notification information, wherein the display control device further comprises a content display control unit that controls the display of the content. According to the above display device, the notification mode of the notification information to be notified to the user is determined based on the preference information, so that the notification information can be presented in accordance with the user's preferences.

[0106] (Aspect 11) A television receiver comprising the display device according to embodiment 10. According to the above configuration, notification information can be presented to a user watching the television in accordance with the user's preferences.

[0107] (Aspect 12) A display control system comprising: a receiving device that receives content data relating to content; a display control device that controls the display of the content indicated by the content data; an input device that accepts input of user identification information relating to a user; a notification information control device that changes the notification mode of notification information notified to the user in accordance with the user identification information; and an output device that outputs at least one of the content and the notification information. According to the above configuration, notification information can be presented in accordance with the user's preferences.

[0108] (Aspect 13) A display control system comprising: an acquisition unit that acquires content data; a first control unit that controls the display of content indicated by the content data; and a second control unit that outputs speech content generated by a trained model to one or more users, and generates preference information indicating the preferences of the users by referring to the users' responses to the speech content. According to the above configuration, since the user's preference information is generated by referring to the user's response, the preference information can be updated and accumulated.

[0109] (Aspect 14) A display control program for causing a computer to function as the first acquisition unit, the second acquisition unit, the preference information acquisition unit, and the control unit according to any one of aspects 1 to 8. According to the above configuration, notification information can be presented in accordance with the user's preferences.

[0110] The display control device according to each aspect of the present invention may be realized by a computer. In this case, the program of the display control device that realizes the display control device on a computer by causing the computer to operate as each part (software element) of the display control device, and the computer-readable recording medium on which the program is recorded, also fall within the scope of the present invention.

[0111] The present invention is not limited to the above-described embodiments, and various modifications are possible within the scope of the claims. Embodiments obtained by appropriately combining the technical means disclosed in different embodiments are also included in the technical scope of the present invention. Furthermore, new technical features can be formed by combining the technical means disclosed in each embodiment.

[0112] For example, in the above-described embodiments, the user may be able to select whether or not to display text data. Similarly, the user may be able to select whether or not to display the avatar's speech. Furthermore, the display control device 100 may be implemented as a standalone device, such as an electronic device, including a set-top box as an example. In this case, the electronic device may have at least some of the functions shown in FIG. 1, FIG. 4, or FIG. 9, with the remaining functions being provided elsewhere. Alternatively, the electronic device may have all of these functions. In particular, the illustrated server device 200 may have multiple trained models LM. In this case, all of the multiple trained models LM may be provided in the electronic device, or at least some may be provided in the electronic device, with the other trained models LM provided elsewhere. Note that the "avatar" in the above-described embodiments refers to an object represented based on a character, which can be primarily selected or created by the user. The avatar does not necessarily have to be a user's alter ego as the general term implies; it may simply correspond to a predetermined image. [Explanation of symbols]

[0113] 1...Display device 100 Display control device 11 First acquisition section 12 First control section 21 Second acquisition unit (preference information acquisition unit) 22 Second control section 13. Synthesis section 10. Receiving unit 20 Sensing unit 23 Preference Information DB 30...Display section 40 Speaker 200 Server device

Claims

1. a first acquisition unit that acquires content data related to the content; a second acquisition unit that acquires user identification information that identifies a user; a preference information acquisition unit that acquires preference information corresponding to the user identification information; a control unit that determines a notification mode of notification information to be notified to the user based on the preference information; A display control device comprising:

2. the notification information includes at least one of a predetermined image and a speech via the image, the notification manner includes at least one of a display manner of the image and an utterance manner of the utterance, The control unit determines at least one of a display mode of the image and a speech mode via the image based on the preference information. The display control device according to claim 1 .

3. The user identification information includes at least one of imaging data and voice data including speech by the user. The display control device according to claim 2 .

4. The control unit A process of determining a display mode of the image based on the preference information; and A process of generating an utterance by determining the utterance style based on the preference information. Execute at least one of the following: The display control device according to claim 3 .

5. The control unit updating preference information indicating the preferences of the user based on the user identification information; The display control device according to claim 4 .

6. The control unit generating input data including information obtained from the user identification information and the preference information; The generated input data is input into a trained model to determine the content of the utterance to the user via the image. The display control device according to any one of claims 2 to 5.

7. the second acquisition unit acquires user reaction information regarding a reaction of the user; The system further includes a preference information generating unit that generates the preference information based on the user reaction information. The display control device according to any one of claims 1 to 5.

8. The preference information generating unit generates preference information by linking the content of the utterance to the user with the user reaction information. The display control device according to claim 7 .

9. a first acquisition unit that acquires content data related to the content; a third acquisition unit that acquires user reaction information regarding a reaction of the user; a second control unit that generates preference information related to the user's preferences based on the user reaction information; A display control device comprising:

10. A display control device according to any one of claims 1 to 5; a display unit that displays at least one of the content and the notification information; Equipped with The display device further includes a content display control unit that controls the display of the content.

11. A television receiver comprising the display device according to claim 10.

12. a receiving device for receiving content data relating to the content; a display control device that controls display of the content indicated by the content data; an input device that accepts input of user identification information relating to a user; a notification information control device that changes a notification mode of notification information to be notified to the user in accordance with the user identification information; an output device that outputs at least one of the content and the notification information; A display control system comprising:

13. an acquisition unit that acquires content data; a first control unit that controls display of the content indicated by the content data; a second control unit that outputs utterance content generated by the trained model to one or more users and generates preference information indicating preferences of the users by referring to responses of the users to the utterance content; A display control system comprising:

14. A display control program for causing a computer to function as the first acquisition unit, the second acquisition unit, the preference information acquisition unit, and the control unit according to claim 1 .

Citation Information

Patent Citations

  • Media preferences

    JP2011504710A