Information processing apparatus, display apparatus, television receiver, and information processing system
The information processing device tailors television content and notifications to individual user preferences, improving satisfaction by integrating personalized display and audio modes.
Patent Information
- Application Number
- JP2025083437
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-05-19
- Publication Date
- 2026-02-24
AI Technical Summary
Existing television technologies struggle to provide personalized and suitable information to users based on their identification, leading to suboptimal user satisfaction.
An information processing device that includes units to acquire content and user information, controlling the notification mode of information based on user data, integrating content and personalized notification through a synthesis process.
Enables the presentation of suitable information tailored to individual users, enhancing user satisfaction by adapting display and audio modes to user preferences and context.
Smart Images

Figure 2026031391000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to an information processing device, a display device, a television receiver, and an information processing system. [Background technology]
[0002] Conventionally, television devices have been proposed that automatically provide personal media preferences based on user identification information detected by a camera. For example, in the technology described in Patent Document 1, a television device identifies a user based on pre-registered user information and a facial image detected by a camera, and displays media content related to the identified user or automatically changes settings. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Special Publication 2011-504710 Summary of the Invention [Problem to be solved by the invention]
[0004] In order to improve user satisfaction in television viewing, it is important to consider what information to provide to users, but the technology described in Patent Document 1 has had problems in this regard.
[0005] An object of one aspect of the present invention is to realize a technology that can present suitable information to a user. [Means for solving the problem]
[0006] In order to solve the above problem, an information processing device according to one aspect of the present invention includes a first acquisition unit that acquires content data related to content, a second acquisition unit that acquires user information related to a user, and a control unit that changes the notification mode of notification information notified to the user in accordance with the user information.
[0007] In order to solve the above-described problems, an information processing system according to one aspect of the present invention includes a receiving device that receives content data related to content, a display control device that controls display of content indicated by the content data, an input device that accepts input of user information related to a user, a notification information control device that changes a notification mode of notification information to be notified to the user in accordance with the user information, and an output device that outputs at least one of the content and the notification information. [Effects of the Invention]
[0008] According to one aspect of the present invention, it is possible to present suitable information to a user. [Brief explanation of the drawings]
[0009] [Figure 1] 1 is a block diagram showing a configuration of a display device according to a first embodiment of the present invention. [Figure 2] FIG. 1 is a diagram illustrating a display device according to a first embodiment of the present invention. [Figure 3] 1 is a diagram showing an example of a display on the display device according to the first embodiment of the present invention. [Figure 4] 1 is a block diagram showing a configuration of a display control device according to a first embodiment of the present invention. [Figure 5] 3 is a flowchart showing a processing flow by the display control device according to the first embodiment of the present invention. [Figure 6] 3 is a flowchart showing a processing flow by the display control device according to the first embodiment of the present invention. [Figure 7] 3 is a flowchart showing a processing flow by the display control device according to the first embodiment of the present invention. [Figure 8] 1 is a diagram showing an example of a display on the display device according to the first embodiment of the present invention. [Figure 9] FIG. 2 is a diagram for explaining processing by the display control device according to the first embodiment of the present invention. [Figure 10]3 is a flowchart showing a processing flow by the display control device according to the first embodiment of the present invention. [Figure 11] FIG. 10 is a block diagram showing an example of the configuration of a display control device according to a second embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0010] [Embodiment 1] Hereinafter, one embodiment of the present invention will be described in detail. <Display device 1> FIG. 1 is a block diagram showing the configuration of a display device 1 according to this embodiment. As shown in FIG. 1, the display device 1 includes a receiving unit 10, a sensing unit 20, a display control device 100, a display unit 30, and a speaker 40. Here, the display control device 100 is sometimes referred to as an information processing device. Note that the components shown in FIG. 1 are merely examples of the configuration of the display device 1 and are not limited to these components. For example, the display device 1 may include various components such as a remote control device operated by a user, an operation signal receiving unit that receives an operation signal from the remote control device, and a storage unit that stores video data. Furthermore, the components included in the display device 1 may be distributed across multiple devices via a network, for example. In such a case, the display device 1 may be referred to as a display system, and the display control device 100 may be referred to as a display control system or an information processing system. Note that, in the example shown in FIG. 1, the display device 1 includes the display control device 100, but this embodiment is not limited thereto. The display control device 100 may also be implemented as a set-top box connected to the display device 1.
[0011] (Receiver 10) The receiving unit 10 receives content data and supplies the received content data to the display control device 100. The content data received by the receiving unit 10 includes, for example, encoded video data, encoded audio data, and related information accompanying the video data, but this does not limit the present embodiment. Furthermore, examples of the encoded video data include data (e.g., TS (Transport Steam)) encoded by various video encoding technologies such as MPEG2, MPEG4, H.264, and H.265, but this does not limit the present embodiment.
[0012] Furthermore, the related information may include, for example, at least one of program information related to the content, program guide data including the program information, information on data broadcasting provided in relation to the content, and explanatory information related to the content, but these examples do not limit the present embodiment. Furthermore, for example, the related information may be information that can be extracted from the content data without performing a decoding process using a video encoding technology such as MPEG2, as described above. However, this example does not limit the present embodiment.
[0013] The receiving unit 10 may also be configured to acquire the content data from the Internet via wireless or wired communication. The receiving unit 10 may also be configured to acquire the content data from broadcast waves and include a tuner for selecting one of multiple channels included in the broadcast waves. When the receiving unit 10 includes a tuner, the display device 1 is also referred to as a television receiver (or simply a television).
[0014] (Sensing unit 20) The sensing unit 20 senses one or more users who use the display device 1. As an example, the sensing unit 20 includes one or more cameras and captures images of the one or more users using the cameras. The sensing unit 20 then supplies sensing data including image data captured by the cameras to the display control device 100. The sensing unit 20 may also include, for example, a laser scanner (LiDAR device), millimeter-wave radar, or ultrasonic sensor that detects users by reflecting laser light, and may be configured to include scan data acquired by the laser scanner in the sensing data and supply the data to the display control device 100. The sensing unit 20 may also include one or more microphones that collect speech from the one or more users, and may be configured to include voice data collected by the microphones in the sensing data and supply the data to the display control device 100.
[0015] (Display section 30) The display unit 30 displays the display data generated by the display control device 100. As will be described later, the display data includes, for example, at least one of video data of the content indicated by the content data (content video data) and notification information. The display unit 30 may also be configured to display at least one of video of the content indicated by the content data (content video) and notification information for the user.
[0016] As an example, the display unit 30 may be configured to include a display panel and a driver that drives the display panel based on image data of the display image, but this does not limit the present embodiment. Furthermore, a liquid crystal panel or an organic EL panel may be used as the display panel, but this does not limit the present embodiment. Display examples by the display unit 30 will be described later.
[0017] (Speaker 40) The speaker 40 outputs the audio of the content indicated by the content data to the user. If the notification information includes audio (also called notification audio), the speaker 40 outputs the notification audio to the user. Note that in this embodiment, the term "audio" may include a human voice, but is not limited to this, and refers to any sound propagating through a medium such as air.
[0018] (Example of using display device 1) FIG. 2 is a diagram showing an example of how the display device 1 is used. In the example shown in FIG. 2, the display device 1 is realized as a stationary display device with legs. As shown in FIG. 2, the display device 1 is configured to include a sensing unit 20 that senses a user U, a display unit 30 that displays display data, and speakers 40 (two in the example of FIG. 2) that output audio of content and audio of notification information. The display device 1 may also be realized as a wall-mounted display device.
[0019] <Display control device 100> Returning to FIG. 1, the configuration of each unit of the display control device 100 included in the display device 1 will be described. As shown in FIG. 1, the display control device 100 includes a first acquisition unit 11, a first control unit 12, a synthesis unit 13, a second acquisition unit 21, and a second control unit 22. Note that the designations "first," "second," and the like do not limit this embodiment. For example, either or both of the "first acquisition unit 11" and the "second acquisition unit 21" may be simply referred to as an "acquisition unit," and either or both of the "first control unit 12" and the "second control unit 22" may be simply referred to as a "control unit." Furthermore, the first control unit 12 may be referred to as a content display control unit.
[0020] (First acquisition unit 11) The first acquisition unit 11 acquires content data related to content. As an example, the first acquisition unit 11 acquires content data received by the above-mentioned receiving unit 10. The first acquisition unit 11 supplies the acquired content data to the first control unit 12.
[0021] (First control unit 12) The first control unit 12 controls the display of the content indicated by the content data supplied from the first acquisition unit 11 (in other words, the display of the content video). As an example, the first control unit 12 decodes video data included in the content data, and supplies the content video obtained by the decoding process to the synthesis unit 13. The content video constitutes part of the display data displayed by the display unit 30. The first control unit 12 also decodes audio data included in the content data, and supplies the content audio obtained by the decoding process to the synthesis unit 13. The content audio constitutes part of the audio output by the speaker 40. The first control unit 12 also extracts the above-mentioned related information from the content data, and supplies the extracted related information to the second control unit 22, for example.
[0022] (Second acquisition unit 21) The second acquisition unit 21 acquires sensing data of one or more users who use the display device 1 from the sensing unit 20. Here, as described above, the sensing data includes at least one of imaging data related to the users and audio data including speech by the users. The second acquisition unit 21 supplies the sensing data to the second control unit 22.
[0023] The sensing data is an example of user information related to the user of the display device 1. Therefore, the second acquisition unit 21 may be expressed as acquiring user information related to the user of the display device 1. In this embodiment, the user information includes imaging data and audio data acquired by the sensing unit 20 or another device.
[0024] The sensing data is an example of user identification information for identifying the user of the display device 1. Therefore, the second acquisition unit 21 may be expressed as acquiring user identification information for identifying the user of the display device 1.
[0025] The second acquiring unit 21 may be configured to acquire preference information corresponding to the user identification information. As an example, the second acquiring unit 21 may be configured to function as a preference information acquiring unit that acquires preference information corresponding to the user identification information from a preference information storage unit (preference information DB) that stores preference information of a plurality of users.
[0026] (Second control unit 22) The second control unit 22 executes an analysis process of analyzing at least one of the content data, the content video, the content audio, the related information, and the sensing data, and performs various processes by referring to the analysis results. As an example, the second control unit 22 executes at least one of a process of identifying the content indicated by the content data by analyzing the content video data or the related information, and a process of identifying utterances by the one or more users or a state of the one or more users by analyzing the sensing data, and performs various processes by referring to the results of these processes.
[0027] As an example, the second control unit 22 determines (generates) the notification mode of the notification information to be notified to the user by referring to the sensing data. In other words, the second control unit 22 changes the notification mode of the notification information to be notified to the user in accordance with the user information. Here, the notification information includes at least one of a predetermined image and speech via the image. For example, the notification information includes an object whose display mode and / or speech mode changes in accordance with at least one of the content data and the sensing data. Here, the object is an example of a predetermined image. Furthermore, as an example, the object may be a character generated by the character generation unit 204 described below, or a CG-processed image obtained from an external source by the sensing unit 20. Therefore, it may be expressed that the notification mode of the notification information includes at least one of a display mode of a predetermined image in accordance with the user information and a speech mode of speech in accordance with the user information.
[0028] Further, as one example, the second control unit 22 may perform a process of determining a speech style suitable for personal information identified based on the user information. For example, the second control unit 22 may perform a process of determining a speech style (such as a speech style for elderly men / a speech style for children) according to the age group or gender, etc., of the individual identified based on the user information (such as an elderly man / a child).
[0029] Further, for example, the second control unit 22 may refer to the imaging data to identify the relative position of the user, and change at least one of the display mode and speech mode of the alarm information according to the identified relative position. Here, the relative position indicates the position of the user relative to the display device 1, and includes, for example, the direction from the display device 1 and the distance from the display device 1.
[0030] The speech style may include speech content, which is the content of the speech, and the second control unit 22 may determine the speech content to the user by referring to at least one of content and user information. The second control unit 22 may also be configured to control the timing of speech to the user based on at least one of content data and user information.
[0031] As an example, the second control unit 22 may refer to the sensing data to perform at least one of the following processes: a process of identifying the number of one or more users; and a process of determining whether at least one of the one or more users is a registered user; and, depending on the results of the executed processes, may perform a process of determining the content of a speech to be made to the user (for example, from a character) via a specified image.
[0032] Furthermore, second control unit 22 may control the speech of the character by determining whether to speak, the timing of the speech, and / or the content of the speech. Second control unit 22 may also be configured to control the display mode of the character in response to the control of the speech from the character.
[0033] The second control unit 22 may also be configured to control speech via an object (e.g., speech from a character) based on content or content data. Here, "control based on content or content data" includes, for example, control based on at least one of content video, content audio, and related information obtained from the content data. Also, for example, the second control unit 22 may be configured to control speech based on analysis results related to the content obtained by analyzing the content or content data. Here, "analyzing content" refers to analyzing content video and content audio, but is not limited to this and may also include analyzing related information. Related information may include various data specified by broadcasting standards, data related to an electronic program guide, data obtained from a web search based on data obtained from a broadcast, and the like. The analysis results may be analyzed by the display control device 100 or the display device 1, or may be obtained from the cloud.
[0034] The second control unit 22 may be configured to determine the notification mode of the notification information to be notified to the user based on preference information. In other words, the second control unit 22 may be configured to determine the notification mode of the notification information in accordance with the preference information of the user identified with reference to the sensing data. As an example, the second control unit 22 may be configured to determine at least one of the display mode and speech mode of the character included in the notification information with reference to the preference information.
[0035] The second control unit 22 may also be configured to perform at least one of a process for determining the display mode of an image based on the preference information and a process for generating speech based on the preference information. As an example, the second control unit 22 may also be configured to perform at least one of a process for determining a character to be displayed by referring to the preference information and a process for determining a tone of voice for the character's speech by referring to the preference information. Note that data related to the character, such as the tone of voice for the character's speech or the display mode, may be selected from data pre-stored in the display device 1.
[0036] (Synthesis section 13) The synthesis unit 13 synthesizes the notification information determined (generated) by the second control unit 22 with the content video data supplied from the first control unit 12. More specifically, the synthesis unit 13 synthesizes the video data included in the notification information determined (generated) by the second control unit 22 with the content video data supplied from the first control unit 12 to generate display data, and supplies the generated display data to the display unit 30.
[0037] In addition, the synthesis unit 13 generates output audio including the notification audio and the content audio by integrating the notification audio included in the above notification information with the content audio supplied from the first control unit 12, and supplies the generated output audio to the speaker 40. Note that the above example does not limit the present embodiment, and as an example, the display image may be configured to include one of a content video and a notification video but not the other. More specifically, as an example, the display image displayed by the display unit 30 may be configured to include a content video or a notification video. Similarly, the output audio may be configured to include one of a content audio and a notification audio but not the other. More specifically, as an example, the output display audio output by the speaker 40 may be configured to include a content audio or a notification audio.
[0038] FIG. 3 shows a display example of a display image generated by the synthesis unit 13. In the example shown in the upper part of FIG. 3, the display image generated by the synthesis unit 13 and displayed by the display unit 30 includes a first region R1 displaying content video data and a second region R2 displaying notification information. The second region R2 also includes a character CR included in the notification information, text data UC indicating the content of the utterance by the character CR, and text data UU indicating the utterance by the user (corresponding to "you" in FIG. 3). At least one of the display manner and the utterance manner of the character changes depending on at least one of the content data and the sensing data. Note that the example shown in the upper part of FIG. 3 does not limit this embodiment. As an example, as shown in the lower part of FIG. 3, the display image may be configured not to include the first region R1.
[0039] Although specific examples of the character CR do not limit the present embodiment, as an example, the character may be any of a character simulating a user, a character simulating a living body such as an animal, and a character simulating a non-living body such as a robot, or may be a character other than the above. Note that the character CR may also be referred to as an avatar.
[0040] As described above, each unit included in the display device 1 may be distributed across multiple devices via a network, for example. In such a case, the display device 1 may be referred to as a display system, and the display control device 100 may be referred to as a display control system or an information processing system. The display control system according to this embodiment may be described as including, for example, a receiving device (receiving unit 10) that receives content data related to content, a display control device (first control unit 12) that controls the display of content indicated by the content data, an input device that accepts input of user information related to the user, a notification information control device (second control unit 22) that changes the notification mode of notification information to be notified to the user in accordance with the user information, and an output device (display unit 30) that outputs at least one of the content and the notification information. The input device may be realized, for example, as the second acquisition unit 21 described above or an operation signal receiving unit that receives an operation signal from a remote control device operated by a user.
[0041] (Specific Configuration Example of Display Control Device 100) Next, a specific configuration example of the display control device 100 according to this embodiment will be described with reference to Fig. 4. Fig. 4 is a block diagram showing a specific configuration example of the display control device 100. Note that in the following description, overlapping descriptions of matters that have already been described about the display control device 100 may be omitted.
[0042] (First control unit 12) As shown in FIG. 4, the first control unit 12 includes, for example, a content playback unit 121. The content playback unit 121 decodes video data included in the content data and supplies the decoded content video to the display unit 30. The content playback unit 121 generates decoded content video data using a decoding process compliant with various video encoding technologies such as MPEG2, MPEG4, H.264, and H.265, for example. The content playback unit 121 also extracts related information included in the content data from the content data and supplies the extracted related information to the second control unit 22. The content playback unit 121 also decodes audio data included in the content data and supplies the decoded content audio to the speaker 40. Note that, as described in FIG. 1, the display control device 100 may include a synthesis unit 13, and the first control unit 12 may supply the content video and content audio to the synthesis unit 13.
[0043] (Second control unit 22) 4, the second control unit 22 includes an analysis unit 221 and a notification information generation unit 222. The analysis unit 221 analyzes the sensing data supplied from the second acquisition unit 21, the decoded content video data supplied from the content playback unit 121, and the related information supplied from the content playback unit 121. The analysis process by the analysis unit 221 may include, for example, a process of identifying the relative positions of one or more users by analyzing imaging data included in the sensing data, a process of identifying the states (postures, facial expressions, emotions, etc.) of one or more users by analyzing imaging data included in the sensing data, a process of identifying the speech content of one or more users by analyzing audio data included in the sensing data, a process of identifying the relative positions of one or more users by analyzing audio data included in the sensing data, a process of identifying the content of the content at each point in time (scenes, characters, actions of characters, utterances of characters, etc.) by analyzing the content video data, and a process of identifying the title, characters, plot, etc. of the content indicated by the content video data by analyzing the related information; however, these examples do not limit the present embodiment.
[0044] Note that, for example, the various analytical processes performed by the analysis unit 221 can use a machine-learned inference model (prediction model) that executes various algorithms such as an object detection algorithm or an utterance extraction algorithm, but this does not limit the present embodiment. Furthermore, for example, the inference model may be configured as part of a trained model LM (described later), or may be realized as a model separate from the trained model LM (described later). Furthermore, the inference model may be configured to be included in the server device 200 (described later), or may be configured to be included in the display control device 100. The results of the above analytical processes performed by the analysis unit 221 are supplied to the alarm information generation unit 222.
[0045] The notification information generation unit 222 refers to the analysis result by the analysis unit 221 and generates notification information according to the analysis result. The notification video included in the generated notification information is supplied to the display unit 30 and displayed. Note that the display control device 100 may be configured to include a synthesis unit 13, which synthesizes the notification video and the content video. Furthermore, audio data included in the generated notification information is output from the speaker 40.
[0046] As shown in FIG. 4, the notification information generating unit 222 includes a prompt generating unit 201, an utterance content acquiring unit 202, a voice generating unit 203, and a character generating unit 204, for example.
[0047] (Prompt generation unit 201) The prompt generation unit 201 generates input data (prompt) to be input to the trained model LM by referring to the analysis result by the analysis unit 221, and inputs the generated input data to the trained model LM. Here, the trained model LM may be configured, for example, to be provided in a server device 200 connected to the display control device 100 via a network N, as shown in FIG. 4, or may be configured to be provided in the display control device 100. Furthermore, the trained model LM may be, for example, a large-scale language model, but is not limited to this. Any machine-learned generative model can be used as the trained model LM.
[0048] The specific prompt generated by the prompt generation unit 201 is not limited to this embodiment, but as an example, the prompt may include reference information including the content of the dialogue between the user and the character up to the present time, the content of the current content, and the current state of the user, and instruction information to generate speech content for the user based on the above reference information.
[0049] (Utterance content acquisition unit 202) The utterance content acquisition unit 202 acquires the utterance content generated by the trained model LM based on the prompt generated by the prompt generation unit 201. As an example, the utterance content acquisition unit 202 acquires the utterance content in the form of text data, but this does not limit the present embodiment.
[0050] The processing by the prompt generation unit 201 and the speech content acquisition unit 202 can also be described as a process of generating input data (prompt) according to the results of the analysis processing by the analysis unit 221, and inputting the generated input data (prompt) into a trained model, thereby determining the speech content from the character to the user.
[0051] (Speech generation unit 203) The voice generation unit 203 generates voice data indicating the content of the utterance acquired by the utterance content acquisition unit 202. Here, the voice generation unit 203 may determine the tone of voice and accent of the utterance of the voice data by referring to the analysis result by the analysis unit 221. Alternatively, the prompt generated by the prompt generation unit 201 may include instruction information for specifying the tone of voice, and the voice generation unit 203 may determine the tone of voice and accent of the utterance of the voice data by referring to the response of the trained model LM to the instruction. The voice data generated by the voice generation unit 203 is supplied to the speaker 40.
[0052] (Character generation unit 204) The character generation unit 204 generates notification information by referring to the analysis result by the analysis unit 221. As an example, the character generation unit 204 generates (determines) video data of a character (character video data) to be included in the notification information by referring to the analysis result by the analysis unit 221. As an example, the character generation unit 204 determines the character's appearance, costume, and movements, etc., by referring to the analysis result by the analysis unit 221, and supplies a character video representing the determined content to the display unit 30. As described in FIG. 1, the display control device 100 may be configured to include the synthesis unit 13, and the character generation unit 204 may be configured to supply the character video to the synthesis unit 13. Furthermore, the character generation unit 204 may be configured to generate, together with the character video data, image data of a speech bubble that displays the content of the character's speech in text. Furthermore, the character generation unit 204 may be configured to create at least a portion of the character's video data using a generation AI. Alternatively, the character generation unit 204 may generate character images using images obtained by applying CG processing to images acquired externally by the sensing unit 20, or images created by an image generation AI from text data related to speech.
[0053] (Processing flow by the display control device 100) FIG. 5 is a flow diagram showing part of the processing flow by the display control device 100.
[0054] (Step S21) 5, first, in step S21, the second acquiring unit 21 acquires sensing data from the sensing unit 20. Specific examples of sensing data have been described above, and therefore will not be described here.
[0055] (Step S221) Subsequently, in step S221, the analysis unit 221 analyzes the sensing data acquired in step S21. A specific example of the analysis process by the analysis unit 221 has been described above, and therefore a description thereof will be omitted here.
[0056] (Step S222) Subsequently, in step S222, the notification information generator 222 generates notification information by referring to the analysis result in step S221.
[0057] (Specific processing example 1) FIG. 6 is a flow chart showing a specific processing example 1 by the display control device 100. In FIG.
[0058] (Step S221A) First, in step S221A, the analysis unit 221 refers to the sensing data obtained by the sensing unit 20 and detects a user positioned in front of the display device 1.
[0059] (Step S221B) Next, in step S221B, the analysis unit 221 executes a process of identifying the number of users positioned in front of the display device 1. If there is one user, the process proceeds to step S221C, and if there are multiple users, the process proceeds to step S222C.
[0060] (Step S221C) If there is one user, in step S221C, the analysis unit 221 performs personal identification processing of the user. As an example, the analysis unit 221 performs personal identification processing of the user by analyzing a face image of the user indicated by the imaging data included in the sensing data.
[0061] (Step S221D) Subsequently, in step S221D, the analysis unit 221 determines whether the user identified in step S221C corresponds to any one of a plurality of pre-registered users. In other words, the analysis unit 221 determines whether the user identified in step S221C is a registered user. Then, if the user identified in step S221C is a registered user, the process proceeds to step S222B; otherwise, the process proceeds to step S222A.
[0062] (Step S222A) If it is determined in step S221D that the user is not a registered user, in step S222A, the notification information generation unit 222 generates utterance content that does not include personal information and includes the generated utterance content in the notification information. As an example, in this step, the prompt generation unit 201 generates a prompt that includes at least one of the content of the dialogue between the user and the character up to the current time, the content of the content at the current time, the user's current state, and information that the user is not a registered user, and supplies the generated prompt to the trained model LM. Then, the utterance content acquisition unit 202 acquires the utterance content generated by the trained model LM to which the prompt is input. The utterance content generated in this manner is utterance content that does not include the user's personal information.
[0063] The specific examples of speech content generated in this step do not limit this embodiment, but as an example, the user may be referred to as "you" instead of by name, and may include weather-related content such as "It's hot today, isn't it?"
[0064] (Step S222B) On the other hand, if it is determined in step S221D that the user is a registered user, in step S222B, the notification information generation unit 222 generates utterance content including the user's personal information and includes the generated utterance content in the notification information. More specifically, the notification information generation unit 222 references personal information associated with the user, including the user's name, the user's preferences, the user's status history, viewing history, etc., and generates utterance content including the personal information. As an example, in this step, the prompt generation unit 201 generates a prompt including at least one of the content of the dialogue between the user and the character up to the current time, the content of the current content, and the user's current status, as well as the user's personal information, and supplies the generated prompt to the trained model LM. Then, the utterance content acquisition unit 202 acquires the utterance content generated by the trained model LM to which the prompt is input. The utterance content generated in this manner becomes utterance content including the user's personal information.
[0065] Although specific examples of speech content generated in this step do not limit this embodiment, one example may be to call the user by name and include content that reflects the user's condition history, such as "Yamada-san, have you recovered from your cold?"
[0066] (Step S222C) On the other hand, if it is determined in step S221B that there are multiple users, in step S222C, the notification information generation unit 222 generates utterances for multiple users and includes the generated utterance content in the notification information. As an example, in this step, the prompt generation unit 201 generates a prompt including at least one of the content of the dialogue between the user and the character up to the current time, the content of the content at the current time, the user's current state, and information that there are multiple users, and supplies the generated prompt to the trained model LM. Then, the utterance content acquisition unit 202 acquires the utterance content generated by the trained model LM to which the prompt is input. The utterance content generated in this way becomes the utterance content for multiple users.
[0067] Although specific examples of speech content generated in this step do not limit this embodiment, as an example, the above multiple users may be referred to as "everyone" and may include speech content suitable for multiple users, such as "What kind of gathering are you having today?"
[0068] Also, although not shown in the drawings, in step S221B, as another example, there is a case where the user is not detected by the camera. In this case, for example, when the character tries to speak, a gesture of searching for the user (the person to speak to) is displayed. Alternatively, the character may mutter to itself, "I wonder if there is anyone here," or may display gestures such as leaving, closing its eyes, or crying. In this way, when the user is not detected by the camera or the like, a display or speech different from that when the user is detected may be executed.
[0069] In this example, the second control unit 22 performs at least one of the following processes by referring to the sensing data from the sensing unit 20: a process of identifying the number of users; and a process of determining whether at least one of the one or more users is a registered user; and, depending on the results of the processes performed, a process of determining the content of the speech to be made to the user by a character, which is a specified image.
[0070] (Specific processing example 2) FIG. 7 is a flow chart showing a specific processing example 2 by the display control device 100. In FIG.
[0071] (Step S221F) In step S221F, the analysis unit 221 refers to the sensing data obtained by the sensing unit 20 and detects the direction of the user who has spoken.
[0072] (Step S222) Next, in step S222, the notification information generator 222 changes the video data of the character so that the character's body and line of sight are directed toward the user who has made the utterance.
[0073] (Step S221E) In step S221E, the analysis unit 221 determines whether or not the user has spoken (whether or not the user has spoken to the character) by referring to the sensing data from the sensing unit 20. If the user has spoken, the process proceeds to step S221F; if not, the process proceeds to step S222.
[0074] Fig. 8 shows an example of display by the display unit 30 of the display device 1 that performs the processing according to this example. In the upper part of Fig. 8, the character CR utters "This movie is interesting" (UC in the upper part of Fig. 8), and after the user responds to the utterance with "That's right" (UU in the upper part of Fig. 8), the user moves.
[0075] In this case, the character video data of the character CR is changed so that the body and line of sight are directed toward the user after the movement, as shown in the lower part of Fig. 8. Then, the character CR utters "This scene is interesting" (UC in the lower part of Fig. 8) while facing the user after the movement.
[0076] The above processing example is an example of processing in which the second control unit 22 refers to the imaging data included in the sensing data to identify the relative position of the user, and changes at least one of the display mode and speech mode of the alarm information according to the identified relative position.
[0077] (Specific processing example 3) The display control device 100 may perform the processing described below in addition to the above-mentioned specific processing examples 1 and 2. Fig. 9 is a diagram for explaining this processing example. The processing example shown on the left side of Fig. 9 shows processing when the analysis unit 221 refers to imaging data included in the sensing data from the sensing unit 20 and determines that the user is located closer than a predetermined distance (or when it determines that the size of the user in the imaging data is larger than a predetermined size). In this case, the analysis unit 221 performs processing to detect a user gesture, etc., using the imaging data as is.
[0078] On the other hand, the processing example shown on the right side of FIG. 9 illustrates processing performed when the analysis unit 221, with reference to the imaging data included in the sensing data obtained by the sensing unit 20, determines that the user is located farther than a predetermined distance (or determines that the size of the user in the imaging data is smaller than a predetermined size). In this case, the analysis unit 221 identifies the position of the user (recognition target) in the imaging data and cuts out the user from the imaging data. In this example, the sensing data obtained by the sensing unit 20 may include both imaging data with relatively low resolution and imaging data with relatively high resolution. Alternatively, the resolution of the imaging data included in the sensing data obtained by the sensing unit 20 may be dynamically changeable. In such a case, for example, the analysis unit 221 may refer to the low-resolution imaging data to determine that the user is located farther than a predetermined distance, identify the position of the user (recognition target) in the low-resolution imaging data, and cut out the user from the high-resolution imaging data. The analysis unit 221 then refers to the partial image cut out as described above to perform processing to detect user gestures, etc.
[0079] FIG. 10 is a flow diagram showing the processing flow in this processing example.
[0080] (Step S221G) In step S221G, the analysis unit 221 refers to the sensing data obtained by the sensing unit 20 and detects a user positioned in front of the display device 1.
[0081] (Step S221H) Subsequently, in step S221H, the analysis unit 221 refers to the sensing data and estimates the distance from the display device 1 to the user and the relative position of the user.
[0082] (Step S221I) Next, in step S221I, the analysis unit 221 determines whether the user is within a distance where face detection or gesture detection is possible. If the user is within a distance where face detection or gesture detection is possible, the process proceeds to step S221L; if not, the process proceeds to step S221J.
[0083] (Step S221J) If the user is located farther away than the distance at which face detection or gesture detection is possible, in step S221J, the analysis unit 221 identifies the position of the recognition target (whole body, face, hands, etc.) from the low-resolution image included in the sensing data.
[0084] (Step S221K) Then, in step S221K, the analysis unit 221 cuts out a partial image including the recognition target (whole body, face, hand, etc.) from the high-resolution image included in the sensing data.
[0085] (Step S221L) Then, in step S221L, face detection and gesture detection are performed by referring to the partial image cut out in step S221K (if NO in step S221I) or the image acquired in step S221G (if YES in step S221I). Then, the notification information generation unit 222 generates notification information according to the detection result.
[0086] In this example, when the user's relative position does not satisfy a predetermined standard (for example, when the user is located farther than a predetermined distance), the second control unit 22 extracts one or more partial images including specific parts of the user from imaging data that includes the user in its field of view and has a higher resolution than the imaging data (high-resolution imaging data included in the sensing data), and performs a process of determining the display mode of the alarm information based on the image analysis results of the extracted partial images.
[0087] Generally, analysis processing that references high-resolution imaging data requires computational costs, but according to the above processing example, it is possible to dynamically determine whether to use high-resolution data or continue processing with low-resolution data depending on the distance to the user, so that analysis processing can be performed efficiently while suppressing increases in computational costs.
[0088] (Other specific processing examples) The specific processing by the display control device 100 is not limited to the above example. It can also be expressed as the notification information generation unit 222 monitoring the user's state by referring to the sensing data from the sensing unit 20 and changing the timing and content of speaking to the user depending on the user's state. As an example, the following processing may be performed. The notification information generation unit 222 identifies the user's line of sight, and if the user is looking at the display device 1, speaks to the user, and if not, does not speak to the user. The notification information generation unit 222 identifies the user's facial expression, and, by referring to the facial expression, determines whether the user is serious, and if the user is serious, does not speak to the user. The notification information generation unit 222 identifies the user's facial expression, and, by referring to the facial expression, determines whether the user is smiling, and if the user is smiling, speaks to the user. The notification information generation unit 222 estimates the user's age, and if it is determined that the user is an elderly person above a predetermined age, speaks slowly. The notification information generating unit 222 may execute a process of estimating the age of the user, selecting a topic according to the estimated age, and making an utterance including the selected topic.
[0089] (Effects of Display Device 1) As described above, the display device 1 according to this embodiment is configured to include a first acquisition unit 11 that acquires content data, a first control unit 12 that controls the display of content indicated by the content data, a second acquisition unit 21 that acquires sensing data of one or more users, and a second control unit 22 that determines the notification mode of notification information to be notified to the user by referring to the sensing data. According to the display device 1 configured as described above, the notification mode of the notification information is determined by referring to the sensing data, so that suitable information can be presented to the user.
[0090] [Embodiment 2] Another embodiment of the present invention will now be described. Fig. 11 is a block diagram showing a specific configuration example of a display control device 100A according to this embodiment. The display control device 100A shown in Fig. 11 has the same configuration as the display control device 100 shown in Fig. 4. Furthermore, the display control device 100A has a generation unit 200A, and the generation unit 200A has a language model LM. The other configurations are the same as those of the display control device 100 shown in Fig. 4, so duplicated explanations will be omitted.
[0091] [Software implementation example] The functions of the display device 1 (hereinafter referred to as the "device") can be realized by a program that causes a computer to function as the device, and a program that causes a computer to function as each control block of the device (particularly each part included in the first control unit 12 and the second control unit 22).
[0092] In this case, the device includes a computer having at least one control device (e.g., a processor) and at least one storage device (e.g., a memory) as hardware for executing the program. The control device and storage device execute the program, thereby realizing the functions described in each of the above embodiments.
[0093] The program may be non-transitory and may be recorded on one or more computer-readable recording media. The recording media may or may not be included in the device. In the latter case, the program may be supplied to the device via any wired or wireless transmission medium.
[0094] Furthermore, some or all of the functions of the control blocks can be realized by logic circuits. For example, an integrated circuit in which a logic circuit that functions as each of the control blocks is formed is also included in the scope of the present invention. In addition, the functions of the control blocks can also be realized by, for example, a quantum computer.
[0095] Furthermore, each process described in each of the above embodiments may be executed by AI (Artificial Intelligence). In this case, the AI may run on the control device or on another device (for example, an edge computer or a cloud server).
[0096] 〔summary〕 The present specification describes at least the following aspects.
[0097] (Aspect 1) a first acquisition unit that acquires content data related to the content; a second acquisition unit that acquires user information related to the user; a control unit that changes a notification mode of notification information to be notified to the user in accordance with the user information; An information processing device comprising:
[0098] According to the above configuration, the notification mode of the notification information notified to the user is changed in accordance with the user information, so that suitable information can be presented to the user.
[0099] (Aspect 2) the notification information includes at least one of a predetermined image and a speech via the image, The notification mode includes at least one of a display mode of the image corresponding to the user information and an utterance mode of the utterance corresponding to the user information. 2. The information processing device according to claim 1.
[0100] (Aspect 3) the user information includes at least one of imaging data and audio data; determining a speech style suitable for the personal information identified based on the user information; 3. The information processing device according to aspect 2.
[0101] (Aspect 4) The control unit Identifying the relative position of the user by referring to the imaging data; At least one of the display mode and the speech mode of the notification information is changed according to the identified relative position. The information processing device according to aspect 3.
[0102] (Aspect 5) The control unit If the identified relative position does not satisfy a predetermined criterion, extracting one or more partial images including a specific part of the user from imaging data that includes the user in its angle of view and has a higher resolution than the imaging data; A display mode of the notification information is determined according to the result of image analysis of the extracted partial image. 5. The information processing device according to claim 4.
[0103] (Aspect 6) The speech style includes speech content, which is the content of the speech, The control unit The content of the speech to the user is determined based on the user information. 6. The information processing device according to any one of aspects 2 to 5.
[0104] (Aspect 7) The control unit refers to the sensing data of the user included in the user information. determining the number of one or more users; and A process for determining whether at least one of the one or more users is a registered user. Execute at least one of the following processes, The content of the speech to be given to the user via the image is determined according to the result of the executed processing. An information processing device according to aspect 6.
[0105] (Aspect 8) The control unit generating input data according to the results of the executed processing; The generated input data is input into a trained model to determine the content of the utterance to be given to the user. An information processing device according to aspect 7.
[0106] (Aspect 9) an information processing device according to any one of aspects 1 to 8; a display unit that displays at least one of the content and the notification information; Equipped with The information processing device further includes a content display control unit that controls display of the content.
[0107] (Aspect 10) A television receiver equipped with the display device according to embodiment 9.
[0108] (Aspect 11) a receiving device for receiving content data relating to the content; a display control device that controls display of the content indicated by the content data; an input device that accepts input of user information related to a user; a notification information control device that changes a notification mode of notification information to be notified to the user in accordance with the user information; an output device that outputs at least one of the content and the notification information; A display control system comprising:
[0109] (Aspect 12) A program for causing a computer to function as the information processing device according to aspect 1, the program causing a computer to function as the first acquisition unit, the second acquisition unit, and the control unit.
[0110] The display control device according to each aspect of the present invention may be realized by a computer. In this case, the program of the display control device that realizes the display control device on a computer by causing the computer to operate as each part (software element) of the display control device, and the computer-readable recording medium on which the program is recorded, also fall within the scope of the present invention.
[0111] The present invention is not limited to the above-described embodiments, and various modifications are possible within the scope of the claims. Embodiments obtained by appropriately combining the technical means disclosed in different embodiments are also included in the technical scope of the present invention. Furthermore, new technical features can be formed by combining the technical means disclosed in each embodiment.
[0112] For example, in the above-described embodiments, the user may be able to select whether or not to display text data. Similarly, the user may be able to select whether or not to display the avatar's speech. Furthermore, the display control device 100 may be implemented as a standalone device, such as an electronic device, including a set-top box as an example. In this case, the electronic device may have at least some of the functions shown in FIG. 1, FIG. 4, or FIG. 11, with the remaining functions being provided elsewhere. Alternatively, the electronic device may have all of these functions. In particular, the illustrated server device 200 may have multiple trained models LM. In this case, all of the multiple trained models LM may be provided in the electronic device, or at least some may be provided in the electronic device, with other trained models LM provided elsewhere. Note that the "avatar" in the above-described embodiments refers, for example, to an object represented based on a character, which can be primarily selected or created by the user. The avatar does not have to be a user's alter ego as generally defined; it may simply correspond to a predetermined image. [Explanation of symbols]
[0113] 1...Display device 100 Display control device 11 First acquisition section 12 First control section 21 Second acquisition section 22 Second control section 13. Synthesis section 10. Receiving unit 20 Sensing unit 30...Display section 40 Speaker
Claims
1. a first acquisition unit that acquires content data related to the content; a second acquisition unit that acquires user information related to the user; a control unit that changes a notification mode of notification information to be notified to the user in accordance with the user information; An information processing device comprising:
2. the notification information includes at least one of a predetermined image and a speech via the image, The notification mode includes at least one of a display mode of the image corresponding to the user information and an utterance mode of the utterance corresponding to the user information. The information processing device according to claim 1 .
3. the user information includes at least one of imaging data and audio data; determining a speech style suitable for the personal information identified based on the user information; The information processing device according to claim 2 .
4. The control unit obtaining a relative position of the user; At least one of the display mode and the speech mode of the notification information is changed according to the identified relative position. The information processing device according to claim 3 .
5. The control unit If the acquired relative position does not satisfy a predetermined criterion, extracting a partial image including a specific part of the user from imaging data that includes the user in an angle of view and has a higher resolution than the imaging data; A display mode of the notification information is determined according to the result of image analysis of the extracted partial image. The information processing device according to claim 4 .
6. The speech style includes speech content, which is the content of the speech, The control unit The content of the speech to the user is determined based on the user information. The information processing device according to claim 2 .
7. The control unit refers to the sensing data of the user included in the user information. A process of identifying the number of users; and A process for determining whether the user is a registered user. Execute at least one of the following processes, The content of the speech to be given to the user via the image is determined according to the result of the executed processing. The information processing device according to claim 6 .
8. The control unit generating input data according to the results of the executed processing; The generated input data is input into a trained model to determine the content of the utterance to be given to the user. The information processing device according to claim 7 .
9. An information processing device according to any one of claims 1 to 5; a display unit that displays at least one of the content and the notification information; Equipped with The information processing device further includes a content display control unit that controls display of the content.
10. A television receiver comprising the display device according to claim 9.
11. a receiving device for receiving content data relating to the content; a display control device that controls display of the content indicated by the content data; an input device that accepts input of user information related to a user; a notification information control device that changes a notification mode of notification information to be notified to the user in accordance with the user information; an output device that outputs at least one of the content and the notification information; An information processing system comprising:
Citation Information
Patent Citations
Media preferences
JP2011504710A