Information processing apparatus, display apparatus, television receiver, information processing system, and information processing method

The information processing device addresses the lack of suitable information provision in TVs by using user sensing data to tailor content interactions, enhancing user satisfaction through personalized and context-aware display and speech.

JP2026031090AActive Publication Date: 2026-02-24SHARP KK
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
JP2024134407
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-08-09
Publication Date
2026-02-24
Estimated Expiration
2044-08-09

AI Technical Summary

Technical Problem

Existing television technologies fail to provide suitable information to users alongside content, leading to suboptimal user satisfaction.

Method used

An information processing device that acquires content data and controls speech via objects displayed with the content, using user sensing data to determine appropriate notification modes, including display and speech styles tailored to individual preferences and contexts.

Benefits of technology

Enhances user experience by presenting relevant information alongside content, improving satisfaction through personalized and context-aware interactions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026031090000001_ABST
    Figure 2026031090000001_ABST
Patent Text Reader

Abstract

To provide a technology capable of presenting suitable information to a user together with contents.SOLUTION: A display control device (100) includes a first acquisition unit (11) for acquiring content data related to content, and a control unit (22) for controlling utterance through an object displayed together with the content on the basis of the content or the content data.SELECTED DRAWING: Figure 4
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an information processing device, a display device, a television receiver, an information processing system, and an information processing method. [Background technology]

[0002] Conventionally, television devices have been proposed that automatically provide personal media preferences based on user identification information detected by a camera. For example, in the technology described in Patent Document 1, a television device identifies a user based on pre-registered user information and a facial image detected by a camera, and displays media content related to the identified user or automatically changes settings. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Special Publication 2011-504710 Summary of the Invention [Problem to be solved by the invention]

[0004] In order to improve user satisfaction when watching television, it is important to consider what information to provide to users along with the content, but the technology described in Patent Document 1 had issues in this regard.

[0005] An object of one aspect of the present invention is to realize a technology that can present suitable information to a user along with content. [Means for solving the problem]

[0006] In order to solve the above problem, an information processing device according to one embodiment of the present invention includes a first acquisition unit that acquires content data related to content, and a control unit that controls speech via an object displayed together with the content based on the content or the content data.

[0007] In addition, a display device according to another aspect of the present invention comprises the above-mentioned information processing device and a display unit that displays both the content and the object, and the information processing device further comprises a content display control unit that controls the display of the content.

[0008] A television receiver according to another aspect of the present invention includes the above-described display device.

[0009] In addition, an information processing system according to another aspect of the present invention includes a first acquisition unit that acquires content data related to content, and a control unit that controls speech via an object displayed together with the content based on the content or the content data.

[0010] In addition, an information processing method according to another aspect of the present invention includes a first acquisition step of acquiring content data related to content, and a control step of controlling speech via an object displayed together with the content based on the content or the content data. [Effects of the Invention]

[0011] According to one aspect of the present invention, it is possible to present suitable information to a user together with content. [Brief explanation of the drawings]

[0012] [Figure 1] 1 is a block diagram showing a configuration of a display device according to a first embodiment of the present invention. [Figure 2] FIG. 1 is a diagram illustrating a display device according to a first embodiment of the present invention. [Figure 3]1 is a diagram showing an example of a display on the display device according to the first embodiment of the present invention. [Figure 4] 1 is a block diagram showing a configuration of a display control device according to a first embodiment of the present invention. [Figure 5] 3 is a flowchart showing a processing flow by the display control device according to the first embodiment of the present invention. [Figure 6] 3 is a flowchart showing a processing flow by the display control device according to the first embodiment of the present invention. [Figure 7] 1 is a diagram showing an example of a display on the display device according to the first embodiment of the present invention. [Figure 8] 1 is a diagram showing an example of a display on the display device according to the first embodiment of the present invention. [Figure 9] 1 is a diagram showing an example of a display on the display device according to the first embodiment of the present invention. [Figure 10] 1 is a diagram showing an example of a display on the display device according to the first embodiment of the present invention. [Figure 11] 3 is a flowchart showing a processing flow by the display control device according to the first embodiment of the present invention. [Figure 12] 3 is a flowchart showing a processing flow by the display control device according to the first embodiment of the present invention. [Figure 13] 1 is a diagram showing an example of a display on the display device according to the first embodiment of the present invention. [Figure 14] FIG. 10 is a block diagram showing another example of the configuration of a display control device according to the second embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0013] [Embodiment 1] Hereinafter, one embodiment of the present invention will be described in detail.

[0014] <Display device 1> FIG. 1 is a block diagram showing the configuration of a display device 1 according to this embodiment. As shown in FIG. 1, the display device 1 includes a receiving unit 10, a sensing unit 20, a display control device 100, a display unit 30, and a speaker 40. The components shown in FIG. 1 are merely examples of the configuration of the display device 1, and the present invention is not limited to these components. For example, the display device 1 may include various components, such as a remote control device operated by a user, an operation signal receiving unit that receives an operation signal from the remote control device, or a storage unit that stores video data. Furthermore, the components included in the display device 1 may be distributed across multiple devices via a network, for example. In such a case, the display device 1 may be referred to as a display system, and the display control device 100 may be referred to as a display control system. While the display device 1 includes the display control device 100 in the example shown in FIG. 1, the present embodiment is not limited to this configuration. The display control device 100 may also be implemented as a set-top box connected to the display device 1.

[0015] (Receiver 10) The receiving unit 10 receives content data and supplies the received content data to the display control device 100. The content data received by the receiving unit 10 includes, for example, encoded video data, encoded audio data, and related information accompanying the video data, but this does not limit the present embodiment. Furthermore, examples of the encoded video data include data (e.g., TS (Transport Steam)) encoded by various video encoding technologies such as MPEG2, MPEG4, H.264, and H.265, but this does not limit the present embodiment.

[0016] Furthermore, the related information may include, for example, at least one of program information related to the content, program guide data including the program information, information on data broadcasting provided in association with the content, and explanatory information related to the content, but these examples do not limit the present embodiment. Furthermore, for example, the related information may be information that can be extracted from the content data without performing a decoding process using a video encoding technology such as MPEG2, as described above. However, this example does not limit the present embodiment.

[0017] The receiving unit 10 may also be configured to acquire the content data from the Internet via wireless or wired communication. The receiving unit 10 may also be configured to acquire the content data from broadcast waves and include a tuner for selecting one of multiple channels included in the broadcast waves. When the receiving unit 10 includes a tuner, the display device 1 is also referred to as a television receiver (or simply a television).

[0018] (Sensing unit 20) The sensing unit 20 senses one or more users who use the display device 1. As an example, the sensing unit 20 includes one or more cameras and captures images of the one or more users using the cameras. The sensing unit 20 then supplies sensing data including image data captured by the cameras to the display control device 100. The sensing unit 20 may also include a laser scanner (LiDAR device) that detects users by reflecting laser light, and may be configured to include scan data acquired by the laser scanner in the sensing data and supply the data to the display control device 100. The sensing unit 20 may also include one or more microphones that collect speech from the one or more users, and may be configured to include audio data indicating the audio collected by the microphones in the sensing data and supply the data to the display control device 100.

[0019] (Display section 30) The display unit 30 displays a display image generated by the display control device 100. The display image includes at least one of a still image and a moving image (video). As will be described later, the display image includes, for example, at least one of a video of the content indicated by the content data (also referred to as a content video) and notification information for the user. The display unit 30 may also be configured to display at least one of a video of the content indicated by the content data (content video) and notification information for the user.

[0020] As an example, the display unit 30 may be configured to include a display panel and a driver that drives the display panel based on image data of the display image, but this does not limit the present embodiment. Furthermore, a liquid crystal panel or an organic EL panel may be used as the display panel, but this does not limit the present embodiment. Display examples by the display unit 30 will be described later.

[0021] (Speaker 40) The speaker 40 outputs the audio of the content indicated by the content data to the user. If the notification information includes audio (also called notification audio), the speaker 40 outputs the notification audio to the user. Note that in this embodiment, the term "audio" may include a human voice, but is not limited to this, and refers to any sound propagating through a medium such as air.

[0022] (Example of using display device 1) FIG. 2 is a diagram showing an example of how the display device 1 is used. In the example shown in FIG. 2, the display device 1 is realized as a stationary display device with legs. As shown in FIG. 2, the display device 1 is configured to include a sensing unit 20 that senses a user U, a display unit 30 that displays an image for display, and speakers 40 (two in the example of FIG. 2) that output audio of content and audio of notification information. The display device 1 may also be realized as a wall-mounted display device.

[0023] <Display control device 100> Returning to FIG. 1, the configuration of each unit of the display control device 100 included in the display device 1 will be described. As shown in FIG. 1, the display control device 100 includes a first acquisition unit 11, a first control unit 12, a synthesis unit 13, a second acquisition unit 21, and a second control unit 22. Note that the designations "first," "second," and the like do not limit this embodiment. For example, either or both of the "first acquisition unit 11" and the "second acquisition unit 21" may be simply referred to as an "acquisition unit," and either or both of the "first control unit 12" and the "second control unit 22" may be simply referred to as a "control unit." The display control device 100 may also be referred to as an information processing device. The first control unit 12 may also be referred to as a content display control unit.

[0024] (First acquisition unit 11) The first acquisition unit 11 acquires content data related to content. As an example, the first acquisition unit 11 acquires content data received by the above-mentioned receiving unit 10. The first acquisition unit 11 supplies the acquired content data to the first control unit 12.

[0025] (First control unit 12) The first control unit 12 controls the display of the content indicated by the content data supplied from the first acquisition unit 11 (in other words, the display of the content video). As an example, the first control unit 12 decodes video data included in the content data, and supplies the content video obtained by the decoding process to the synthesis unit 13. The content video forms a part of the display image displayed by the display unit 30. The first control unit 12 also decodes audio data included in the content data, and supplies the content audio obtained by the decoding process to the synthesis unit 13. The content audio forms a part of the audio output by the speaker 40. The first control unit 12 also extracts the above-mentioned related information from the content data, and supplies the extracted related information to the second control unit 22, for example.

[0026] (Second acquisition unit 21) The second acquisition unit 21 acquires sensing data of one or more users who use the display device 1 from the sensing unit 20. Here, as described above, the sensing data includes at least one of imaging data that includes the users in the angle of view and audio data that includes speech by the users. The second acquisition unit 21 supplies the sensing data to the second control unit 22.

[0027] The sensing data is an example of user information related to the user of the display device 1. Therefore, the second acquisition unit 21 may be expressed as acquiring user information related to the user of the display device 1. In addition, in this embodiment, the user information may include imaging data and audio data acquired by the sensing unit 20 or another device.

[0028] The sensing data is an example of user identification information for identifying the user of the display device 1. Therefore, the second acquisition unit 21 may be expressed as acquiring user identification information for identifying the user of the display device 1.

[0029] The second acquiring unit 21 may be configured to acquire preference information corresponding to the user identification information. As an example, the second acquiring unit 21 may be configured to function as a preference information acquiring unit that acquires preference information corresponding to the user identification information from a preference information storage unit (preference information DB) that stores preference information of a plurality of users.

[0030] (Second control unit 22) The second control unit 22 executes an analysis process of analyzing at least one of the content data, the content video, the content audio, the related information, and the sensing data, and performs various processes by referring to the analysis results. As an example, the second control unit 22 executes at least one of a process of identifying the content by analyzing the content video or the related information, and a process of identifying utterances by the one or more users or a state of the one or more users by analyzing the sensing data, and performs various processes by referring to the results of these processes.

[0031] As an example, the second control unit 22 determines (generates) the notification mode of the notification information to be notified to the user by referring to the sensing data. In other words, the second control unit 22 changes the notification mode of the notification information to be notified to the user in accordance with the user information. Here, the notification information includes at least one of a predetermined image and speech via the image. For example, the notification information includes an object whose display mode and / or speech mode changes in accordance with at least one of the content data and the sensing data. Here, the object is an example of a predetermined image. Furthermore, as an example, the object may be a character generated by the character generation unit 204 described below, a CG-processed image obtained from an external source, or an image created by an image generation AI from text data related to speech. Therefore, it may be expressed that the notification mode of the notification information includes at least one of a display mode of a predetermined image in accordance with the user information and a speech mode of speech in accordance with the user information.

[0032] Further, as one example, the second control unit 22 may perform a process of determining a speech style suitable for personal information identified based on the user information. For example, the second control unit 22 may perform a process of determining a speech style (such as a speech style for elderly men / a speech style for children) according to the age group or gender, etc., of the individual identified based on the user information (such as an elderly man / a child).

[0033] Further, for example, the second control unit 22 may refer to the imaging data to identify the relative position of the user, and change at least one of the display mode and speech mode of the alarm information according to the identified relative position. Here, the relative position indicates the position of the user relative to the display device 1, and includes, for example, the direction from the display device 1 and the distance from the display device 1.

[0034] The speech style may include speech content, which is the content of the speech, and the second control unit 22 may determine the speech content to the user by referring to at least one of content and user information. The second control unit 22 may also be configured to control the timing of speech to the user based on at least one of content data and user information.

[0035] As an example, the second control unit 22 may refer to the sensing data to perform at least one of the following processes: a process of identifying the number of one or more users; and a process of determining whether at least one of the one or more users is a registered user; and, depending on the results of the executed processes, may perform a process of determining the content of a speech to the user (for example, from a character) via a specified image.

[0036] Furthermore, second control unit 22 may control the speech of the character by determining whether to speak, the timing of the speech, and / or the content of the speech. Second control unit 22 may also be configured to control the display mode of the character in response to the control of the speech from the character.

[0037] The second control unit 22 may also be configured to control speech via an object (for example, speech from a character) based on content or content data. Here, "control based on content or content data" includes, for example, control based on at least one of content video, content audio, and related information obtained from the content data. Also, for example, the second control unit 22 may be configured to control speech in accordance with analysis results related to the content obtained by analyzing the content or content data. Here, "analyzing content" refers to analyzing content video and content audio, but is not limited to this and may also include analyzing related information. The analysis results may be analyzed by the display control device 100 or the display device 1, or may be obtained from the cloud.

[0038] The second control unit 22 may be configured to determine the notification mode of the notification information to be notified to the user based on preference information. In other words, the second control unit 22 may be configured to determine the notification mode of the notification information in accordance with the preference information of the user identified with reference to the sensing data. As an example, the second control unit 22 may be configured to determine at least one of the display mode and speech mode of the character included in the notification information with reference to the preference information.

[0039] Furthermore, the second control unit 22 may be configured to perform at least one of the following processes: determining the display mode of an image based on preference information; and generating speech from a plurality of voice tone candidates based on preference information. As an example, the second control unit 22 may be configured to perform at least one of the following processes: determining a character to be displayed together with content from a plurality of character candidates by referring to preference information; and selecting a voice tone for the character's speech from a plurality of voice tone candidates by referring to preference information.

[0040] (Synthesis section 13) The synthesis unit 13 synthesizes the notification information determined (generated) by the second control unit 22 with the content video supplied from the first control unit 12. More specifically, the synthesis unit 13 synthesizes the notification video included in the notification information determined (generated) by the second control unit 22 with the content video supplied from the first control unit 12 to generate a display image including the content video and the notification video, and supplies the generated display image to the display unit 30.

[0041] In addition, the synthesis unit 13 generates output audio including the notification audio and the content audio by integrating the notification audio included in the above notification information with the content audio supplied from the first control unit 12, and supplies the generated output audio to the speaker 40.

[0042] Fig. 3 shows a display example of a display image generated by the synthesis unit 13. As shown in Fig. 3, the display image generated by the synthesis unit 13 and displayed by the display unit 30 includes a first region R1 for displaying a content video and a second region R2 for displaying notification information. The second region R2 also includes a character CR included in the notification information, text data UC indicating the content of the utterance by the character CR, and text data UU indicating the utterance by the user (corresponding to "you" in Fig. 3). At least one of the display manner and the utterance manner of the character changes depending on at least one of the content data and the sensing data.

[0043] Although specific examples of the character CR do not limit the present embodiment, as an example, the character may be any of a character simulating a user, a character simulating a living body such as an animal, a character simulating a non-living body such as a robot, or a character other than the above. Note that the character CR may also be referred to as an avatar.

[0044] In addition, a predetermined website screen may be displayed in the first area R1 together with or instead of the content video. The predetermined website screen here may be, for example, an e-commerce (EC) site where electronic commerce is possible. The EC site may be a website related to the provision of services such as travel arrangements as well as products.

[0045] (Specific Configuration Example of Display Control Device 100) Next, a specific configuration example of the display control device 100 according to this embodiment will be described with reference to Fig. 4. Fig. 4 is a block diagram showing a specific configuration example of the display control device 100. Note that in the following description, overlapping descriptions of matters that have already been described about the display control device 100 may be omitted.

[0046] (First control unit 12) As shown in FIG. 4, the first control unit 12 includes, for example, a content playback unit 121. The content playback unit 121 decodes video data included in the content data and supplies the decoded content video to the synthesis unit 13 and the second control unit 22. The content playback unit 121 generates decoded content video using a decoding process compliant with various video encoding technologies such as MPEG2, MPEG4, H.264, and H.265. The content playback unit 121 also extracts related information included in the content data from the content data and supplies the extracted related information to the synthesis unit 13 and the second control unit 22. The content playback unit 121 also decodes audio data included in the content data and supplies the decoded content audio to the speaker 40. Note that, as described in FIG. 1, the display control device 100 may be configured to include the synthesis unit 13, and the first control unit 12 may be configured to supply the content video and the content audio to the synthesis unit 13.

[0047] (Second control unit 22) 4, the second control unit 22 includes an analysis unit 221 and a notification information generation unit 222. The analysis unit 221 analyzes the sensing data supplied from the second acquisition unit 21, the decoded content video supplied from the content playback unit 121, and the related information supplied from the content playback unit 121. The analysis process by the analysis unit 221 may include, for example, a process of identifying the relative positions of one or more users by analyzing imaging data included in the sensing data, a process of identifying the states (postures, facial expressions, emotions, etc.) of one or more users by analyzing imaging data included in the sensing data, a process of identifying the speech content of one or more users by analyzing audio data included in the sensing data, a process of identifying the relative positions of one or more users by analyzing audio data included in the sensing data, a process of identifying the content of the content at each point in time (scenes, characters, actions of characters, speech of characters, etc.) by analyzing the content video, and a process of identifying the title, characters, plot, etc. of the content shown in the content video by analyzing the related information; however, these examples do not limit the present embodiment.

[0048] Note that, for example, the various analytical processes performed by the analysis unit 221 can use a machine-learned inference model (prediction model) that executes various algorithms such as an object detection algorithm or an utterance extraction algorithm, but this does not limit the present embodiment. Furthermore, for example, the inference model may be configured as part of a trained model LM (described later), or may be realized as a model separate from the trained model LM (described later). Furthermore, the inference model may be configured to be included in the server device 200 (described later), or may be configured to be included in the display control device 100. The results of the above analytical processes performed by the analysis unit 221 are supplied to the alarm information generation unit 222.

[0049] The notification information generation unit 222 refers to the analysis result by the analysis unit 221 and generates notification information according to the analysis result. The notification video included in the generated notification information is supplied to the synthesis unit 13 and synthesized with the content video. The audio included in the generated notification information is output from the speaker 40. Note that the notification information generation unit 222 may be configured to control the volume of the audio of the content via the content playback unit 121 when generating notification information including audio and outputting it from the speaker 40.

[0050] As shown in FIG. 4, the notification information generation unit 222 includes, as an example, a prompt generation unit 201, an utterance content acquisition unit 202, a voice generation unit 203, a character generation unit 204, an utterance control unit 205, and a conversation history management unit 206.

[0051] (Prompt generation unit 201) The prompt generation unit 201 generates input data (prompt) to be input to the trained model LM with reference to the analysis result by the analysis unit 221, and inputs the generated input data to the trained model LM. Here, the trained model LM may be configured, for example, to be provided in a server device 200 connected to the display control device 100 via a network N, as shown in FIG. 4, or may be configured to be provided in the display control device 100. Furthermore, the trained model LM may be, for example, a large-scale language model, but is not limited to this. Any machine-learned generative model can be used as the trained model LM. The trained model LM may be a model combining multiple models, or may be a single multimodal model. Furthermore, the prompt generation unit 201 may be configured to save the generated prompt as a history.

[0052] The specific prompt generated by the prompt generation unit 201 is not limited to this embodiment, but as an example, the prompt can be configured to include reference information including the content of the dialogue between the user and the character up to the present time, the content of the current content, the current state of the user (sensing data), etc., and instruction information to generate speech content for the user based on the above reference information.

[0053] (Utterance content acquisition unit 202) The utterance content acquisition unit 202 acquires the utterance content generated by the trained model LM based on the prompt generated by the prompt generation unit 201. As an example, the utterance content acquisition unit 202 acquires the utterance content in the form of text data, but this does not limit the present embodiment. Note that the utterance content acquisition unit 202 may be configured to link the generated utterance content with utterance destination information indicating the destination of the utterance content. Linking of the utterance destination information may be performed when the trained model LM generates the utterance content. The utterance destination includes, for example, a user watching the content, a character (oneself), or a second character displayed together with the character.

[0054] Furthermore, the speech content acquisition unit 202 may be configured to acquire speech content candidates with reference to at least one of content data and user information. Specifically, the speech content acquisition unit 202 may be configured to acquire speech content candidates from a character with reference to at least one of content and sensing data. In this case, the speech content acquisition unit 202 does not acquire speech content candidates when, for example, the content of the content is serious and does not fit into a conversation with a character, or when the length of the speech content candidates (the time required to read them out loud) is too long for the user's state (unsettled, in a cheerful mood), etc.

[0055] The processing by the prompt generation unit 201 and the speech content acquisition unit 202 can also be described as processing that generates input data (prompt) according to the results of the analysis processing by the analysis unit 221, and determines the speech content from the character to the user by inputting the generated input data (prompt) into a trained model.

[0056] (Speech control unit 205) The speech control unit 205 changes the speech style of the character (object) depending on the speech destination of the character. The speech control unit 205 specifies the speech destination by referring to speech destination information linked to the speech content. Here, when the character or a second character is specified as the speech destination, the speech control unit 205 supplies the speech content to the voice generation unit 203.

[0057] When the user is specified as the speech destination, the speech control unit 205 refers to the sensing data to identify the timing of the user's speech. That is, the speech control unit 205 determines whether the user is about to speak. Then, the speech control unit 205 controls the character not to speak at the timing of the user's speech. That is, the speech control unit 205 does not supply the speech content acquired by the speech content acquisition unit 202 to the speech generation unit 203. On the other hand, when it is not the user's timing to speak (when the speech control unit 205 is waiting to be addressed by the user), the speech control unit 205 supplies the speech content acquired by the speech content acquisition unit 202 to the speech generation unit 203.

[0058] Furthermore, the speech control unit 205 may be configured to identify the state of the user by referring to the sensing data. Specifically, the speech control unit 205 may refer to the analysis result of the sensing data by the analysis unit 221 and determine whether the state of the user is in a predetermined state. Examples of the predetermined state include a state in which the user is concentrating on the content (viewing the content with a serious expression), a state in which the user is performing an action other than viewing the content (such as reading), etc.

[0059] When the speech control unit 205 determines that the user's state is in a predetermined state, it may be configured to, for example, prevent the character from speaking (not supply the speech content to the voice generation unit 203), or to control the character to speak to someone other than the user.

[0060] Furthermore, the speech control unit 205 may be configured to determine whether the content being displayed corresponds to a specific category or a specific scene. The content being displayed corresponds to a specific category or a specific scene, for example, if the content being displayed includes content that is not suitable for viewing with the character (that should be viewed seriously). The speech control unit 205 may be configured to control the character not to speak when it is determined that the content corresponds to a specific category or a specific scene.

[0061] The speech control unit 205 may also be configured to determine, as the speech content, a candidate speech content that is recorded as not having been uttered in the conversation history. That is, speech content that has been generated and acquired but not actually uttered by the character may be uttered later. In this case, the speech control unit 205 performs a process of sequentially determining whether or not it is time to utter the determined speech content. Then, when it is time to utter the utterance content, the speech control unit 205 causes the character to utter the utterance content.

[0062] The speech control unit 205 may be configured to change the content of the acquired candidate by referring to at least one of the content of the candidate, the content, and the sensing data. For example, the speech control unit 205 changes the content of the candidate by performing processing such as shortening the utterance content when the utterance content is too long for the user's situation or the content scene.

[0063] (Conversation history management unit 206) When the conversation history management unit 206 determines not to utter a candidate utterance content, it saves a history indicating that the candidate utterance was not uttered as part of the conversation history between the user and the character. Here, the conversation history management unit 206 accumulates the history in the conversation history database 23 provided in the display control device 100. Note that the conversation history management unit 206 may also be configured to accumulate at least one of the utterance content uttered by the character and the utterance content of the user in the conversation history database 23.

[0064] (Speech generation unit 203) The voice generation unit 203 generates voice data indicating the content of the utterance supplied from the speech control unit 205. Here, the voice generation unit 203 may determine the tone of voice and accent of the utterance of the voice data by referring to the analysis result by the analysis unit 221. Alternatively, the prompt generated by the prompt generation unit 201 may include instruction information for specifying the tone of voice, and the voice generation unit 203 may determine the tone of voice and accent of the utterance of the voice data by referring to the response of the trained model LM to the instruction. The voice generated by the voice generation unit 203 is supplied to the synthesis unit 13. The voice is then supplied from the synthesis unit 13 to the speaker 40 together with the voice of the content.

[0065] (Character generation unit 204) The character generation unit 204 generates the notification information by referring to the analysis result by the analysis unit 221. As an example, the character generation unit 204 generates (determines) video data of a character (character video data) to be included in the notification information by referring to the analysis result by the analysis unit 221. The character may be any character with an empathetic appearance, and is not limited to a human. As an example, the character generation unit 204 determines the character's appearance, including its physical appearance, its clothing, its movements, etc., by referring to the analysis result by the analysis unit 221, and supplies the character video expressing the determined content to the synthesis unit 13. Note that the character generation unit 204 may be configured to generate, along with the character video, an image of a speech bubble displaying the character's speech content in text. Alternatively, the character generation unit 204 may be configured to create at least a part of the character's video using a generation AI. At least a part of the character may be, for example, an image obtained by applying CG processing to an image acquired from an external device by the sensing unit 20, or an image created by an image generation AI from text data related to speech.

[0066] The character generation unit 204 may also be configured to generate a plurality of character images. In this case, the synthesis unit 13 receives a supply of a plurality of character images and displays the plurality of characters together with the content.

[0067] Furthermore, the character generation unit 204 may be configured to change the display mode of a character (object) depending on the speech destination of the character. The character generation unit 204 may identify the speech destination by referring to speech destination information linked to the speech content, or may refer to the speech destination identified by the speech control unit 205. Here, when a user is identified as the speech destination, the character generation unit 204 generates a character image facing in the direction of the user, for example, facing forward. At this time, the character generation unit 204 may be configured to generate, together with the character image, display data indicating that the character is speaking to the user (for example, text display data saying "Calling"). At this time, the character generation unit 204 may be configured to generate a speech bubble image with a display mode (for example, color, text size) different from that when the speech destination is the character or a second character.

[0068] Furthermore, when a second character is specified as the speech destination, the character generation unit 204 generates, for example, a character image facing in the direction in which the second character is displayed (for example, sideways). Furthermore, when the character itself is specified as the speech destination, the character generation unit 204 generates a character image facing in the direction in which the first area R1 (area displaying content) exists.

[0069] Furthermore, the character generation unit 204 may be configured to make the character perform a predetermined action when the speech control unit 205 determines that the character will not speak. The predetermined action may include, for example, a nod, a surprised reaction, or the like.

[0070] (Processing flow by the display control device 100) FIG. 5 is a flow diagram showing part of the processing flow by the display control device 100.

[0071] (Step S11) 5, first, in step S11, the first acquisition unit 11 acquires content data. Specific examples of content data have been described above, and therefore will not be described here.

[0072] (Step S221A) Subsequently, in step S221A, the analysis unit 221 analyzes the content data acquired in step S11. A specific example of the analysis process by the analysis unit 221 has been described above, and therefore a description thereof will be omitted here.

[0073] (Step S21) Subsequently, in step S21, the second acquisition unit 21 acquires sensing data from the sensing unit 20. Specific examples of sensing data have been described above, and therefore will not be described here.

[0074] (Step S221) Subsequently, in step S221, the analysis unit 221 analyzes the sensing data acquired in step S21. A specific example of the analysis process by the analysis unit 221 has been described above, and therefore a description thereof will be omitted here.

[0075] (Step S222) Subsequently, in step S222, the notification information generator 222 generates notification information by referring to the analysis result in step S221.

[0076] (Specific processing example 1) 6 is a flow diagram showing a specific processing example 1 by the display control device 100. The start conditions of this processing are not limited to the above, but the processing may be started, for example, when content is played, when a user operates the television or speaks to call a character, or when the sensing unit 20 detects a user.

[0077] (Step S221A) First, in step S221A, the analysis unit 221 analyzes the content video to identify the content at each time point.

[0078] (Step S222A) Next, in step S222A, the prompt generation unit 201 generates a prompt by referring to the analysis result by the analysis unit 221, and inputs the generated prompt to the trained model LM. Then, the trained model LM generates the content of the character's utterance.

[0079] (Step S222B) Next, in step S222B, the trained model LM or the utterance content acquisition unit 202 links the generated utterance content with utterance destination information.

[0080] (Step S222C) Next, in step S222C, the speech control unit 205 refers to the speech destination information linked to the speech content to identify the speech destination. If the second character is identified as the speech destination, the process proceeds to step S222D. If the content is identified as the speech destination, the process proceeds to step S222E. If the user is identified as the speech destination, the process proceeds to step S222F.

[0081] (Step S222D) If the second character is identified as the speech destination, in step S222D, the notification information generation unit 222 controls the character to speak to the second character. Specifically, the speech control unit 205 supplies the speech content linked to the speech destination information indicating the character as the speech destination to the voice generation unit 203, and the character generation unit 204 generates an image of the character speaking to the second character. As a result, as shown in FIG. 7, for example, the second character CR2, the character CR1 facing the direction of the second character CR2, and text data UC of the speech content to the second character CR2 (for example, "Avatar 2, this movie is interesting") are displayed in the second region R2 of the display unit 30. Then, a voice speaking to the second character is heard from the speaker 40.

[0082] (Step S222E) If the content is identified as the speech destination, in step S222E, the notification information generation unit 222 controls the character to utter a comment about the content. Specifically, the speech control unit 205 supplies the speech content linked to the speech destination information indicating the content as the speech destination to the voice generation unit 203, and the character generation unit 204 generates an image of the character talking to itself. As a result, as shown in FIG. 8, for example, a character facing the direction of the first area R1 and text data UC of a comment about the content (e.g., "This movie is interesting") are displayed in the second area R2 of the display unit 30. Then, the comment is heard from the speaker 40.

[0083] (Step S222F) If the user is specified as the speech destination, in step S222F, the speech control unit 205 refers to the sensing data to specify the timing of the user's speech. If it is not the user's timing to speak, the process proceeds to step S222G, and if it is the user's timing to speak, the process proceeds to step S222H.

[0084] (Step S222G) If it is not the user's timing to speak (if the user is waiting to be spoken to), in step S222G, the notification information generation unit 222 controls the character to speak to the user. Specifically, the speech control unit 205 supplies the speech content linked to speech destination information indicating the user as the speech destination to the voice generation unit 203, and the character generation unit 204 generates an image of the character speaking to the user. As a result, as shown in FIG. 9, the second area R2 of the display unit 30 displays, for example, a character CR facing the direction of the user, text data UC of the content spoken to the user (e.g., "This movie is interesting"), and display data CALL indicating that the character is speaking to the user. Then, a voice speaking to the user is heard from the speaker 40.

[0085] (Step S222H) If it is time for the user to speak, in step S222H, the notification information generation unit 222 controls the character not to speak. Specifically, the speech control unit 205 does not supply the speech content to the voice generation unit 203, and the character generation unit 204 generates an image of the character in a silent state. As a result, in the second area R2 of the display unit 30, as shown in Fig. 10, for example, a character CR facing the direction in which the user is present and text data UC indicating silence (for example, "...") are displayed. Then, the voice of the character is not heard from the speaker 40 (only the audio of the content is heard).

[0086] (Step S222J) Next, in step S222J, it is determined whether or not the content is being played back. If it is determined that the content is being played back, the process returns to step S221A, and if it is determined that the content is not being played back, the process ends.

[0087] (Specific processing example 2) 11 is a flow diagram showing a specific processing example 2 by the display control device 100. Processing example 2 has the same flow as processing example 1 from step S221A to step S222C.

[0088] (Step S222K) If the user is identified as the utterance destination in step S222C, in step S222K, the utterance control unit 205 identifies the user's state by referring to the sensing data (or the analysis result thereof) by the sensing unit 20. Then, if it is identified that the user's state is a predetermined state, the process proceeds to step S222E or step S222H, and if it is identified that the user's state is not a predetermined state, the process proceeds to step S222G.

[0089] (Specific processing example 3) 12 is a flow diagram showing a specific processing example 3 by the display control device 100. Processing example 3 has the same flow as processing example 1 from step S221A to step S222B.

[0090] (Step S222L) After step S222B, in step S222L, speech control unit 205 determines whether the content being displayed corresponds to a specific category or a specific scene. If it is determined that the content corresponds to a specific category or a specific scene, the process proceeds to step S222H, and if it is determined that the content does not correspond to a specific category or a specific scene, the process proceeds to step S222C.

[0091] (Modifications of specific processing examples 1 to 3) In addition, the second control unit 22 of the display control device 100 may be configured to execute a process, during the processes shown in the above specific processing examples 1 to 3, or separately from the above processes, to control the display mode of an object so that the object (character) faces the display area of ​​the content when the content is being displayed (including when the content is started or resumed).

[0092] Specifically, when content display is started or resumed, the character generation unit 204 generates an image of the character CR facing the direction in which the first region R1 exists, as shown in FIG. 13, for example. If the character CR was displayed before the content display was started or resumed, the character generation unit 204 generates an image of the character CR facing the direction in which the first region R1 exists, more so than the character CR that was displayed up until then. Note that, as the image of the character CR facing the direction in which the first region R1 exists, the character generation unit 204 may generate an image in which a portion of the character CR (for example, only the head and above, as shown in FIG. 13) faces that direction, or may generate an image in which the entire character CR faces that direction. As a result, the image displayed on the entire display unit 30 visually appears as if the character CR is viewing the content displayed in the first region R1 together with the user, as shown in FIG. 13. This enhances the sense of unity when viewing content together with the avatar.

[0093] Note that while the character CR faces the direction in which the first region R1 exists (while the content is being viewed), the second control unit 22 may be configured to store the details of the displayed content in a storage unit (not shown). In this way, for example, after the display of the content has finished, it becomes possible to generate speech content based on the stored content content and have the character CR speak (converse with the user about the content). The period during which the content is being viewed may also include a state in which the character CR is not facing the direction in which the first region R1 exists but is only listening to the audio.

[0094] Furthermore, when the display of the content is finished or interrupted, the character generation unit 204 generates an image of the character CR facing in a direction different from the direction in which the first region R1 exists (for example, toward the side where the user exists).

[0095] The character generation unit 204 may be configured to generate an image of the character CR facing the direction in which the first region R1 exists when content in which the user is interested is displayed or resumed. Whether the content is of interest to the user can be determined based on, for example, the content name or genre (registered by the user) pre-stored in a storage unit (not shown), information on the user's preferences, etc. Whether the content is of interest to the user can also be determined based on the results of an analysis by the analysis unit 221 based on sensing data (image data and audio data acquired by the sensing unit 20 or another device) (for example, whether the user has an expression showing interest in the content, whether the user has spoken to the effect that they are interested in the content, etc.).

[0096] (Effects of Display Device 1) As described above, the display device 1 according to this embodiment is configured to include a first acquisition unit that acquires content data related to the content, and a control unit that controls speech via an object displayed together with the content based on the content or the content data. According to the display device 1 configured as described above, the notification mode of the notification information to be displayed together with the content is determined by referring to the sensing data, so that suitable information can be presented to the user together with the content.

[0097] Furthermore, according to the display device 1 of this embodiment, the character speaks to the user at the appropriate timing, so the character will not accidentally speak to the user when there is no need for the character to speak (for example, when the user is not watching content) or when the user does not want to be disturbed by the character (for example, when the user is watching content intently), etc. As a result, the user and the character can communicate smoothly.

[0098] [Embodiment 2] Another embodiment of the present invention will now be described. Fig. 14 is a block diagram showing a specific configuration example of a display control device 100A according to this embodiment. The display control device 100A shown in Fig. 14 has the same configuration as the display control device 100 shown in Fig. 4. Furthermore, the display control device 100A has a generation unit 200A, and the generation unit 200A has a language model LM. Other configurations are the same as those of the display control device 100 shown in Fig. 4, so duplicated explanations will be omitted.

[0099] [Software implementation example] The functions of the display device 1 (hereinafter referred to as the "device") can be realized by a display control program that causes a computer to function as the device, and that causes a computer to function as each control block of the device (particularly each part included in the first control unit 12 and the second control unit 22).

[0100] In this case, the device includes a computer having at least one control device (e.g., a processor) and at least one storage device (e.g., a memory) as hardware for executing the display control program. The display control program is executed by the control device and the storage device, thereby realizing the functions described in each of the above embodiments.

[0101] The display control program may be non-transitory and may be recorded on one or more computer-readable recording media. The recording media may or may not be included in the device. In the latter case, the display control program may be supplied to the device via any wired or wireless transmission medium.

[0102] Furthermore, some or all of the functions of the control blocks can be realized by logic circuits. For example, an integrated circuit in which a logic circuit that functions as each of the control blocks is formed is also included in the scope of the present invention. In addition, the functions of the control blocks can also be realized by, for example, a quantum computer.

[0103] Furthermore, each process described in each of the above embodiments may be executed by AI (Artificial Intelligence). In this case, the AI ​​may run on the control device or on another device (for example, an edge computer or a cloud server).

[0104] 〔summary〕 The present specification describes at least the following aspects.

[0105] (Aspect 1) a first acquisition unit that acquires content data related to the content; a second acquisition unit that acquires user information related to the user; a control unit that changes a notification mode of notification information to be notified to the user in accordance with the user information; An information processing device comprising:

[0106] According to the above configuration, the notification mode of the notification information notified to the user is changed in accordance with the user information, so that suitable information can be presented to the user.

[0107] (Aspect 2) a first acquisition unit that acquires content data related to the content; a control unit that controls speech via an object displayed together with the content based on the content or the content data; An information processing device comprising:

[0108] (Aspect 3) The control unit The speech is controlled in accordance with an analysis result regarding the content obtained by analyzing the content or the content data. 3. The information processing device according to aspect 2.

[0109] (Aspect 4) The control unit When the content is being displayed (including when it is started or resumed), the display mode of the object is controlled so that the object faces the display area of ​​the content. The information processing device according to aspect 3.

[0110] (Aspect 5) The control unit Displaying a plurality of said objects together with said content; The display mode or speech mode of the object is changed depending on the speech destination of the object. The information processing device according to aspect 3.

[0111] (Aspect 6) The destination of the speech includes: A user who views the content; and A second object to be displayed with the object Contains one of the following: An information processing device according to aspect 5.

[0112] (Aspect 7) an information processing device according to any one of aspects 2 to 6; a display unit that displays both the content and the object; Equipped with The information processing device further includes a content display control unit that controls display of the content.

[0113] (Aspect 8) A television receiver comprising the display device according to embodiment 7.

[0114] (Aspect 9) a first acquisition unit that acquires content data related to the content; a control unit that controls speech via an object displayed together with the content based on the content or the content data; An information processing system comprising:

[0115] (Aspect 10) a first acquisition step of acquiring content data relating to the content; a control step of controlling utterances via an object displayed together with the content based on the content or the content data; An information processing method comprising:

[0116] The display control device according to each aspect of the present invention may be realized by a computer. In this case, the program of the display control device that realizes the display control device on a computer by causing the computer to operate as each part (software element) of the display control device, and the computer-readable recording medium on which the program is recorded, also fall within the scope of the present invention.

[0117] The present invention is not limited to the above-described embodiments, and various modifications are possible within the scope of the claims. Embodiments obtained by appropriately combining the technical means disclosed in different embodiments are also included in the technical scope of the present invention. Furthermore, new technical features can be formed by combining the technical means disclosed in each embodiment.

[0118] For example, in the above-described embodiments, the user may be able to select whether or not to display text data. Similarly, the user may be able to select whether or not to display the avatar's speech. Furthermore, the display control device 100 may be implemented as a standalone device, such as an electronic device, including a set-top box as an example. In this case, the electronic device may have at least some of the functions shown in FIG. 1, FIG. 4, or FIG. 14, with the remaining functions being provided elsewhere. Alternatively, the electronic device may have all of these functions. In particular, the illustrated server device 200 may have multiple trained models LM. In this case, all of the multiple trained models LM may be provided in the electronic device, or at least some may be provided in the electronic device, with the other trained models LM provided elsewhere. Note that the "avatar" in the above-described embodiments refers to an object represented based on a character, which can be primarily selected or created by the user. The avatar does not necessarily have to be a user's alter ego as the general term implies; it may simply correspond to a predetermined image. [Explanation of symbols]

[0119] 1...Display device 100 Display control device 11 First acquisition section 12 First control section 21 Second acquisition section 22 Second control section 13. Synthesis section 10. Receiving unit 20 Sensing unit 30...Display section 40 Speaker

Claims

1. a first acquisition unit that acquires content data related to the content; a control unit that controls speech via an object displayed together with the content based on the content or the content data; An information processing device comprising:

2. The control unit The speech is controlled in accordance with an analysis result regarding the content obtained by analyzing the content or the content data. The information processing device according to claim 1 .

3. The control unit When the content is displayed, the display mode of the object is controlled so that the object faces the display area of ​​the content. The information processing device according to claim 2 .

4. The control unit Displaying a plurality of said objects together with said content; The display mode or speech mode of the object is changed depending on the speech destination of the object. The information processing device according to claim 2 .

5. The destination of the speech includes: A user who views the content; and A second object to be displayed together with the object Contains one of the following: The information processing device according to claim 4 .

6. An information processing device according to any one of claims 1 to 5; a display unit that displays both the content and the object; Equipped with The information processing device further includes a content display control unit that controls display of the content.

7. A television receiver comprising the display device according to claim 6.

8. a first acquisition unit that acquires content data related to the content; a control unit that controls speech via an object displayed together with the content based on the content or the content data; An information processing system comprising:

9. a first acquisition step of acquiring content data relating to the content; a control step of controlling utterances via an object displayed together with the content based on the content or the content data; An information processing method comprising:

Citation Information

Patent Citations

  • Interactive operation-supporting system, interactive operation-supporting method and recording medium

    JP2002041276A

  • Character control system for television receiver

    JP2003061007A

  • Remote operation method, system, user terminal, and viewing terminal

    JP2015194864A

  • Control device, control method, and control program

    JP2018180472A

  • Object control system and object control method

    JP2018187712A