Information processing device, display device, television receiver, information processing system, and information processing method

The information processing apparatus enhances user satisfaction by providing tailored information and speech modes based on user and content analysis, addressing the inadequacies of existing television systems in presenting relevant content-related information.

JP7716547B1Active Publication Date: 2025-07-31SHARP KK
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2024134407
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2024-08-09
Publication Date
2025-07-31
Estimated Expiration
2044-08-09

AI Technical Summary

Technical Problem

Existing television systems fail to provide suitable information to users alongside content, leading to reduced user satisfaction.

Method used

An information processing apparatus that acquires content data and controls the display and speech of objects alongside the content based on user and content analysis, allowing for tailored notification information and speech modes.

Benefits of technology

Enables the presentation of suitable information to users, improving user satisfaction by ensuring appropriate timing and mode of object interaction with the content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007716547000001_ABST
    Figure 0007716547000001_ABST
Patent Text Reader

Abstract

To realize a technology that can present suitable information to a user along with content. [Solution] The display control device (100) includes a first acquisition unit (11) that acquires content data related to the content, and a control unit (22) that controls speech via an object displayed together with the content based on the content or the content data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an information processing apparatus, a display apparatus, a television receiver, an information processing system, and an information processing method.

Background Art

[0002] Conventionally, a television apparatus that automatically provides an individual's media preference based on identification information of a user detected by a camera has been proposed. For example, in the technique described in Patent Document 1, the television apparatus identifies a user based on pre-registered user information and a face image detected by a camera, and displays media content related to the identified user or automatically changes settings.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In order to improve the user's satisfaction in watching television, it is important what information to provide to the user together with the content. However, the technique described in Patent Document 1 has a problem in this regard.

[0005] One aspect of the present invention aims to realize a technique capable of presenting suitable information to a user together with content.

Means for Solving the Problems

[0006] To solve the above problems, an information processing apparatus according to an aspect of the present invention includes a first acquisition unit that acquires content data related to content, and a control unit that controls speech via an object to be displayed together with the content based on the content or the content data.

[0007] Further, a display device according to another aspect of the present invention includes the above information processing apparatus and a display unit that displays both the content and the object, and the information processing apparatus further includes a content display control unit that controls the display of the content.

[0008] Further, a television receiver according to another aspect of the present invention includes the above display device.

[0009] Further, an information processing system according to another aspect of the present invention includes a first acquisition unit that acquires content data related to content, and a control unit that controls speech via an object to be displayed together with the content based on the content or the content data.

[0010] Further, an information processing method according to another aspect of the present invention includes a first acquisition step of acquiring content data related to content, and a control step of controlling speech via an object to be displayed together with the content based on the content or the content data.

Effects of the Invention

[0011] According to one aspect of the present invention, it is possible to present suitable information to the user together with the content.

Brief Description of the Drawings

[0012]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Mode for Carrying Out the Invention

[0013] 〔Embodiment 1〕 Hereinafter, an embodiment of the present invention will be described in detail.

[0014] <Display Device 1> FIG. 1 is a block diagram showing the configuration of a display device 1 according to the present embodiment. As shown in FIG. 1, the display device 1 includes a receiving unit 10, a sensing unit 20, a display control device 100, a display unit 30, and a speaker 40. Note that each unit shown in FIG. 1 is merely an example of the configuration of the display device 1 and is not limited to this configuration. For example, the display device 1 may include various configurations such as a remote control device operated by a user, an operation signal receiving unit that receives an operation signal from the remote control device, or a storage unit for storing video data. In addition, each unit included in the display device 1 may be distributed among a plurality of devices via a network as an example. In such a case, the display device 1 may be expressed as a display system, or the display control device 100 may be expressed as a display control system. Note that in the example shown in FIG. 1, the display device 1 includes the display control device 100, but the present embodiment is not limited to this, and the display control device 100 may be realized as a set-top box connected to the display device 1.

[0015] (Receiving Unit 10) The receiving unit 10 receives content data and supplies the received content data to the display control device 100. The content data received by the receiving unit 10 includes, as an example, encoded video data, encoded audio data, and related information associated with the video data, but this does not limit the present embodiment. In addition, examples of the encoded video data include data (e.g., TS (Transport Stream)) encoded by various video encoding techniques such as MPEG2, MPEG4, H.264, and H.265, but this also does not limit the present embodiment.

[0016] In addition, as an example, the related information may include at least any one of program information regarding the content, program guide data including the program information, information on data broadcasting provided in relation to the content, and explanatory information regarding the content, etc., but these examples do not limit the present embodiment either. Also, as an example, the related information may be extractable from the content data without performing decoding processing by video encoding techniques such as MPEG2 described above. However, this example does not limit the present embodiment either.

[0017] Further, the receiving unit 10 may be configured to acquire the content data from the Internet via wireless or wired communication. Also, the receiving unit 10 may be configured to acquire the content data from a broadcast wave and may be configured to include a tuner for selecting any one of a plurality of channels included in the broadcast wave. When the receiving unit 10 includes a tuner, the display device 1 is also referred to as a television receiver (or simply a TV).

[0018] (Sensing unit 20) The sensing unit 20 senses one or more users who use the display device 1. As an example, the sensing unit 20 includes one or more cameras and images the one or more users with the cameras. Then, the sensing unit 20 supplies sensing data including the imaging data by the cameras to the display control device 100. Also, the sensing unit 20 may be configured to include a laser scanner (LiDAR device) that detects a user by reflection of laser light and supply the scan data acquired by the laser scanner to the display control device 100 included in the sensing data. Further, the sensing unit 20 may be configured to include one or more microphones that collect the utterances of the one or more users and supply voice data indicating the voice collected by the microphones to the display control device 100 included in the sensing data.

[0019] (Display unit 30) The display unit 30 displays the display image generated by the display control device 100. The display image includes at least one of a still image and a moving image (video). As an example, the display image includes at least one of the video of the content indicated by the content data (also referred to as content video) and the notification information to the user, as will be described later. Further, the display unit 30 may be configured to display at least one of the video of the content indicated by the content data (content video) and the notification information to the user.

[0020] As an example, the display unit 30 can be configured to include a display panel and a driver that drives the display panel based on the image data of the display image, but this does not limit the present embodiment. Further, as the display panel, a liquid crystal panel or an organic EL panel can be used, but this also does not limit the present embodiment. Note that the display example by the display unit 30 will be described later.

[0021] (Speaker 40) The speaker 40 outputs the voice of the content indicated by the content data to the user. Further, when the notification information includes voice (also referred to as notification voice), the notification voice is output to the user. In the present embodiment, the term "voice" can include a human voice, but is not limited thereto, and refers to sound in general that propagates through a medium such as air.

[0022] (Usage example of the display device 1) FIG. 2 is a diagram showing a usage example of the display device 1. In the example shown in FIG. 2, the display device 1 is realized as a stationary display device having legs. Further, as shown in FIG. 2, the display device 1 includes a sensing unit 20 that senses the user U, a display unit 30 that displays a display image, and speakers 40 (two in the example of FIG. 2) that output the voice of the content and the voice of the notification information. Note that the display device 1 may be realized as a wall-mounted display device.

[0023] <Display control device 100> Returning to FIG. 1, the configuration of each part of the display control device 100 included in the display device 1 will be described. As shown in FIG. 1, the display control device 100 includes a first acquisition unit 11, a first control unit 12, a synthesis unit 13, a second acquisition unit 21, and a second control unit 22. Note that the designations such as "first" and "second" do not limit this embodiment. For example, either one or both of the "first acquisition unit 11" and the "second acquisition unit 21" may be simply referred to as the "acquisition unit", and either one or both of the "first control unit 12" and the "second control unit 22" may be simply referred to as the "control unit". Also, the display control device 100 may be referred to as an information processing device. Further, the first control unit 12 may be expressed as a content display control unit.

[0024] (First acquisition unit 11) The first acquisition unit 11 acquires content data related to the content. As an example, the first acquisition unit 11 acquires the content data received by the above-described reception unit 10. The first acquisition unit 11 supplies the acquired content data to the first control unit 12.

[0025] (First control unit 12) The first control unit 12 controls the display of the content indicated by the content data supplied from the first acquisition unit 11 (in other words, the display of the content video). As an example, the first control unit 12 decodes the video data included in the content data, and supplies the content video obtained by the decoding process to the synthesis unit 13. The content video constitutes a part of the display image to be displayed by the display unit 30. Also, the first control unit 12 decodes the audio data included in the content data, and supplies the content audio obtained by the decoding process to the synthesis unit 13. The content audio constitutes a part of the audio output by the speaker 40. Further, the first control unit 12 extracts the above-described related information from the content data, and supplies the extracted related information to the second control unit 22 as an example.

[0026] (Second acquisition unit 21) The second acquisition unit 21 acquires sensing data of one or more users who use the display device 1 from the sensing unit 20. Here, as described above, the sensing data includes at least one of imaging data including the user in the viewing angle and voice data including the speech of the user. The second acquisition unit 21 supplies the sensing data to the second control unit 22.

[0027] Note that the sensing data is an example of user information regarding the user of the display device 1. Therefore, the second acquisition unit 21 may be expressed as acquiring user information regarding the user of the display device 1. Also, in the present embodiment, the user information may include imaging data and voice data acquired by the sensing unit 20 or another device.

[0028] Also, the sensing data is an example of user identification information for identifying the user of the display device 1. Therefore, the second acquisition unit 21 may be expressed as acquiring user identification information for identifying the user of the display device 1.

[0029] Also, the second acquisition unit 21 may be configured to acquire preference information corresponding to the user identification information. As an example, the second acquisition unit 21 may function as a preference information acquisition unit that acquires preference information corresponding to the user identification information from a preference information storage unit (preference information DB) that stores preference information of a plurality of users.

[0030] (Second control unit 22) The second control unit 22 executes an analysis process for analyzing at least any one of the content data, the content video, the content audio, the related information, and the sensing data, and performs various processes with reference to the analysis result. As an example, the second control unit 22 performs a process of specifying the content of the content by analyzing the content video or the related information, and a process of specifying the speech or the state of the one or more users by analyzing the sensing data, and executes various processes with reference to the results of these processes.

[0031] As an example, the second control unit 22 determines (generates) the notification mode of the notification information to be notified to the user with reference to the sensing data. In other words, the second control unit 22 changes the notification mode of the notification information to be notified to the user according to the user information. Here, the notification information includes at least any one of a predetermined image and speech via the image. For example, the notification information includes an object whose display mode and / or speech mode changes according to at least any one of the content data and the sensing data. Here, the object is an example of a predetermined image. Further, the object may be, for example, a character generated by the character generation unit 204 described later, or an image obtained by CG processing an image acquired from the outside, or an image created by an image generation AI from text data related to speech. Therefore, it may be expressed that the notification mode of the notification information includes at least any one of the display mode of a predetermined image according to the user information and the speech mode of the speech according to the user information.

[0032] Further, as an example, the second control unit 22 may perform a process of determining a speech mode suitable for the personal information specified based on the user information. For example, the second control unit 22 may perform a process of determining a speech mode (such as a speech mode for an elderly male / a speech mode for children, etc.) according to the age group or gender, etc. (for example, elderly male / child, etc.) of the individual specified based on the user information.

[0033] Further, as an example, the second control unit 22 may perform processing of identifying the relative position of the user with reference to the imaging data and changing at least one of the display mode and the speech mode of the notification information according to the identified relative position. Here, the relative position refers to the position of the user starting from the display device 1, and includes, as an example, the direction from the display device 1 and the distance from the display device 1.

[0034] Further, the speech mode includes the speech content which is the content of the speech. The second control unit 22 may determine the speech content to the user with reference to at least one of the content and the user information. Also, the second control unit 22 may be configured to control the timing of the speech to the user based on at least one of the content data and the user information.

[0035] As an example, the second control unit 22 may execute at least one of the processing of identifying the number of one or more users with reference to the sensing data and the processing of determining whether at least one of the one or more users is a registered user, and determine the speech content to the user via a predetermined image (from a character, for example) according to the result of the executed processing.

[0036] Further, the second control unit 22 may make at least one of the determination of whether to speak, the speech timing, and the speech content as the control of the speech by the character. Also, the second control unit 22 may be configured to control the display mode of the character in correspondence with the control of the speech from the character.

[0037] Further, the second control unit 22 may be configured to control the utterance through an object (for example, an utterance from a character) based on the content or content data. Here, "control based on the content or content data" includes, for example, controlling based on at least any one of the content video, content audio, and related information obtained from the content data. Also, for example, the second control unit 22 may be configured to control the utterance according to the analysis result regarding the content obtained by analyzing the content or content data. Here, "analyze the content" refers to analyzing the content video and content audio, but is not limited thereto, and may also include analyzing related information. Note that the above analysis result may be analyzed by the display control device 100 or the display device 1, or may be obtained from the cloud.

[0038] Further, the second control unit 22 may be configured to determine the notification mode of the notification information to be notified to the user based on the preference information. In other words, the second control unit 22 may be configured to determine the notification mode of the notification information according to the preference information of the user specified by referring to the sensing data. For example, the second control unit 22 may be configured to determine at least any one of the display mode and the utterance mode of the character included in the notification information by referring to the preference information.

[0039] Further, the second control unit 22 may be configured to execute at least any one of the process of determining the display mode of the image based on the preference information and the process of generating an utterance based on the preference information from a plurality of voice and sound candidates. For example, the second control unit 22 may be configured to execute at least any one of the process of determining the character to be displayed together with the content by referring to the preference information from a plurality of character candidates and the process of selecting the voice and sound of the character's utterance by referring to the preference information from a plurality of voice and sound candidates.

[0040] (Composition unit 13) The synthesizing unit 13 synthesizes the notification information determined (generated) by the second control unit 22 and the content video supplied from the first control unit 12. More specifically, the synthesizing unit 13 synthesizes the notification video included in the notification information determined (generated) by the second control unit 22 and the content video supplied from the first control unit 12, thereby generating a display image including the content video and the notification video, and supplying the generated display image to the display unit 30.

[0041] Also, the synthesizing unit 13 integrates the notification voice included in the above notification information and the content voice supplied from the first control unit 12, thereby generating an output voice including the notification voice and the content voice, and supplying the generated output voice to the speaker 40.

[0042] FIG. 3 shows a display example of the display image generated by the synthesizing unit 13. As shown in FIG. 3, the display image generated by the synthesizing unit 13 and displayed by the display unit 30 includes a first area R1 for displaying the content video and a second area R2 for displaying the notification information. Further, the second area R2 includes a character CR included in the above notification information, text data UC indicating the speech content by the character CR, and text data UU indicating the speech by the user (corresponding to "you" in FIG. 3). Here, at least one of the display mode and the speech mode of the character changes according to at least one of the content data and the sensing data.

[0043] Although a specific example of the character CR does not limit this embodiment, as an example, it may be any of a character simulating a user, a character simulating a living body such as an animal, a character simulating a non-living body such as a robot, or a character other than the above. Note that the character CR may also be expressed as an avatar.

[0044] In addition, a predetermined website screen may be displayed together with or instead of the content video in the first region R1. The predetermined website screen referred to here is, for example, an EC (e-commerce) site capable of electronic commerce. The EC site may be a website related to the provision of services such as travel arrangements as well as goods.

[0045] (Specific configuration example of the display control device 100) Subsequently, with reference to FIG. 4, a specific configuration example of the display control device 100 according to the present embodiment will be described. FIG. 4 is a block diagram showing a specific configuration example of the display control device 100. In the following description, redundant explanations may be omitted for matters already described with respect to the display control device 100.

[0046] (First control unit 12) As shown in FIG. 4, as an example, the first control unit 12 includes a content playback unit 121. The content playback unit 121 decodes the video data included in the content data, and supplies the decoded content video to the synthesis unit 13 and the second control unit 22. As an example, the content playback unit 121 generates a decoded content video using decoding processing compliant with various video encoding technologies such as MPEG2, MPEG4, H.264, and H.265. Further, the content playback unit 121 extracts related information included in the content data from the content data, and supplies the extracted related information to the synthesis unit 13 and the second control unit 22. Also, the content playback unit 121 decodes the audio data included in the content data, and supplies the decoded content audio to the speaker 40. As described with reference to FIG. 1, the display control device 100 may be configured to include the synthesis unit 13, and the first control unit 12 may be configured to supply the content video and the content audio to the synthesis unit 13.

[0047] (Second control unit 22) As shown in FIG. 4, the second control unit 22 includes an analysis unit 221 and a notification information generation unit 222. The analysis unit 221 analyzes the sensing data supplied from the second acquisition unit 21, the decoded content video supplied from the content playback unit 121, and the related information supplied from the content playback unit 121. Examples of the analysis processing by the analysis unit 221 include processing for identifying the relative position of each of one or more users by analyzing the imaging data included in the sensing data, processing for identifying the state (posture, expression, emotion, etc.) of each of one or more users by analyzing the imaging data included in the sensing data, processing for identifying the speech content of each of one or more users by analyzing the voice data included in the sensing data, processing for identifying the relative position of each of one or more users by analyzing the voice data included in the sensing data, processing for identifying the content (scene, characters, actions of characters, speech of characters, etc.) at each time point of the content by analyzing the content video, and processing for identifying the title, characters, summary, etc. of the content shown in the content video by analyzing the related information. However, these examples do not limit the present embodiment.

[0048] Note that, as an example, various analysis processes by the analysis unit 221 can use a machine-learned inference model (prediction model) that executes various algorithms such as an object detection algorithm and a speech extraction algorithm. However, this does not limit the present embodiment. Also, the inference model may be configured as a part of the learned model LM described later, or may be realized as a model separate from the learned model LM described later. Further, the inference model may be provided as a configuration of the server device 200 described later, or may be provided as a configuration of the display control device 100. The results of the above analysis processing by the analysis unit 221 are supplied to the notification information generation unit 222.

[0049] The notification information generation unit 222 refers to the analysis result by the analysis unit 221 and generates notification information according to the analysis result. The notification video included in the generated notification information is supplied to the composition unit 13 and composited with the content video. Also, the voice included in the generated notification information is output from the speaker 40. Note that when the notification information generation unit 222 generates notification information including voice and outputs it from the speaker 40, it may be configured to control to suppress the volume of the voice of the content via the content playback unit 121.

[0050] As shown in FIG. 4, as an example, the notification information generation unit 222 includes a prompt generation unit 201, a speech content acquisition unit 202, a voice generation unit 203, a character generation unit 204, a speech control unit 205, and a conversation history management unit 206.

[0051] (Prompt Generation Unit 201) The prompt generation unit 201 refers to the analysis result by the analysis unit 221, generates input data (prompt) for input to the learned model LM, and inputs the generated input data to the learned model LM. Here, as an example, the learned model LM may be configured to be included in the server device 200 connected to the display control device 100 via the network N as shown in FIG. 4, or may be configured to be included in the display control device 100. Also, as an example, the learned model LM may be a large language model, but is not limited thereto. Any machine-learned generation model can be used as the learned model LM. The learned model LM may be a model that combines a plurality of models, or may be a single multimodal model. Also, the prompt generation unit 201 may be configured to save the generated prompt as a history.

[0052] The specific prompt generated by the prompt generation unit 201 is not limited to this embodiment. As an example, it can be configured to include reference information such as the content of the conversation between the user and the character up to the current time, the content of the current content, the current state of the user (sensing data), etc., and instruction information for generating the speech content to the user based on the above reference information.

[0053] (Speech content acquisition unit 202) The speech content acquisition unit 202 acquires the speech content generated by the learned model LM based on the prompt generated by the prompt generation unit 201. As an example, the speech content acquisition unit 202 acquires the above speech content in the form of text data, but this is not limited to this embodiment. Note that the speech content acquisition unit 202 may be configured to associate the generated speech content with speech destination information indicating the speech destination of the speech content. The association of the speech destination information may be performed when the learned model LM generates the speech content. The speech destination includes, for example, any one of the user who views the content, the character (itself), and the second character displayed together with the character.

[0054] In addition, the speech content acquisition unit 202 may be configured to acquire candidates for the speech content by referring to at least one of the content data and the user information. Specifically, the speech content acquisition unit 202 may be configured to acquire candidates for the speech content from the character by referring to at least one of the content and the sensing data. In this case, the speech content acquisition unit 202 does not acquire candidates for the speech content, for example, when the content is serious and not familiar with the conversation with the character, or when the length of the candidate for the speech content (the time required to read aloud) is too long for the state of the user (uneasy, in a cheerful mood), etc.

[0055] The processing by the prompt generation unit 201 and the speech content acquisition unit 202 can also be expressed as a process of generating input data (prompt) according to the result of the analysis processing by the analysis unit 221, and determining the speech content from the character to the user by inputting the generated input data (prompt) into the learned model.

[0056] (Speech control unit 205) The speech control unit 205 changes the speech mode of the character (object) according to the speech destination of the character. The speech control unit 205 specifies the speech destination by referring to the speech destination information associated with the speech content. Here, when the character or the second character is specified as the speech destination, the speech control unit 205 supplies the speech content to the speech generation unit 203.

[0057] When the user is specified as the speech destination, the speech control unit 205 specifies the user's speech timing by referring to the sensing data. That is, the speech control unit 205 determines whether the user is about to speak. Then, at the user's speech timing, the speech control unit 205 controls so that the character does not speak. That is, the speech content acquired by the speech content acquisition unit 202 is not supplied to the speech generation unit 203. On the other hand, when it is not the user's speech timing (waiting for the user to speak), the speech control unit 205 supplies the speech content acquired by the speech content acquisition unit 202 to the speech generation unit 203.

[0058] Also, the speech control unit 205 may be configured to specify the user's state by referring to the sensing data. Specifically, the speech control unit 205 may refer to the analysis result of the sensing data by the analysis unit 221 and determine whether the user's state is a predetermined state. The predetermined state includes, for example, a state where the user is concentrating on the content (watching the content with a serious expression), a state where the user is performing an operation other than watching the content (such as reading a book), and the like.

[0059] When the speech control unit 205 determines that the user's state is a predetermined state, for example, it may be configured to prevent the character from speaking (not supply the speech content to the speech generation unit 203), or to control the character to speak to someone other than the user.

[0060] Further, the speech control unit 205 may be configured to determine whether the content being displayed belongs to a specific category or a specific scene. Whether the content being displayed belongs to a specific category or a specific scene means, for example, that the content being displayed contains content that is not preferably viewed (should be viewed seriously) with the character. Then, when the speech control unit 205 determines that the content belongs to a specific category or a specific scene, it may be configured to control so that the character does not speak.

[0061] Further, the speech control unit 205 may be configured to determine, as the speech content, a candidate for speech content in which it is recorded that no speech has occurred in the conversation history. That is, speech content that has been generated / acquired but the character has not actually spoken may be spoken later. In this case, the speech control unit 205 performs a process of sequentially determining whether it is the timing to speak the determined speech content. Then, when it is the timing to speak, the speech control unit 205 causes the character to speak the speech content.

[0062] Further, the speech control unit 205 may be configured to change the acquired candidate content by referring to at least any one of the candidate content, the content, and the sensing data. For example, when the speech content is too long for the user's situation or the scene of the content, the speech control unit 205 changes the candidate content by performing processing such as shortening the speech content.

[0063] (Conversation history management unit 206) When the conversation history management unit 206 determines not to utter a candidate for the utterance content, it stores a history indicating that the candidate was not uttered as part of the conversation history between the user and the character. Here, the conversation history management unit 206 stores the history in the conversation history database 23 provided in the display control device 100. Note that the conversation history management unit 206 may be configured to store at least one of the utterance content uttered by the character and the utterance content of the user in the conversation history database 23.

[0064] (Voice generation unit 203) The voice generation unit 203 generates voice data indicating the utterance content supplied from the utterance control unit 205. Here, the voice generation unit 203 may determine the voice tone and accent of the utterance of the voice data with reference to the analysis result by the analysis unit 221. Further, the prompt generated by the above-described prompt generation unit 201 may include instruction information for designating a voice tone, and the voice generation unit 203 may determine the voice tone and accent of the utterance of the voice data by referring to a response by the learned model LM to the instruction. The voice generated by the voice generation unit 203 is supplied to the synthesis unit 13. Then, the voice is supplied from the synthesis unit 13 to the speaker 40 together with the voice of the content.

[0065] (Character generation unit 204) The character generation unit 204 generates notification information with reference to the analysis result by the analysis unit 221. As an example, the character generation unit 204 generates (determines) character video data (character video data) to be included in the notification information with reference to the analysis result by the analysis unit 221. The character may have a form that allows for empathy and is not limited to a human. As an example, the character generation unit 204 determines the appearance including the character's appearance, the character's clothing, the character's movements, etc., with reference to the analysis result by the analysis unit 221, and supplies a character video expressing the determined content to the composition unit 13. Note that the character generation unit 204 may be configured to generate an image of a speech balloon that displays the speech content of the character as text together with the character video. Further, the character generation unit 204 may be configured to create at least a part of the character video with a generation AI. At least a part of the character may be, for example, an image obtained by subjecting an image acquired from the outside by the sensing unit 20 to CG processing, or an image created by an image generation AI from text data related to speech.

[0066] Further, the character generation unit 204 may be configured to generate a plurality of character videos. In this case, the composition unit 13 that receives the supply of the plurality of character videos will display the plurality of characters together with the content.

[0067] Also, the character generation unit 204 may be configured to change the display mode of a character (object) according to the destination of the speech of the character. The character generation unit 204 may identify the destination of the speech by referring to the destination information associated with the speech content, or may refer to the destination of the speech identified by the speech control unit 205. Here, when the user is identified as the destination of the speech, the character generation unit 204 generates a character video facing the direction where the user exists, for example, the front. At this time, the character generation unit 204 may be configured to generate display data (for example, character display data of "calling") indicating that the character is talking to the user together with the character video. Also, at this time, the character generation unit 204 may be configured to generate a balloon image with a display mode (for example, color, character size) different from the case where the destination of the speech is a character or a second character.

[0068] Also, when the second character is identified as the destination of the speech, the character generation unit 204 generates a character video facing the direction where the second character appears (for example, horizontal), for example. Also, when the character itself is identified as the destination of the speech, the character generation unit 204 generates a video of the character facing the direction where the first region R1 (the region for displaying the content) exists.

[0069] Also, the character generation unit 204 may be configured to cause the character to perform a predetermined action when the speech control unit 205 determines that the character does not speak. The predetermined action includes, for example, nodding, a surprised reaction, and the like.

[0070] (Flow of processing by the display control device 100) FIG. 5 is a flowchart showing a part of the flow of processing by the display control device 100.

[0071] (Step S11) In the example shown in FIG. 5, first, in step S11, the first acquisition unit 11 acquires content data. Since specific examples of the content data have been described above, the description is omitted here.

[0072] (Step S221A) Subsequently, in step S221A, the analysis unit 221 analyzes the content data acquired in step S11. Since specific examples of the analysis process by the analysis unit 221 have been described above, the description is omitted here.

[0073] (Step S21) Subsequently, in step S21, the second acquisition unit 21 acquires sensing data from the sensing unit 20. Since specific examples of the sensing data have been described above, the description is omitted here.

[0074] (Step S221) Subsequently, in step S221, the analysis unit 221 analyzes the sensing data acquired in step S21. Since specific examples of the analysis process by the analysis unit 221 have been described above, the description is omitted here.

[0075] (Step S222) Subsequently, in step S222, the notification information generation unit 222 generates notification information with reference to the analysis result in step S221.

[0076] (Specific processing example 1) FIG. 6 is a flowchart showing a specific processing example 1 by the display control device 100. The start condition of this processing is not limited to this, but for example, it is started when content is played, an operation to call a character is made by a user's operation or speech on the TV, or the sensing unit 20 detects the user.

[0077] (Step S221A) First, in step S221A, the analysis unit 221 specifies the content at each time point of the content by analyzing the content video.

[0078] (Step S222A) Subsequently, in step S222A, the prompt generation unit 201 refers to the analysis result by the analysis unit 221 to generate a prompt, and inputs the generated prompt into the learned model LM. Then, the learned model LM generates the speech content of the character.

[0079] (Step S222B) Subsequently, in step S222B, the learned model LM or the speech content acquisition unit 202 associates the destination information with the generated speech content.

[0080] (Step S222C) Subsequently, in step S222C, the speech control unit 205 refers to the destination information associated with the speech content to identify the destination. When the second character is identified as the destination, it proceeds to step S222D. When the content is identified as the destination, it proceeds to step S222E. When the user is identified as the destination, it proceeds to step S222F.

[0081] (Step S222D) When the second character is identified as the destination, in step S222D, the notification information generation unit 222 controls so that the character talks to the second character. Specifically, the speech control unit 205 supplies the speech content associated with the destination information indicating the character as the destination to the speech generation unit 203, and the character generation unit 204 generates the image of the character when talking to the second character. As a result, in the second area R2 of the display unit 30, as shown in FIG. 7, for example, the second character CR2, the character CR1 facing the direction where the second character CR2 exists, and the text data UC of the content (for example, "Avatar 2, this screen is interesting") for talking to the second character CR2 are displayed. And a voice talking to the second character can be heard from the speaker 40.

[0082] (Step S222E) When the content is specified as the destination, in step S222E, the notification information generation unit 222 controls the character to comment on the content. Specifically, the utterance control unit 205 supplies the utterance content associated with the utterance destination information indicating the content as the utterance destination to the voice generation unit 203, and the character generation unit 204 generates an image of the character who makes a monologue. As a result, in the second area R2 of the display unit 30, as shown in FIG. 8, for example, a character facing the direction where the first area R1 exists and text data UC of a comment on the content (for example, "This movie is interesting") are displayed. And a comment can be heard from the speaker 40.

[0083] (Step S222F) When the user is specified as the destination, in step S222F, the utterance control unit 205 specifies the user's utterance timing with reference to the sensing data. If it is not the user's utterance timing, the process proceeds to step S222G, and if it is the user's utterance timing, the process proceeds to step S222H.

[0084] (Step S222G) When it is not the user's utterance timing (waiting for the user to speak), in step S222G, the notification information generation unit 222 controls the character to speak to the user. Specifically, the utterance control unit 205 supplies the utterance content associated with the utterance destination information indicating the user as the utterance destination to the voice generation unit 203, and the character generation unit 204 generates an image of the character when speaking to the user. As a result, in the second area R2 of the display unit 30, as shown in FIG. 9, for example, a character CR facing the direction where the user exists, text data UC of the content to speak to the user (for example, "This movie is interesting"), and display data CALL indicating that the character is speaking to the user are displayed. And a voice speaking to the user can be heard from the speaker 40.

[0085] (Step S222H) If it is the user's speech timing, in step S222H, the notification information generation unit 222 controls so that the character does not speak. Specifically, the speech control unit 205 does not supply the speech content to the voice generation unit 203, and the character generation unit 204 generates an image of a character in a silent state. As a result, in the second area R2 of the display unit 30, as shown in FIG. 10, for example, a character CR facing the direction where the user exists and text data UC indicating silence (for example, "···") are displayed. And no voice of the character can be heard from the speaker 40 (only the voice of the content can be heard).

[0086] (Step S222J) Subsequently, in step S222J, it is determined whether the content is being played. If it is determined that the content is being played, the process returns to step S221A, and if it is determined that the content is not being played, the process ends.

[0087] (Specific processing example 2) FIG. 11 is a flowchart showing a specific processing example 2 by the display control device 100. The flow from step S221A to step S222C in processing example 2 is common to processing example 1.

[0088] (Step S222K) If the user is specified as the speech destination in step S222C, in step S222K, the speech control unit 205 refers to the sensing data (or its analysis result) by the sensing unit 20 to specify the state of the user. And if it is specified that the state of the user is a predetermined state, the process proceeds to step S222E or step S222H, and if it is specified that the state of the user is not a predetermined state, the process proceeds to step S222G.

[0089] (Specific processing example 3) FIG. 12 is a flowchart showing a specific processing example 3 by the display control device 100. The flow from step S221A to step S222B in processing example 3 is common to processing example 1.

[0090] (Step S222L) After step S222B, in step S222L, the speech control unit 205 determines whether the content being displayed belongs to a specific category or a specific scene. If it is determined that the content belongs to a specific category or a specific scene, the process proceeds to step S222H; if it is determined that it does not belong, the process proceeds to step S222C.

[0091] (Modifications of specific processing examples 1 to 3) Note that the second control unit 22 of the display control device 100 may be configured to execute a process of controlling the display mode of an object (character) so that the object faces the display area side of the content when the content is being displayed (including when it is started or resumed), either among the processes shown in the above specific processing examples 1 to 3 or separately from the above processes.

[0092] Specifically, when the display of the content starts or resumes, the character generation unit 204 generates an image of the character CR facing the direction in which the first region R1 exists, as shown in FIG. 13 for example. If the character CR has been displayed before the display of the content starts or resumes, the character generation unit 204 generates an image of the character CR facing the direction in which the first region R1 exists, rather than the character CR that has been displayed so far. Note that the character generation unit 204 may generate an image in which only a part (for example, only the upper part from the neck as shown in FIG. 13) faces the direction as the image of the character CR facing the direction in which the first region R1 exists, or may generate an image in which the whole faces the direction. As a result, the image projected onto the entire display unit 30 looks like the user is watching the content in which the character CR is visually reflected in the first region R1 as shown in FIG. 13. Thereby, the sense of unity in enjoying the content together with the avatar is improved.

[0093] Note that while the character CR faces the direction in which the first region R1 exists (while watching the content), the second control unit 22 may be configured to store the content of the displayed content in a storage unit (not shown). In this way, for example, after the display of the content ends, it becomes possible to generate speech content based on the stored content of the content and cause the character CR to speak (have a conversation with the user about the content). Regarding the time while watching the content, it may include a state where only the sound is heard even if the character CR does not face the direction in which the first region R1 exists.

[0094] Also, when the display of the content ends or is interrupted, the character generation unit 204 generates an image of the character CR facing a direction different from the direction in which the first region R1 exists (for example, the side where the user exists).

[0095] Note that the character generation unit 204 may be configured to generate an image of the character CR facing the direction in which the first region R1 exists when the content in which the user is interested is displayed or resumed. The determination as to whether the content is the content in which the user is interested can be made based on, for example, the content name or genre (registered by the user) stored in advance in a storage unit (not shown), information regarding the user's preferences, and the like. Further, the determination as to whether the content is the content in which the user is interested can be made based on the result of analysis by the analysis unit 221 (for example, the user has an expression indicating interest in the content, a statement indicating interest in the content is made, etc.) based on the sensing data (imaging data and audio data acquired by the sensing unit 20 or other devices).

[0096] (Effect of the display device 1) As described above, in the display device 1 according to the present embodiment, a configuration is adopted in which the display device 1 includes a first acquisition unit that acquires content data regarding content, and a control unit that controls speech via an object to be displayed together with the content based on the content or the content data. According to the display device 1 configured as described above, since the notification mode of the notification information to be displayed together with the content is determined with reference to the sensing data, it is possible to present suitable information to the user together with the content.

[0097] Further, according to the display device 1 according to the present embodiment, since the character talks at a timing when it is appropriate to talk to the user, for example, when there is no need for the character to talk (for example, when the user is not viewing the content), or when the user does not want to be interrupted by the character (for example, when the user is seriously viewing the content), etc., the character will not accidentally talk to the user. As a result, the user and the character can communicate smoothly.

[0098] 〔Embodiment 2〕 Next, other embodiments of the present invention will be described. FIG. 14 is a block diagram showing a specific configuration example of the display control device 100A according to the present embodiment. The display control device 100A shown in FIG. 14 has the same configuration as the display control device 100 shown in FIG. 4. Further, the display control device 100A includes a generation unit 200A, and the generation unit 200A has a language model LM. Since the other configurations are the same as those of the display control device 100 shown in FIG. 4, redundant explanations are omitted.

[0099] [Example of Realization by Software] The functions of the display device 1 (hereinafter referred to as the "device") can be realized by a display control program for causing a computer to function as the device, and a display control program for causing a computer to function as each control block of the device (particularly each part included in the first control unit 12 and the second control unit 22).

[0100] In this case, the device includes a computer having at least one control device (for example, a processor) and at least one storage device (for example, a memory) as hardware for executing the display control program. By executing the display control program with this control device and storage device, each function described in each of the above embodiments is realized.

[0101] The display control program may be recorded on one or more computer-readable recording media, not temporarily. This recording medium may or may not be provided in the device. In the latter case, the display control program may be supplied to the device via any wired or wireless transmission medium.

[0102] Also, part or all of the functions of each of the above control blocks can be realized by a logic circuit. For example, an integrated circuit in which a logic circuit functioning as each of the above control blocks is formed is also included in the scope of the present invention. In addition to this, for example, it is also possible to realize the functions of each of the above control blocks by a quantum computer.

[0103] Also, each process described in each of the above embodiments may be executed by AI (Artificial Intelligence). In this case, the AI may operate in the above control device, or may operate in another device (for example, an edge computer or a cloud server, etc.).

[0104] 〔Summary〕 At least the following aspects are described in this specification.

[0105] (Aspect 1) A first acquisition unit that acquires content data related to content, A second acquisition unit that acquires user information related to a user, And a control unit that changes a notification mode of notification information to be notified to the user according to the user information An information processing apparatus comprising.

[0106] According to the above configuration, since the notification mode of the notification information to be notified to the user is changed according to the user information, it is possible to present suitable information to the user.

[0107] (Aspect 2) A first acquisition unit that acquires content data related to content, And a control unit that controls speech via an object to be displayed together with the content based on the content or the content data An information processing apparatus comprising.

[0108] (Aspect 3) The control unit,[[]] Controls the speech according to an analysis result related to the content obtained by analyzing the content or the content data The information processing apparatus according to Aspect 2.

[0109] (Aspect 4) The control unit,[[]] When the content is being displayed (including when it is started or resumed), control the display mode of the object so that the object faces the display area side of the content. The information processing apparatus according to Mode 3.

[0110] (Mode 5) The control unit Displays a plurality of the objects together with the content, Change the display mode or speech mode of the object according to the speech destination of the object. The information processing apparatus according to Mode 3.

[0111] (Mode 6) The speech destination includes A user who views the content, and A second object displayed together with the object Either one is included. The information processing apparatus according to Mode 5.

[0112] (Mode 7) The information processing apparatus according to any one of Modes 2 to 6, and A display unit that displays the content and the object together Comprising The information processing apparatus further includes a content display control unit that controls the display of the content. A display device

[0113] (Mode 8) A television receiver including the display device according to Mode 7.

[0114] (Mode 9) A first acquisition unit that acquires content data related to the content, A control unit that controls the speech via the object to be displayed together with the content based on the content or the content data An information processing system comprising

[0115] (Mode 10) A first acquisition step of acquiring content data related to content; A control step of controlling speech via an object to be displayed together with the content based on the content or the content data; An information processing control method comprising the above.

[0116] The display control device according to each aspect of the present invention may be realized by a computer. In this case, a program for realizing the display control device by operating the computer as each part (software element) provided in the display control device, and a computer-readable recording medium recording the same also fall within the scope of the present invention.

[0117] The present invention is not limited to the above-described embodiments, and various modifications are possible within the scope shown in the claims. Embodiments obtained by appropriately combining the technical means disclosed in different embodiments are also included in the technical scope of the present invention. Furthermore, by combining the technical means disclosed in each embodiment, new technical features can be formed.

[0118] For example, in the above-described embodiment, the presence or absence of display of text data may be selectable by the user. Similarly, the presence or absence of the avatar's uttered voice may be selectable by the user. Further, the display control device 100 may be realized as an electronic device including, for example, a set-top box as a stand-alone device. In this case, the electronic device may have at least some of all the functions shown in FIG. 1, FIG. 4, or FIG. 14, and other functions may be provided outside the electronic device. Alternatively, the electronic device may have all of these functions. In particular, there may be a plurality of learned models LM of the illustrated server device 200. In this case, all of the plurality of learned models LM may be provided in the electronic device, or at least a part thereof may be provided in the electronic device, and other learned models LM may be provided outside the electronic device. Note that the "avatar" in the above-described embodiment is an object represented based on a character, and mainly, can be selected or created by the user. The avatar does not have to be the user's own alter ego as a general term, and may correspond to a predetermined image.

Explanation of Reference Numerals

[0119] 1 ··· Display device 100 ··· Display control device 11 ··· First acquisition unit 12 ··· First control unit 21 ··· Second acquisition unit 22 ··· Second control unit 13 ··· Synthesis unit 10 ··· Reception unit 20 ··· Sensing unit 30 ··· Display unit 40 ··· Speaker

Claims

1. A first acquisition unit that acquires content data related to content, A control unit that controls speech via an object to be displayed together with the content based on the content or the content data Comprising, The control unit, Can identify the presence or absence of the user's speech according to the destination of the speech via the object, When the destination of the speech is other than the user, regardless of the presence or absence of the user's speech, the object is controlled so as to speak according to the content, an information processing device.

2. The control unit, Controls the speech according to the analysis result related to the content obtained by analyzing the content or the content data The information processing device according to claim 1.

3. A first acquisition unit that acquires content data related to content, A control unit that controls speech via an object to be displayed together with the content based on the content or the content data Comprising, The control unit, Can identify the presence or absence of the user's speech, Even if there is no speech from the user, the object is controlled so as to speak according to the content, Controls the speech according to the analysis result related to the content obtained by analyzing the content or the content data, When the content is being displayed, controls the display mode of the object so that the object faces the display area side of the content Information processing device.

4. A first acquisition unit that acquires content data related to content, A control unit that controls speech via an object to be displayed together with the content based on the content or the content data Comprising, The control unit, Can identify the presence or absence of the user's speech, Even if there is no speech from the user, the object is controlled so as to speak according to the content, Controls the speech according to the analysis result related to the content obtained by analyzing the content or the content data, Displays a plurality of the objects together with the content, Changes the display mode or speech mode of the object according to the destination of the speech of the object Information processing device.

5. For the destination of the speech, The user, and A second object that is displayed together with the object is included in any of The information processing apparatus according to claim 4.

6. An information processing apparatus according to any one of claims 1 to 5, a display unit that displays both the content and the object comprising The information processing apparatus further includes a content display control unit that controls the display of the content, and is a display device.

7. A television receiver including the display device according to claim 6.

8. A first acquisition unit that acquires content data related to content, a control unit that controls speech via an object to be displayed together with the content based on the content or the content data comprising The control unit can identify the presence or absence of a user's speech according to the destination of the speech via the object, and when the destination is other than the user, controls the object so that the object speaks according to the content regardless of the presence or absence of the user's speech. An information processing system.

9. A first acquisition step of acquiring content data related to content, a control step of controlling speech via an object to be displayed together with the content based on the content or the content data comprising In the control step, the presence or absence of a user's speech can be identified according to the destination of the speech via the object, and when the destination is other than the user, controls the object so that the object speaks according to the content regardless of the presence or absence of the user's speech. An information processing method.

10. A first acquisition unit that acquires content data related to content, a control unit that controls speech via an object to be displayed together with the content based on the content or the content data comprising The control unit controls the speech according to an analysis result related to the content obtained by analyzing the content or the content data, and controls the display mode of the object so that the object faces the display area side of the content when the content is being displayed. An information processing apparatus.

11. A first acquisition unit that acquires content data related to content, A control unit that controls speech via an object to be displayed together with the content based on the content or the content data is provided, wherein the control unit controls the speech according to the analysis result regarding the content obtained by analyzing the content or the content data, displays a plurality of the objects together with the content, and an information processing apparatus that changes the display mode or speech mode of the object according to the speech destination of the object.

Citation Information

Patent Citations

  • Interactive operation-supporting system, interactive operation-supporting method and recording medium

    JP2002041276A

  • Character control system for television receiver

    JP2003061007A

  • Media preferences

    JP2011504710A

  • Control device, control method, and control program

    JP2018180472A

  • Object control system and object control method

    JP2018187712A