Information processing device, processing method, and program

The information processing apparatus and method enhance user engagement by analyzing user interactions to generate impactful content through a language model, addressing the issue of irrelevant information provision in existing systems.

JP2026047693APending Publication Date: 2026-03-16SHARP KK
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-09-04
Publication Date
2026-03-16

AI Technical Summary

Technical Problem

Existing information processing systems fail to provide information that significantly impacts user behavior, often providing irrelevant content based on conversation history, leading to user dissatisfaction.

Method used

An information processing apparatus and method that includes an input unit, recognition unit, prompt generation unit, and output unit to analyze user interactions and generate prompts based on a language model, determining the degree of influence on the user and generating significant impacts when the threshold is met.

Benefits of technology

Enables the provision of information with a substantial impact on users, ensuring relevance and user satisfaction by filtering and prioritizing content that aligns with user preferences and behaviors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026047693000001_ABST
    Figure 2026047693000001_ABST
Patent Text Reader

Abstract

To provide an information processing device that can present information that has a significant impact on the user. [Solution] An information processing device comprising: an input unit for inputting content including at least an image or sound; a recognition unit for recognizing the content; a prompt generation unit for generating a prompt to be input to a language model; an output data acquisition unit for acquiring the output from the language model as output data; and an output unit for generating and outputting a response sentence from the output data, wherein the prompt generation unit determines the degree of impact on the user from the content recognized by the recognition unit, and generates the prompt including the content recognized by the recognition unit when the degree of impact is above a threshold.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to an information processing apparatus and the like.

Background Art

[0002] For example, as shown in Patent Document 1, there is disclosed an information processing system that can provide information suitable for a user's preference by storing and analyzing a conversation exchanged via an avatar as an action history.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] One of the objects of the present disclosure is to provide an information processing apparatus that can present information having a great influence on a user.

Means for Solving the Problems

[0005] The information processing apparatus of the present disclosure includes an input unit that inputs content including at least an image or voice, a recognition unit that recognizes the content of the content, a prompt generation unit that generates a prompt from the content of the content recognized by the recognition unit, an output data acquisition unit that acquires an output from the language model as output data, and an output unit that generates and outputs a response sentence from the output data. The prompt generation unit determines the degree of influence on the user from the content of the content recognized by the recognition unit, and generates the prompt when the degree of influence is equal to or greater than a threshold value.

[0006] The processing method in the information processing device of the present disclosure comprises an input step of inputting content including at least an image or sound; a recognition step of recognizing the content of the information processing device; a prompt generation step of generating a prompt to be input to a language model; an output data acquisition step of acquiring the output from the language model as output data; and an output step of generating and outputting a response statement from the output data. The prompt generation step calculates the degree of impact from the content recognized by the recognition unit, determines that the impact on the user is significant when the degree of impact is greater than or equal to a threshold, and generates the prompt including the content recognized by the recognition unit.

[0007] The program of this disclosure includes an input function for inputting content including at least images or sounds into a computer; a recognition function for recognizing the content of the content; a prompt generation function for generating a prompt to be input to a language model; an output data acquisition function for acquiring the output from the language model as output data; and an output function for generating and outputting a response sentence from the output data. The prompt generation function calculates the degree of impact from the content recognized by the recognition unit, determines that the impact on the user is significant when the degree of impact is above a threshold, and generates the prompt including the content recognized by the recognition unit. [Effects of the Invention]

[0008] This disclosure enables the provision of information processing equipment that can present information with a significant impact on users. [Brief explanation of the drawing]

[0009] [Figure 1] This is a diagram illustrating the system overview in the first embodiment. [Figure 2] This figure illustrates the hardware configuration of the display device in the first embodiment. [Figure 3] This is a diagram illustrating the software configuration in the first embodiment. [Figure 4]This figure illustrates (a) the configuration of the storage unit and (b) an example of content information in the first embodiment. [Figure 5] This diagram illustrates the basic processing flow in the first embodiment. [Figure 6] This figure illustrates (a) an example of a display screen, (b) an example of a prompt, and (c) an example of output data in the first embodiment. [Figure 7] This diagram illustrates the processing flow in the first embodiment. [Figure 8] This figure illustrates an example of operation in the first embodiment, showing (a) an example of a display screen and (b) an example of a display screen. [Figure 9] This figure illustrates an example of operation in the first embodiment, showing (a) an example of a display screen and (b) an example of a display screen. [Figure 10] This diagram illustrates the processing flow in the second embodiment. [Figure 11] This figure illustrates an example of operation in the second embodiment, showing (a) an example of a display screen and (b) an example of a display screen. [Figure 12] This figure illustrates an example of user attributes in the third embodiment. [Figure 13] This diagram illustrates the processing flow in the third embodiment. [Figure 14] This figure illustrates an example of operation in the third embodiment, showing (a) an example of a display screen and (b) another example of a display screen. [Figure 15] This figure illustrates an example of operation in the third embodiment, showing (a) an example of a display screen and (b) another example of a display screen. [Modes for carrying out the invention]

[0010] Generally, there is a known technology that recognizes a user's voice in an information processing device or the like, obtains a conversation history, and analyzes the obtained conversation history to provide information that affects the user's behavior. As information that affects the user's behavior, there is a system that identifies the user's preferences, for example, by analyzing an avatar displayed on a display device equipped with an information processing device or a conversation history of a conversation in a virtual space.

[0011] However, by analyzing the conversation history, there may be provided information that has little possibility of affecting the user's behavior (information that is not very necessary for the user). Obtaining information that has little possibility of affecting the user's behavior (information with little influence) and providing it to the user means that the information required by the user is not provided, and the user feels bothered. Thus, there has been a problem that even information regarding the user's preferences provided based on the conversation history may not provide the information required by the user.

[0012] Thus, the information processing device and the like of the present disclosure that solve the above-described problems will be described in the following embodiments while referring to the drawings. Note that the following embodiments describe the invention described in the claims as an example, and the technical scope of the present invention is not limited to the description of the following embodiments. Further, in the following embodiments, the case where the information processing device of the present disclosure is applied to a display device will be described, but the information processing device is not limited to the display device. For example, the information processing device may be a stand-alone device of a set-top box (STB) having a tuner function, or may be applied to a recording device using a hard disk or the like, a display control device that reproduces a recording medium, a projector, or the like. Further, the information processing device may be an information processing device such as a smartphone, a tablet, or a computer that a user uses as a device capable of displaying content, or may be an in-vehicle device such as a car navigation.

[0013] [1. First Embodiment] [1.1 About the System] [1.1.1 System Overview] FIG. 1 is a diagram for explaining an overview of system 1. System 1 has a display device 10 and is connected to a network NW. Also, the display device 10 may be able to recognize a user U who is around the device, for example. The display device 10 incorporates a camera 12 and a microphone 14 and may recognize the user U in front of the display device 10.

[0014] Also, system 1 may include a server device 20. The server device 20 is connected to the network NW, for example. Further, the server device 20 may store a large language model 22 as a language model.

[0015] The large language model 22 is one of the language models and is generally called LLM (Large Language Models). Although the large language model 22 is stored in the server device 20, it may also be stored in the display device 10. The display device 10 may store a small-scale language model (SLM) as a language model different from the large language model 22, for example, an edge LLM. That is, as a language model, a so-called large language model (LLM) with a large amount of learning data may be stored in the server device 20, and a language model (SLM) with less learning data than the large language model 22 may be stored in the display device 10 or other devices. Also, in the following specification, the language model stored in the display device 10 may be referred to as an edge language model.

[0016] The large language model 22 may be prepared by a service provider constructing system 1 or may use an external service. For example, large language models such as GPT (Generative Pre-trained Transformers, GPT-3, GPT-4, GPT-4o), PaLM (Pathways Language Model), LLaMA (Large Language Model Meta AI), and tsuzumi may be used.

[0017] Furthermore, the display device 10 can acquire and display content. The display device 10 may acquire content from the broadcasting station 30, from the distribution device 40, or from a recording medium. The broadcasting station 30 may also transmit content to the display device 10 via a network NW. For example, the broadcasting station 30 may transmit content to the display device 10 using terrestrial digital broadcasting, BS broadcasting, or CS broadcasting. The broadcasting station 30 may also transmit content using IP multicast broadcasting via a network NW.

[0018] In this embodiment, an "avatar" is used to output the response sentence, but for example, the response sentence may be displayed on the display screen of the display device 10 using a language model without using an avatar, or the display device 10 may read the response sentence aloud. Furthermore, the display device 10 may be a display system that includes a device that simply displays an external video signal and a device that outputs a video signal. For example, the display device 10 may include a system that includes a display and a recording and / or playback device connected via HDMI®.

[0019] [1.1.2 Terminology] Hereafter, terms used in this specification will be used in a manner that can be understood by those skilled in the art, for example as follows:

[0020] An "avatar" is an object represented based on a character. Users can primarily select or create them. An avatar only needs to correspond to a given character; for example, it could represent a virtual friend, secretary, or advisor—someone to converse with. Although called an avatar, it doesn't necessarily have to be a representation of the user themselves.

[0021] "Conversation" refers to the exchange of sentences, and in this embodiment in particular, it refers to the exchange and combination of input sentences and response sentences. Here, the input sentence is a sentence based on the voice spoken by the user. The response sentence is a sentence output or spoken by the avatar. These sentences are sometimes also called messages. When the avatar speaks a sentence (response sentence), it means, for example, that the control unit displays the sentence (message) or outputs the sentence (message) as sound. In this case, by linking with the display of the avatar, the user can interact with the avatar as if it were actually speaking.

[0022] A "prompt" is information used to instruct a language model (large-scale language model, small-scale language model) to generate output data that includes a response sentence.

[0023] "User state" refers to the user's state, which may, for example, be the user's state when they are in front of the display device 10, but preferably it may be the user's state when they are viewing content. For example, the user state may include the user's facial expression (laughing, surprised, crying, etc.), gaze direction, and face orientation that can be recognized from the image. The user state may also include the user's emotions (wow, amazing, I want it, etc.) that can be recognized from their voice. Furthermore, the user state may include the user's age, gender, and the number of users.

[0024] "Content" includes program content broadcast from broadcasting station 30 and distributed content distributed from distribution device 40. Content may also be content recorded on a recording medium (e.g., Blu-ray). Content includes video containing one or more images and audio, and may also include explanatory information, attributes, and subtitle information related to the content. While this explanation primarily uses video content as an example, it may also include still images, audio, text, and web pages as needed.

[0025] [1.2 Hardware Configuration] Next, the hardware configuration of the device in this embodiment will be described. The display device 10, the server device 20, and the distribution device 40 each include at least a control unit, a storage unit, and a communication unit. In addition, each device has general functions as needed. Figure 2 illustrates the hardware configuration of the display device 10.

[0026] The control unit 100 controls the entire display device 10. The control unit 100 realizes various functions by reading and executing various programs stored in the memory unit 110 (e.g., ROM 110A, storage 110C). The control unit 100 may be implemented by one or more control devices / arithmetic units (CPU (Central Processing Unit), SoC (System on a Chip)). Alternatively, the control unit 100 may be composed of a control circuit.

[0027] The memory unit 110 is one or more storage devices that store necessary data or programs. The memory unit 110 temporarily or permanently stores data and programs. For example, the memory unit 110 includes a ROM 110A, a RAM 110B, and a storage device 110C.

[0028] ROM110A is a non-volatile memory that can retain programs and data even when the power is turned off.

[0029] RAM110B is the main memory primarily used by the control unit 100 during processing. RAM110B is a rewritable memory that temporarily holds data including programs read from ROM110A and storage 110C, as well as execution results.

[0030] Storage 110C is a non-volatile storage device capable of storing programs and data. For example, it may consist of storage devices such as HDDs (Hard Disk Drives) or SSDs (Solid State Drives). Alternatively, Storage 110C may be configured as an externally connectable USB memory stick. Furthermore, Storage 110C may be, for example, a storage area located in the cloud.

[0031] The broadcast control unit 120 receives broadcast waves transmitted by the broadcasting station selected by the user, decodes video data from the broadcast waves and outputs it to the display unit 140, and decodes audio data from the broadcast waves and outputs it to the audio output unit 160. The broadcast control unit 120 may have components such as a digital tuner unit (terrestrial / BS / CS, etc.), an OFDM demodulation unit, a DEMUX unit, and an MPEG2 decoding unit. Furthermore, there may be multiple broadcast control units 120.

[0032] The operation unit 130 receives operations from the user, issues operation instructions to each function unit, and notifies the control unit 100 of operation signals corresponding to the received operations. For example, the operation unit 130 receives operation signals from a remote control or the like and operates the display device 10 according to the received operation signals. Alternatively, the operation unit 130 may perform control based on operations received using, for example, operation buttons on the display device 10 or software keys using a touch panel.

[0033] The display unit 140 is a display device capable of displaying the video of the received program and various information. The display unit 140 may be, for example, a liquid crystal display (LCD) or an organic electroluminescent (OLED) display, or other device capable of displaying video. Alternatively, the display unit 140 may be a projection device such as a projector.

[0034] The audio input unit 150 is an input device capable of receiving sounds from the surrounding area where the display device 10 is installed, such as a microphone. The audio input unit 150 may also consist of multiple input devices (for example, a microphone array). The audio input unit 150 is primarily used for inputting the voice of the user listening, but it can also receive general sounds such as ambient noise. The display device 10 can recognize the user's state from the audio input from the audio input unit 150.

[0035] The audio output unit 160 outputs the audio contained in the content. The audio output unit 160 may be a device such as a speaker or headphones. The audio output unit 160 only needs to output sound, and can output general sounds such as music or ambient sounds. There may be multiple audio output units 160.

[0036] The imaging unit 170 is an imaging device that captures images of the area surrounding the display device 10, and is, for example, a camera. The imaging unit 170 may consist of one or more imaging devices. The imaging unit 170 outputs the captured images as image signals. The imaging unit 170 may also output one or more images as a continuous video. From the images captured by the imaging unit 170, the display device 10 can recognize the user's state.

[0037] The communication unit 180 is a communication interface for communicating with other devices. For example, the communication unit 180 may be a network interface that can connect to a wireless LAN, or a network interface that can connect to Ethernet (registered trademark) via a wired connection. Alternatively, the communication unit 180 may be a communication device that can connect to a mobile communication network such as LTE / 4G / 5G / 6G.

[0038] Furthermore, the configuration shown in Figure 2 only needs to include the necessary components in the embodiment. For example, if the user's status is not acquired from the camera, the imaging unit 170 may not be necessary.

[0039] Furthermore, the display device 10 may connect to an external device via an interface unit, and the above-mentioned functions may be implemented by the external device. For example, the imaging unit 170 may be implemented by a camera device connected via a USB interface.

[0040] [1.3 Software Configuration] [1.3.1 Configuration of the Control Unit] Figure 3 is a diagram illustrating the configuration of the control unit within the software configuration. For example, each configuration in Figure 3 is realized when the control unit 100 executes a program stored in the memory unit 110. The control unit 100 also functions as a display control unit that displays content, avatars, etc., on the display unit 140. The display control unit may also include functions such as a content image output unit 1006 and an avatar image output unit 1046.

[0041] (content) The content acquisition unit 1002 acquires content from the broadcast control unit 120 or the communication unit 180. Alternatively, the content acquisition unit 1002 may acquire content stored in the storage unit 110.

[0042] The content playback unit 1004 plays (displays) the content acquired by the content acquisition unit 1002. The content playback unit 1004 separates the multiplexed image data from the content and outputs it to the content image output unit 1006. The content playback unit 1004 also separates the multiplexed audio data from the content and outputs it to the content audio output unit 1008. Any known method can be used to separate the image data (image signal) and audio data (audio signal) from the content.

[0043] The content image output unit 1006 outputs the image data contained in the content to the display unit 140 as an image (or a video as a series of images). Here, the display unit 140 displays the image of the content as the first image.

[0044] Furthermore, the content audio output unit 1008 outputs the audio data contained in the content as audio to the audio output unit 165. Here, the audio output unit 165 outputs the audio of the content as the first audio.

[0045] Furthermore, the content recognition unit 1010 recognizes image data, audio data, and content information contained in the content (for example, program information and subtitle information in the case of broadcast waves) and outputs them to the prompt generation unit 1032. The information about the content recognized by this content recognition unit 1010 is called content information.

[0046] In this embodiment, the content is recognized from the audio or images contained in the content and output to the content information storage area 1106. Here, the content includes information recognized in real time from the audio or images of the content based on the content, and information associated with the content, such as program information of the content and information that describes the content.

[0047] (User perception) The user recognition unit 1022 recognizes a user who is a viewer of the content based on the image captured by the shooting unit 155, and recognizes information about that user. Information about the user is called user information, and the user recognition unit 1022 can recognize the following as user information.

[0048] (1) The number of users. For example, the user recognition unit 1022 recognizes how many people are watching the content. Alternatively, the user recognition unit 1022 may recognize the number of people watching the content by using the direction of the face and the direction of the gaze.

[0049] (2) Identifying the user. For example, the user recognition unit 1022 recognizes who the viewer of the content is. The user recognition unit 1022 can identify the user by acquiring an image of the user's face (face image) and matching it with a face image of a user that has been registered in advance.

[0050] (3) User's facial expression. For example, the user recognition unit 1022 can recognize the user's facial expression.

[0051] The speech recognition unit 1024 recognizes the conversation (input text) spoken by the user based on the voice input unit 150. The speech recognition unit 1024 outputs the input text to the conversation processing unit 1030.

[0052] Furthermore, the speech recognition unit 1024 may recognize the user's emotions (for example, interjections or clapping sounds) as user information.

[0053] In other words, the user recognition unit 1022 can recognize user information based on an image, and the voice recognition unit 1024 can recognize user information based on voice.

[0054] The conversation processing unit 1030 outputs the input sentence, the past conversation history, and any other necessary supplementary information to the large-scale language model 22, and retrieves output data including the response sentence from the large-scale language model 22. This allows the conversation processing unit 1030 to perform a conversation using the input sentence entered by the user and the response sentence output by the avatar.

[0055] To explain in more detail, the conversation processing unit 1030 includes a prompt generation unit 1032 and an output data acquisition unit 1034.

[0056] The prompt generation unit 1032 generates prompts to be fed into the language model (e.g., the large-scale language model 22) based on the input information. The prompts may be in natural language format or in a predetermined format corresponding to the language model 1102 using tags, etc. The prompt generation unit 1032 generates prompts based on the necessary information from the input sentence, conversation history, and supplementary information. Hereinafter, information other than conversation-related information may be referred to as supplementary information.

[0057] Here, supplementary information other than the input text and conversation history may include, for example, the following: • Content information. This may include, for example, the name of the content, the content itself (such as content recognized from the image of the content, content recognized from the audio of the content, etc.). • User information. For example, in addition to basic user information (age, gender, number of people, etc.), it may also include information such as the user's facial expressions, movements, interests, and possessions. • Environmental information. This may include information about the user's environment, such as the date and time, room environment, and weather information.

[0058] • Product information. For example, this may include one or more pieces of information about a product on a designated shopping site. In addition to basic information such as product name and price, the product information may also include information such as product features, weight, size, and product images that can be obtained from the product sales site.

[0059] The prompt generation unit 1032 may, for example, generate a prompt to input to the large-scale language model 22 by inputting the above-mentioned information to the edge LLM once.

[0060] The conversation processing unit 1030 (prompt generation unit 1032) inputs a prompt containing the input sentence as an example of a language model to the large-scale language model 22 of the server device 20. The large-scale language model 22 outputs output data to the conversation processing unit 1030 according to the input prompt. The large-scale language model 22 is based on the one in the server device 20, but if it is stored in the storage unit 110, the control unit 100 can also use the one stored in the storage unit 110. Furthermore, the control unit 100 may choose to use the large-scale language model 22 or the language model stored in the storage unit 110 depending on the content of the conversation and the type of output data requested.

[0061] The conversation processing unit 1030 (output data acquisition unit 1034) acquires output data from the large-scale language model 22. Here, the output data may include the following information.

[0062] • Response to the input statement • Information that controls the information processing device, such as the avatar's facial expressions. • Product recommendation information The conversation processing unit 1030 outputs, for example, the input sentence included in the prompt and the response sentence included in the output data as a conversation to the conversation history storage processing unit 1050. The conversation history storage processing unit 1050 stores the conversation, including the input sentence and response sentence, and information related to the conversation as conversation history information in the conversation history information storage area 1110. The conversation history storage processing unit 1050 also outputs the conversation history to the user attribute acquisition unit 1060.

[0063] Furthermore, the conversation processing unit 1030 outputs the conversation to the avatar output control unit 1040. The conversation processing unit 1030 mainly outputs the response sentences contained in the output data and information about the avatar contained in the output data (for example, the avatar's facial expressions, etc.) to the avatar output control unit 1040.

[0064] The avatar output control unit 1040 functions as an avatar generation unit that generates avatars. The avatar output control unit 1040 controls, for example, the avatar's facial expressions, movements, and appearance. The avatar output control unit 1040 also determines what the avatar will say. The avatar output control unit 1040 then outputs a signal to the avatar image generation unit 1042 for generating an image of the avatar. Here, the avatar image may include not only an image representing the avatar but also a message spoken by the avatar. The avatar image generated by the avatar image generation unit 1042 is then output to the display unit 140 via the avatar image output unit 1046.

[0065] Furthermore, the avatar output control unit 1040 outputs audio signals and text data for generating the voice spoken by the avatar to the avatar voice generation unit 1044. The avatar voice generation unit 1044 generates the avatar's voice in accordance with the response sentence spoken by the avatar. The avatar voice output unit 1048 outputs the generated avatar voice via the voice output unit 160.

[0066] Here, the avatar image generation unit 1042 may acquire information about the voice emitted by the avatar from the avatar voice generation unit 1044. Alternatively, the avatar image generation unit 1042 may acquire information about the voice emitted by the avatar from the avatar output control unit 1040.

[0067] Here, the avatar image may change its facial expression or mouth shape in response to the response text. For example, the avatar can receive a response text from the avatar image generation unit 1042 or the avatar voice generation unit 1044 and change the display of the avatar in accordance with the response text.

[0068] Furthermore, the avatar image may be displayed superimposed on the content image, or it may be displayed separately in a different area.

[0069] The user attribute acquisition unit 1060 analyzes the conversation history input from the conversation history storage processing unit 1050 to acquire user attributes. Here, user attributes are attributes that include information contained in input sentences and response sentences included in the conversation history, and include information that can identify the user, such as the user's address, interests, and possessions. Alternatively, the user attribute acquisition unit 1060 may acquire attributes from user information recognized by the user recognition unit 1022, for example. The user attribute acquisition unit 1060 stores the acquired user attributes in the user attribute storage area 1112, for example, for each user.

[0070] [1.3.2 Memory Unit Configuration] The configuration of the memory unit 110 will be explained with reference to Figure 4(a).

[0071] Language model 1102 is a language model available on the display device 10. Language model 1102 may be, for example, an SLM or an LLM. Language model 1102 may function as an edge LLM when using the large-scale language model 22. Alternatively, language model 1102 may be used instead of the large-scale language model 22.

[0072] The user information storage area 1104 stores user information, which is information about the user. User information may include, for example, pre-stored users who use the display device 10, or it may be stored as a user each time it is recognized. User information may include, for example, user identification information such as the user's name, an icon representing the user, or an image representing the user. Furthermore, user information may also include the user's age, age group, and gender.

[0073] The content information storage area 1106 stores content information, which is information about the content. The content information may store information about the content being played, or it may store content information corresponding to the content stored in the storage unit 110.

[0074] Figure 4(b) shows an example of content information stored in the content information storage area 1106. The content information may include, for example, the content name (e.g., "News at X o'clock"), program information as information describing the content (e.g., content attribute "News", information about the cast "XX", content description "..."), and the content itself. The control unit 100 may store program information, for example, information obtained from an EPG (Electronic Program Guide).

[0075] The content may include, for example, speech recognition information based on the results of recognizing the audio contained in the content (e.g., "Today's temperature was 25 degrees.") and image recognition information based on the results of recognizing the images in the content (e.g., "building, tree, sky" as results of objects recognized from the image). Note that the speech recognition information may also be based on subtitle information instead of the results of speech recognition.

[0076] The content may be acquired and stored by the control unit 100 each time it acquires content (for example, while the control unit 100 is displaying content). The content may be stored temporarily or in chronological order. Furthermore, the content may be stored for a predetermined period of time (for example, 5 seconds, 10 seconds, 1 minute, etc.), which can be arbitrarily set by the user.

[0077] Furthermore, in this embodiment, the control unit 100 stores the result of image recognition as content, but it may also store the image itself. Also, the control unit 100 stores the speech recognition information as a sentence as content, but it may also store it as words or phrases.

[0078] The avatar information memory area 1108 stores avatar information, which is information about the avatar. The avatar information may include, for example, gender (male, female), age, face shape, arrangement of facial features, type of clothing, and type of accessories. The avatar information may also include attributes such as whether the avatar is human, cat, dog, or anime character. Furthermore, the avatar information may include information about the avatar's emotions and the sounds the avatar makes.

[0079] The conversation history information storage area 1110 stores conversation history information, which is information about the history of a conversation. The conversation history information storage area 1110 functions, for example, as a history storage unit. The conversation history information includes input sentences and response sentences as a history of a conversation. In addition to the conversation history, the conversation history information may store one or more necessary pieces of related information, such as user information, content information, and environment information. Furthermore, the conversation history information may store the conversation history as a single unit from the start of content viewing to the end of content viewing, or it may store the conversation history chronologically by date.

[0080] The user attribute memory area 1112 stores user attributes. User attributes are stored for each user. User attributes may also include information about user attributes, such as basic information like the user's identification information and address (viewing location), interest information indicating the user's preferences, and information about the user's possessions (possession information). Here, interest information may include an index (e.g., level of interest) for words and phrases (keywords) that the user has shown interest in. Possession information may include, for example, the type of possessions the user owns, the brand of the possessions, and the manufacturing date of the possessions.

[0081] The score information storage area 1114 stores a score for each keyword. For example, the score information storage area stores a score (e.g., "10") for a word (e.g., "tomorrow") as a keyword. The score information storage area 1114 may be set by the user, or it may have keywords and scores stored in advance. The score information may also store information obtained from a network, for example. Furthermore, the score information storage area 1114 may store a trained model that has been trained using machine learning. In this case, when the control unit 100 inputs a word, phrase, etc., into the trained model stored in the score information storage area 1114, the score corresponding to the input word, phrase, etc., is output.

[0082] [1.4 Overall Explanation] The following describes the overall flow of the conversation between the user and the avatar in this embodiment.

[0083] [1.4.1 Conversation Processing Flow] Figure 5 illustrates the general flow of conversation processing. While it is preferable that the following processes be performed by any of the configurations described in Figure 3, for the sake of explanation, these processes will be described as being performed by the control unit 100.

[0084] The control unit 100 receives the content (S10). The control unit 100 recognizes the content and acquires content information (S12).

[0085] When the control unit 100 detects that the user has spoken (S14; Yes), it recognizes the speech and creates an input sentence (S16). The control unit 100 also obtains necessary information from user information and conversation history information (S18). Here, the control unit 100 only needs to obtain at least the information necessary for the processing described later.

[0086] The control unit 100 selects and acquires one or more necessary pieces of information from the input text, content information, user information, conversation history information, and other supplementary information (S18). Then, the control unit 100 generates a prompt through a prompt generation process (S20). The control unit 100 inputs the generated prompt into the language model (for example, the large-scale language model 22) (S21).

[0087] When the control unit 100 obtains output data corresponding to a prompt from a language model (e.g., a large-scale language model 22) (S22; Yes), it generates a response sentence from the output data (S24).

[0088] Then, the control unit 100 outputs a response sentence from the avatar (S26). The control unit 100 repeatedly executes the process until the conversation processing is completed (S28; No → S10).

[0089] Furthermore, although Figure 5 illustrates the example of the user initiating the speech, the avatar may also initiate the speech. For example, the control unit 100 may execute processing from the prompt generation process (S20) in S20. In this case, the avatar will output a response before the user initiates the speech.

[0090] [1.4.2 Interaction between the display area and the large language model] Figure 6(a) shows an example of a display screen W10. The display screen W10 has a first display area R10 for displaying content, a second display area R12 for displaying an avatar, and a third display area R13 for displaying dialogue. The second display area R12 and the third display area R14 may be configured as a single unit.

[0091] Furthermore, it is preferable to display the conversation in a way that distinguishes between the user's utterance (input) and the avatar's utterance (response). For example, Figure 6(a) displays the user's utterance and the avatar's utterance in a way that distinguishes between them based on the direction of the speech bubble. Alternatively, the user's utterance and the avatar's utterance may be distinguished by, for example, background color, text color, font, etc.

[0092] Furthermore, in the display screen W10 of Figure 6(a), the first display area R10, the second display area R12, and the third display area R14 are displayed, but these display areas can be switched. For example, the user can configure the display screen W10 to display only the avatar by enabling only the second display area R12. In this case, the avatar may be adjusted to a size appropriate for the screen size of display screen W10. Similarly, for example, the user can configure the display screen W10 to display only the dialogue by enabling only the third display area R14. In this case, the dialogue may be adjusted to a size appropriate for the screen size of display screen W10, the font may be enlarged, or more dialogue than usual may be displayed. In addition, multiple display areas may be selected and displayed, such as the first display area R10 and the third display area R14, or the second display area R12 and the third display area R14. In the following display screens as well, the avatar, content, and dialogue can each be selectively displayed.

[0093] Furthermore, for the sake of explanation, the avatar, content, and dialogue were described using the example of switching display areas, but the control unit 100 may also switch the display of the content itself on and off while keeping the display area the same. Also, the display area includes the display screen. In addition, the control unit 100 may output the content displayed in each display area (display screen) to different displays. For example, the content displayed in the first display area R10 may be displayed on the display unit 140 of the display device 10, and the avatar and dialogue displayed in the second display area R12 and the third display area R14 may be output to an external display or to another terminal device (e.g., a smartphone) connected via the communication unit 180.

[0094] Figure 6(b) schematically shows an example of a prompt for input to a language model (e.g., large-scale language model 22) as natural language. Figure 6(c) schematically shows an example of output data output from the language model (e.g., large-scale language model 22) as natural language.

[0095] The prompt preferably includes the input sentence. For example, Figure 6(b) includes the input sentence, "Where is Tenri?". The prompt may also include one or more pieces of information, such as conversation history information, content information, user information, and user attribute information. Furthermore, the prompt may include this information that has been optimized once by the edge LLM.

[0096] The output data preferably includes a response statement. For example, Figure 6(c) includes the response statement, "Northern Nara Prefecture, are there many tourists?" The output data may also include one or more items such as avatar information (avatar's facial expressions, etc.), user attribute information, site information, and supplementary information.

[0097] [1.5 Processing Flow] Figure 8 is a diagram illustrating the processing flow in this embodiment. While it is preferable that each process be performed by the configuration described in Figure 3, for the sake of explanation, it will be described as being performed by the control unit 100.

[0098] [1.5.1 Prompt Generation Process] The control unit 100 obtains content information recognized by the content recognition unit 1010 from the content information storage area 1106 as the content content (S102). Next, the control unit 100 determines from the obtained content information whether the program information attribute (genre) is a specific attribute (S104). Here, as a specific attribute, for example, a classification or genre of content that contains a lot of facts may be specified. Also, the specific attribute may be predetermined or set by the user.

[0099] For example, as a specific content genre, in the classification of programs under the digital broadcasting standard, one could specify "news / current affairs," "sports," or "information / talk shows." In this case, the classifications "drama," "music," "variety," "movie," "anime / special effects," "documentary / educational," "theater / performance," "hobbies / education," "welfare," and "other" would represent programs that do not fall under these specific attributes.

[0100] The control unit 100 may, for example, obtain keywords from content information indicating the content if the content belongs to a specific attribute (genre) (S104; Yes → S106). For example, the control unit 100 may obtain keywords from content information recognized by the content recognition unit 1010, or it may obtain keywords from content information stored in the content information storage area 1106. The control unit 100 may obtain one or more keywords from the content information, for example, from program information, speech recognition information, and image recognition information.

[0101] (1) Acquisition of keywords based on the audio of the content The control unit 100 acquires keywords based on the audio of the content. For example, if the speech recognition information included in the content information contains the sentence "Snow was observed in Hokkaido," the control unit 100 first performs morphological analysis to break down the sentence into words such as "Hokkaido," "in," "snow," "ga," "observed," "sare," "mashita," "ta," and "." The control unit 100 may acquire these words as keywords.

[0102] Furthermore, the control unit 100 may perform syntactic and semantic analysis after morphological analysis to obtain meaningful phrases and words as keywords. For example, the control unit 100 may obtain keywords for each phrase, such as "In Hokkaido," "Snow," "Was observed." Morphological analysis, syntactic analysis, and semantic analysis can utilize any known method or algorithm.

[0103] Furthermore, the control unit 100 may acquire words of a specific part of speech from the analyzed words as keywords. For example, the control unit 100 may acquire specific parts of speech such as nouns, numerals, verbs, adjectives, and adjectival nouns as keywords.

[0104] Furthermore, the control unit 100 may, for example, acquire a word as a keyword if the noun is a proper noun (such as Osaka, Ikoma Mountain, or Yamato River) or a numeral (for example, a phrase containing a number (such as 161 yen, 1st place, or 3.25 times)).

[0105] Furthermore, the control unit 100 may acquire not only single words, but also combinations of multiple words or phrases as keywords. For example, if the words are "fireworks" and "festival," the control unit 100 may acquire "fireworks festival" as a single keyword. Also, if the words are "pollen," "is," and "a lot," the control unit 100 may acquire "a lot of pollen" as a single keyword.

[0106] Furthermore, when acquiring keywords, the control unit 100 may acquire keywords based on their relationship with other words. For example, the control unit 100 may acquire a word analyzed from the speech recognition information as a keyword if related words are included in the sentence. For example, when the word "pollen" is in a sentence, the control unit 100 will acquire the word "pollen" as a keyword if seasonal words such as "tomorrow" or "today" are also included in the same sentence. Also, when the word "pollen" is in a sentence, the control unit 100 will acquire the word "pollen" as a keyword if quantitative words such as "a lot" or "a little" are also included in the same sentence. In this way, when the first word and a related second word are included in the same sentence, the control unit 100 will acquire the first word as a keyword. The control unit 100 may also acquire the second word as a keyword in this case. In other words, the control unit 100 will acquire the first word and the second word as keywords when they are included in the same sentence, but will not acquire them as keywords if they are not included. Furthermore, the control unit 100 may determine whether the second word is included in the preceding sentence or in sentences within a predetermined time (for example, within a time frame of 10 seconds, 30 seconds, or a number of sentences such as 3 or 5), even if they are not the same sentence.

[0107] (2) Acquisition of keywords based on images of the content Furthermore, the control unit 100 may acquire keywords based on the image of the content. For example, the control unit 100 may acquire keywords from words stored as image recognition information as part of the content.

[0108] (3) Acquisition of keywords based on program information of content Furthermore, the control unit 100 may obtain keywords from the program information as part of the content.

[0109] (4) Get keywords from interrupt information The control unit 100 may obtain keywords from interrupt information that is displayed in the content being judged, such as disaster prevention information, temporary information, earthquake early warning, evacuation information, and breaking news.

[0110] Next, the control unit 100 calculates the impact of the content (S108). Here, if the content has a significant impact on the user (high impact), the control unit 100 generates a prompt (S110; Yes → S112). That is, if the impact is high, the control unit 100 generates a prompt and outputs it to the large-scale language model 22. When the large-scale language model 22 receives a prompt, it outputs output data in response to the prompt to the display device 10. Based on the output data, the control unit 100 outputs, for example, a response sentence included in the output data.

[0111] On the other hand, the control unit 100 does not generate a prompt when the content has little impact on the user (low impact) (S110; No). Since no prompt is generated, the control unit 100 has no prompt to output to the large language model 22, and no output data is output from the large language model 22. In other words, the output of response sentences with low impact on the user is suppressed (not output).

[0112] Here, "high impact" means that the impact is above a predetermined threshold (greater than the threshold). On the other hand, "low impact" means that the impact is below a predetermined threshold.

[0113] In this embodiment, the control unit 100 calculates the impact based on whether or not the information is new. The impact is calculated on a per-sentence basis, specifically for each sentence (sentence stored in the speech recognition information) contained within the audio of a single piece of content. The control unit 100 can calculate the impact using the following methods.

[0114] (1) Calculate the impact for each keyword. The control unit 100 calculates a score for each keyword acquired in S106. The control unit 100 then calculates the sum of the scores as the degree of influence on the content. Based on the score information stored in the score information storage area 1114, the control unit 100 calculates a score for each keyword included in the content. The control unit 100 may assign a score of 0 to keywords not included in the score information.

[0115] In particular, it is preferable that the control unit 100 calculates scores so that certain keywords indicating newness receive higher scores. For example, the scores corresponding to certain keywords may be set as follows: 10 for "yesterday", 30 for "now", 20 for "tomorrow", and 10 for "the day after tomorrow". In this embodiment, for example, the further back in time a certain keyword is, the lower the score may be set as it represents older information, while the higher the score may be set for keywords that are present or future.

[0116] (2) Add up the scores based on the combination of keywords to calculate the degree of influence. Furthermore, the control unit 100 may add a score value when calculating the impact of the content if the keywords are in a specific combination. For example, the control unit 100 may add an additional score if the keywords are in the combination of "tomorrow" and "pollen". For example, if the control unit 100 has a score of 20 for "tomorrow" and a score of 5 for "pollen", it may add a score of "30" to the combination of "tomorrow" and "pollen". As a result, the control unit 100 will calculate that the impact of the content would be 25 if it were just "tomorrow" and "pollen", but becomes 55 because it contains two keywords.

[0117] (3) Change the score and calculate the impact For example, the control unit 100 may calculate the degree of impact by changing the score obtained from the score information. Suppose the content includes, for example, a combination of a specific date and time as a keyword and the word "concert". The control unit 100 may, for example, set the highest score for "one month before" the ticket sales start date, which has a large impact on users. Furthermore, the control unit 100 may set the score lower as it approaches the day of the event, when the impact on users decreases. In this way, the control unit 100 may calculate the degree of impact of the content by changing the score set in the score information.

[0118] Furthermore, the control unit 100 may, for example, change the keyword score based on the difference between the current time and a reference time (reference time). For example, the control unit 100 may change the score obtained from the score information to 50% if the difference between the date and time indicated by a particular keyword and the reference time is one day ago, 30% if it is two days ago, and 10% if it is one week ago. This allows the control unit 100 to be set so that newer information has a higher score.

[0119] The control unit 100 may calculate the degree of impact by using one or any combination of the above methods as it sees fit. The control unit 100 may also determine that information that is not new, such as "hello," or information that remains unchanged, has little impact on the user.

[0120] In this way, the control unit 100 determines whether the content contains new information. This is because new content has a greater impact on the user. For example, the control unit 100 determines whether the content contains new information based on whether the impact level calculated in S108 is above a predetermined threshold. For example, when the impact level threshold is set to 50, the control unit 100 determines that the content contains new information if the impact level is 50 or higher. The impact level threshold may also be set by the user. For example, if the user wants to be less strict in determining new information, they may set the impact level threshold low to 30, and if they want to be stricter in determining new information, they may set the impact level threshold high to 80.

[0121] Furthermore, the impact threshold may be changed depending on the attributes of the content. For example, if the content attribute (genre) is "news / news reporting," which is estimated to have a large impact on users, the impact threshold may be set to 30, while if the content attribute (genre) is "information / talk show," which is estimated to have a small impact on users, the impact threshold may be set to 70.

[0122] Furthermore, the control unit 100 may set the impact threshold for each user. Alternatively, the control unit 100 may estimate and change the impact threshold based on the user's reactions and conversation history.

[0123] [1.6 Example of Operation] Figures 8 and 9 illustrate an example of the operation of the display device 10 in this embodiment. Figure 8 shows the operation of the display device 10 based on the content it displays and the output data including the response statement.

[0124] Figure 8(a) shows that avatar A110 is displayed on the display screen W110 of the display device 10. Avatar A110 can, for example, speak response sentences in response to content or user input. The threshold for determining whether content is new is set to 30. The control unit 100 performs speech recognition on the content's audio and stores "There will be a fireworks display in XX village next week" in the content's speech recognition information. The control unit 100 then obtains keywords from the speech recognition information. Here, for example, the control unit 100 obtains "next week," "XX village," "fireworks display," and "there will be" as keywords.

[0125] The control unit 100 calculates the influence of the acquired keywords on the content. For example, the control unit 100 calculates an influence of 40 from the score of "next week" which is 40. The control unit 100 determines that the content is new because the influence of the keywords is high. The control unit 100 generates a prompt based on the content and inputs it into the large-scale language model 22, obtaining output data that includes the response sentence "There will be a fireworks display in XX village next week." The control unit 100 then displays the response sentence as message M110, reacting to the content.

[0126] In this way, when the influence of the content exceeds a threshold, the control unit 100 displays a response sentence based on the content, or the avatar's facial expression, to the user via avatar A110, or informs the user via voice. To the user, it appears as though avatar A110 is reacting to the content.

[0127] Figure 8(b) shows that avatar A120 is displayed on the display screen W120 of the display device 10. The threshold for determining whether content is new is set to 30. The control unit 100 performs speech recognition on the caster's voice and stores "Last year, there was a fireworks display in XX town" in the speech recognition information of the content information. The control unit 100 acquires "last year" as a keyword, for example.

[0128] The control unit 100 calculates the influence of the acquired keywords on the content. For example, the control unit 100 calculates an influence of 5 from the score of 5 for "last year". Next, the control unit 100 determines that the content is not new because the calculated influence of the content is low. Since "last year" is older information than "next week", the control unit 100 does not generate a prompt for the content. As a result, avatar A120 ignores the content displayed in area R120.

[0129] In this way, the control unit 100 does not operate avatar A120, for example, when its impact on the content is low. From the user's perspective, avatar A120 appears to be ignoring the content.

[0130] Figure 9(a) shows that avatar A130 is displayed on the display screen W130 of the display device 10. The control unit 100 recognizes the caster's voice and stores "It will rain nationwide tomorrow afternoon" in the voice recognition information of the content information. The control unit 100 determines from the program information of the content information that the content attribute is "news / reporting" and sets the threshold for determining the content as new to 80. The control unit 100 obtains "tomorrow" and "rain" as keywords from the voice recognition information of the content information.

[0131] The control unit 100 calculates the degree of influence on the content from the acquired keywords. For example, the control unit 100 adds a score of 50 for "tomorrow". Here, the control unit 100 determines that the keywords "tomorrow" and "rain" are a specific combination, and adds another 40 to the score, calculating an influence of 90. Next, the control unit 100 determines that the content is new because the calculated influence of the content is high. The control unit 100 generates a prompt based on the content, inputs it to the large-scale language model 22, and obtains output data that includes the response sentence "It looks like it will rain tomorrow afternoon". The control unit 100 displays the response sentence as message M130.

[0132] Figure 9(b) shows avatar A140 on display screen W140 of the display device 10. Here, a drama program is displayed in area R140 of display screen W140. The control unit 100 determines from the program information of the content information that the content attribute is "drama" and does not obtain keywords from the speech recognition information of the content information. Thus, the control unit 100 determines that the content displayed in area R140 is fiction and not new information, and does not generate a prompt. Avatar A140 does not react to the content.

[0133] In this embodiment, a single threshold was used to determine whether the impact was high or low. However, there is no requirement for only one threshold. For example, a first threshold and a second threshold could be provided. If the impact is greater than or equal to the first threshold, it could be determined to be high, and if it is less than the second threshold, it could be determined to be low. For example, the control unit 100 could determine that the impact is normal if it is greater than or equal to the second threshold and less than the first threshold. For example, if the control unit 100 determines that the impact is high, the avatar may proactively speak about the content. If the control unit 100 determines that the impact is normal, it may conditionally output a response based on the user's input text or conversation history.

[0134] [1.7 Effects, etc.] The control unit 100, for example, determines that the content of the new information will have a significant impact on the user and generates a prompt. This makes it possible to provide an information processing device, for example, in which avatar A110 reacts to new information as information that is likely to influence the user's behavior (i.e., information that is important to the user).

[0135] Furthermore, the control unit 100 does not generate prompts for information with low impact, i.e., old information. Therefore, the output data, including the generated response sentence, does not contain information with little impact, such as past events. Consequently, the avatar A120 does not react to unimportant (low-impact) information, so the user does not experience any annoyance.

[0136] Furthermore, if Avatar A320 reacts every time, users may find it annoying and feel that it is simply reacting mechanically. However, in this embodiment, Avatar A320 reacts only when the content is new, so Avatar A320 does not react uniformly, resulting in a more human-like feel. Generally, people do not engage in conversation for every piece of information, so Avatar A320 in this embodiment also reacts in a human-like manner, giving users a more human-like feeling and a greater sense of familiarity compared to conventional avatars. In addition, Avatar A320, combined with the character's image, reactions, and movements, can provide users with a highly appealing experience.

[0137] [2. Second Embodiment] The second embodiment will now be described. The second embodiment is an embodiment that determines whether information is specific, as it is information that is likely to have a significant impact on user behavior.

[0138] The second embodiment will omit explanations of parts where the hardware and software configuration is the same as that of the first embodiment, and will focus on explaining the differences from the first embodiment.

[0139] [2.1 Processing Flow] Figure 10 is a diagram illustrating the processing flow in this embodiment. Figure 10 replaces Figure 7 of the first embodiment. Instead of S110 in Figure 7, S202 is executed.

[0140] First, the control unit 100 generates a prompt when the content has a significant impact on the user (high impact) (S202; Yes → S112). On the other hand, the control unit 100 does not generate a prompt when the content has a small impact on the user (low impact) (S202; No).

[0141] In this embodiment, the control unit 100 calculates the degree of influence based on whether or not the information is specific. The degree of influence is calculated on a per-sentence basis, specifically for each sentence (sentence stored in the speech recognition information) contained within the audio of a single piece of content. The control unit 100 can calculate the degree of influence in the following ways.

[0142] (1) Calculate the degree of influence from proper nouns or numerical information. The control unit 100 calculates a score for each keyword included in the content based on the score information stored in the score information storage area 1114. In particular, it is preferable for the control unit 100 to calculate a higher score for specific keywords that indicate concrete details, such as proper nouns. For example, the score may be set to 15 for "Company A," 30 for "Mount Fuji," and 20 for "Yamato River."

[0143] Here, the control unit 100 may set a higher score if the proper noun is well-known, or a lower score if the proper noun is not well-known. For example, the control unit 100 may obtain the ranking of search terms on a search engine (trend information) via the network, and based on the trend information, set a score of 30, for example, if the keyword's trend ranking is within the top 10.

[0144] Furthermore, the control unit 100 preferably calculates scores such as numerical information (digital information) keywords to be higher, similar to proper nouns. For example, the score for "161 yen" may be set to 10. Note that for numerical information, if it is variable numerical information such as exchange rates, temperature, or gasoline prices, a higher score may be set, while if it is non-variable numerical information such as telephone numbers, dimensions, or physical constants, a lower score may be set.

[0145] (2) The influence is calculated by adding up scores based on the combination of keywords. For example, the control unit 100 may add a score when calculating the impact of the content if the keywords are in a specific combination. For example, if there is a combination of a company's proper noun and the words "changed" or "replaced", the control unit 100 may add a score of 30 to the combination of the company's proper noun and the words "changed" or "replaced". As a result, the control unit 100 calculates that the impact of the content would be 25 if it only contained the company's proper noun and the words "changed" or "replaced", but becomes 55 because it contains two keywords.

[0146] In addition, if the content information being judged includes combinations such as proper nouns like product models or car types combined with the word "recall," or nouns whose numerical information fluctuates, such as "yen exchange rate," combined with numerical information, a score may be added. The control unit 100 may sequentially record nouns whose numerical information fluctuates in the score information storage area 1114 by learning the content or acquiring them via the network NW. Alternatively, the user may input nouns whose numerical information fluctuates by voice or other means, and record them in the information storage area 1114.

[0147] In this way, the control unit 100 determines whether the content is specific information. This is because when the content is specific, it has a greater impact on the user. As an example, the control unit 100 determines whether the content is specific information based on whether the impact calculated in S108 is above a predetermined threshold.

[0148] In this embodiment, the determination in S202 was performed instead of S110, but as an alternative process, for example, S202 may be executed before or after S110, or a process may be added to perform determinations in parallel with S108 and S110 or S202.

[0149] [2.2 Example of Operation] Figure 11 shows an example of the operation of the display device 10 in this embodiment. Figure 11 shows the operation of the display device 10 based on the content it displays and the output data including the response statement. Here, we will explain using an example related to the yen exchange rate and interest rates, where numerical information fluctuates.

[0150] Figure 11(a) shows that avatar A210 is displayed on the display screen W210 of the display device 10. The threshold for determining the content to be specific is set to 30. The control unit 100 performs speech recognition on the content's audio and stores "The yen exchange rate temporarily fell to 161 yen to the dollar" in the speech recognition information of the content. The control unit 100 acquires keywords such as "yen exchange rate" and "161 yen".

[0151] The control unit 100 calculates the degree of influence on the content from the acquired keywords. For example, the control unit 100 adds a score of 10 to "yen exchange rate". Here, the control unit 100 determines that there is a specific combination of the fluctuating noun "yen exchange rate" and the keyword "161 yen", and adds another 30 to the score, calculating the degree of influence as 40. Because the degree of influence on the content is high, the control unit 100 determines that the content is specific.

[0152] The control unit 100 generates a prompt based on the content, inputs it to the large-scale language model 22, and obtains output data including the response sentence "It seems the yen is depreciating." The control unit 100 then displays the response sentence as message M210.

[0153] Figure 11(b) shows that avatar A220 is displayed on the display screen W220 of the display device 10. The threshold for determining whether content is new is set to 30. The control unit 100 performs speech recognition on the content's audio and stores it in the content information's speech recognition information. The control unit 100 obtains keywords such as "US interest rates," "rise," "yen depreciation," and "progress." For example, the control unit 100 calculates an impact of 5 from the score of "yen depreciation," which is 5.

[0154] The control unit 100 determines that the content is not specific because the impact of the calculated content is low. The control unit 100 does not generate a prompt for the content because "yen depreciation" and "progress" are less specific information than "yen exchange rate" and "161 yen". Avatar A220 does not react to the content displayed in area R220 and can ignore it as information that is unlikely to influence user behavior (information with low impact).

[0155] [2.3 Effects, etc.] The control unit 100, for example, determines that the impact is high if the content is based on specific details, and low if the information is not specific, such as general or immutable information. As a result, avatar A210 reacts to specific details in response to content, but does not react to general statements or immutable information, such as general statements like "the system changes when the president changes" or physical constants like "the speed of light is approximately 300,000 km per second."

[0156] In this way, the control unit 100 generates prompts that include content that has a significant impact on the user, so it can obtain output containing that content from the large-scale language model 22. Furthermore, since the avatar A220 does not react to information that is not important to the user, the user will not feel annoyed.

[0157] Furthermore, if Avatar A320 reacts to every piece of information, users may find it annoying and perceive it as simply reacting mechanically. However, in this embodiment, Avatar A320 reacts to content that has a high impact on the user, resulting in a non-uniform reaction. This creates a more human-like feeling, allowing users to feel a sense of familiarity with Avatar A320. Combined with Avatar A320's reactions and other factors, this provides a highly desirable system for the user.

[0158] [3. Third Embodiment] The third embodiment will now be described. The third embodiment is an embodiment that allows selection of whether a highly influential prompt can be output and determines whether there are user attributes that are likely to influence user behavior.

[0159] The third embodiment will omit explanations of parts where the hardware and software configuration is the same as that of the first embodiment, and will focus on explaining the differences from the first embodiment. [3.1 Software Configuration]

[0160] The user attribute storage area 1112 stores user attributes that identify information influencing user behavior. User attributes are stored, for example, for each user using the display device 10. User attributes may include, for example, location information, interest information, and ownership information to identify the user's attributes. User attributes may also store user information together with user attributes.

[0161] Figure 12 shows an example of user attributes. User attributes include at least location information (e.g., Yamato-Koriyama City, Nara Prefecture), interest information (e.g., keywords "Chinese noodles, ramen," and interest level "80"), and ownership information (e.g., "mobile battery, Company A").

[0162] Here, the level of interest is, for example, a numerical representation of the user's interest in a keyword. For example, the user attribute acquisition unit 1060 analyzes the conversation history and sets the level of interest based on the frequency of its inclusion in the conversation history. As one example, if the user attribute acquisition unit 1060 finds a keyword in the conversation history, and the keyword is not stored in the user attribute storage area 1112, it may store the keyword and add 5 to the level of interest corresponding to the keyword. If the keyword has not been included in the conversation history for a certain period, such as a day or a week, it may subtract 5 from the level of interest corresponding to the keyword.

[0163] [3.2 Processing Flow] Figure 13 is a diagram illustrating the processing flow in this embodiment. Figure 13 is a replacement for Figure 7 of the first embodiment. After S102 in Figure 7, S302 is executed, and S304 is executed instead of S110.

[0164] First, the control unit 100 determines whether or not to generate a prompt based on the degree of influence (S302). The control unit 100 allows the user to choose between a mode in which a prompt is generated when the influence of the content is high, and a mode in which a prompt is generated regardless of the degree of influence. If the control unit 100 determines that a prompt should be output when the influence of the content is high, it determines whether the content information has a specific attribute (S302; Yes → S104). If it determines that a prompt should not be generated based on the degree of influence, it generates a prompt (S302; No → S112).

[0165] In this embodiment, the impact is calculated based on whether or not user attributes are included. The impact is explained as being calculated on a per-sentence basis, specifically for each sentence (sentence stored in the speech recognition information) contained within the audio of a single piece of content. The control unit 100 can calculate the impact based on whether or not user attributes are included, using the following methods.

[0166] (1) Calculate the degree of influence from the user's interests (interest information / level of interest) For example, if the control unit 100 contains keywords in the interest information stored in the user attribute storage area 1112, it adds a score based on the level of interest in the interest information. The control unit 100 may, for example, determine whether a keyword is included in the interest information in the user attribute storage area 1112, and if it determines that it is included, it may add a score based on the level of interest in the specific keyword. For example, if "ramen" is included in the content, and the same "ramen" is stored in the user attribute interest information, and the level of interest in "ramen" is "80", the control unit 100 may add a score of 80.

[0167] (2) Calculate the degree of influence from location information For example, the control unit 100 calculates a score for keywords included in the content from the location information stored in the user attribute memory area 1112. For example, if the location information is "Yamato-Koriyama City, Nara Prefecture" and the acquired keyword matches "Yamato-Koriyama City, Nara Prefecture", the control unit 100 adds a score of 30 to it as a specific keyword. On the other hand, if the keyword is different, such as "Mihama Ward, Chiba City" or "San Francisco City", the control unit 100 may calculate a score of 0. Other patterns for adding a score include the control unit 100 adding a score if the location is the same prefecture, the same regional division, or the same country. Furthermore, the score information memory area 1114 may also store proper nouns indicating places, such as famous places or event halls, associated with location information.

[0168] (3) Calculate the degree of influence from possessions For example, the control unit 100 calculates a score for keywords included in the content from the ownership information stored in the user attribute storage area 1112. For example, if "mobile battery" is stored in the ownership information, the control unit 100 adds 30 to the score if the retrieved keyword is the same as "mobile battery".

[0169] In this way, for example, information that influences user behavior and requires user alerting, such as "there was a problem with the mobile battery," "Company C is recalling its health food products," or "Company B has a recall notice for its cars," can be determined as high-impact information based on ownership information included in user attributes.

[0170] Here, the control unit 100 can calculate the degree of impact from user attributes when interrupt information is included in the content information. When the control unit 100 obtains interrupt information as a keyword, if the obtained keyword contains location information, it compares it with the location information included in the user attributes and adds a score if they match. For example, if the keyword obtained as interrupt information contains "Nara Prefecture", the control unit 100 adds 80 points to the score if the location information in the user attributes is "Nara Prefecture". Here, interrupt information refers to information provided separately from normal content, such as information included in emergency warning broadcasts (emergency warning signals), information included in text superimposed on the content, and information included in the L-shaped area of ​​an L-shaped screen.

[0171] Note that the user attributes to be compared are not limited to location information. For example, if the keyword "elderly evacuation" is obtained from interrupt information, a score may be added if the user's age matches the age previously categorized as elderly, after comparing it with the user's age stored in the user attributes.

[0172] By using one or a combination of the above methods, it is possible to determine that the information satisfies user attributes that influence user behavior.

[0173] In this way, the control unit 100 determines whether the content includes user attributes. This is because when the content includes user attributes, it has a greater impact on the user. As an example, the determination of whether the content includes user attributes is made based on whether the impact level calculated in S108 is above a predetermined threshold.

[0174] In this embodiment, we determined whether user attributes were included, but for example, the processing in S108 of the first embodiment and S202 of the second embodiment may also be performed in parallel.

[0175] [3.3 Example of Operation] Figures 14 and 15 illustrate an example of the operation of the display device 10 in this embodiment. Figure 14 shows the operation of the display device 10 based on the content it displays and the output data including the response statement.

[0176] Figure 14(a) shows that avatar A310 is displayed on the display screen W310 of the display device 10. At this point, user U10, who is viewing the display device 10, says M310, "It looks delicious." The control unit 100 recognizes user U10's speech M310 and displays it as message M312 on the display screen W310.

[0177] The control unit 100 performs speech recognition on the content's audio, stores it as speech recognition information, and obtains a keyword, for example, "ramen." Next, the control unit 100 generates a prompt from the input sentence "Looks delicious" and the content containing "ramen," inputs it to the large-scale language model 22, and obtains output data including the response sentence "That ramen looks delicious." Then, the control unit 100 displays the response sentence as message M314. Subsequently, the control unit 100 inputs the output data to the user attribute acquisition unit 1060 via the conversation history storage processing unit 1050. If the user is interested in "ramen" based on the output data including the input sentence and response sentence, the control unit 100 adds the interest level corresponding to the keyword "ramen" stored in the user attribute interest information and stores it in the user attribute storage area 1112.

[0178] Figure 14(b) shows that avatar A320 is displayed on the display screen W320 of the display device 10. The threshold for determining whether user attributes are included in the content is set to 50. The control unit 100 performs speech recognition on the content's audio and stores it in the speech recognition information of the content information. The control unit 100 acquires "ramen" as a keyword. The control unit 100 calculates an influence score on the content's content from the acquired keyword.

[0179] The control unit 100 determines that the keyword "ramen" is included in the interest information of the viewing user stored in the user attribute storage area 1112, and calculates the influence level as 80 based on the interest level of the interest information. Next, the control unit 100 determines that the user attribute is included in the content because the influence level of the content is high. The control unit 100 generates a prompt based on the content, inputs it to the large-scale language model 22, and obtains output data that includes the response sentence "A new ramen shop has opened." The control unit 100 displays the response sentence as message M320.

[0180] Figure 15(a) shows that avatar A330 is displayed on the display screen W330 of the display device 10. The ○△□ exhibition hall is located in Yamato-Koriyama City, Nara Prefecture, and for example, the user resides in Yamato-Koriyama City, Nara Prefecture. Here, the location information of the ○△□ exhibition hall is stored in the score information storage area 1114 along with the score. The threshold for determining whether content is included in the user attributes is set to 50.

[0181] The control unit 100 acquires a keyword, for example, "○△□ Exhibition Hall," based on the audio output of the content displayed in area R330. The control unit 100 determines that the location information of "○△□ Exhibition Hall" stored in the score information storage area 1114 matches the location information of the viewing user stored in the user attribute storage area 1112, and calculates the influence level as 80. Next, the control unit 100 determines that the content contains user attributes because the calculated influence level of the content is high.

[0182] Next, the control unit 100 generates a prompt based on the content and inputs it into the large-scale language model 22, obtaining output data that includes the response sentence "It seems there is an event at the ○△□ exhibition hall." Then, the control unit 100 displays the response sentence as message M330.

[0183] Figure 15(b) shows that avatar A340 is displayed on display screen W340 of the display device 10. In area R340 of display screen W340, an informational program is displayed as content, and interrupt information is displayed as message M342. The threshold for determining whether content contains user attributes is set to 50.

[0184] The control unit 100 obtains "Nara Prefecture" as a keyword from the interrupt information. The control unit 100 compares the obtained keyword, for example, "Nara Prefecture," with the location information of the viewing user stored in the user attribute storage area 1112 to see if they match. Since "Nara Prefecture" matches the location information stored in the user attribute, the control unit 100 calculates the influence level as 80. Next, the control unit 100 determines that the content contains user attributes because the calculated influence level of the content is high.

[0185] Next, the control unit 100 generates a prompt based on the content including the interrupt information, inputs it to the large-scale language model 22, and obtains output data containing information such as a response sentence, for example, "It looks like a flood is about to occur," and information to change the avatar's facial expression to a "surprised face." Then, the control unit 100 displays the response sentence as message M340 and changes the facial expression of avatar A340 to a "surprised face." In this way, avatar A340 reacts to the information displayed in area R340. Furthermore, avatar A340 can change its facial expression to alert the user depending on the content.

[0186] In this embodiment, an On / Off function for "Influential Conversation Only Mode" is provided, which allows response text to be displayed even for information with low impact via a settings screen or the like. When "Influential Conversation Only Mode" is On (first mode), the control unit 100, as an example of a setting unit, controls the system to generate prompts only when the impact level is above a predetermined threshold, and the avatar becomes a quiet avatar that does not talk about unimportant topics. The user can enjoy conversing with the avatar or concentrate on viewing content depending on their mood and the situation. On the other hand, when "Influential Conversation Only Mode" is Off (second mode), prompts are generated regardless of the impact level, so the avatar becomes a talkative avatar that also talks about unimportant topics. When the setting is Off, the user can enjoy more conversations with the avatar compared to when it is On.

[0187] Furthermore, in this embodiment, user attributes are stored for each user, but for example, the avatar displayed on the screen may individually possess only the interest information from the user attributes. For example, an avatar such as a ramen expert could be set up that determines that the keyword "ramen" has a high influence and generates prompts.

[0188] [3.4 Effects, etc.] The control unit 100 determines the impact level to be high when the content-based information is related to, for example, the user's attributes (e.g., preferences, possessions, location, etc.), and low when it is not related to the user's attributes. As a result, the control unit 100 generates prompts based on the content that has a significant impact on the user, but suppresses the generation of prompts for content that has a low impact. Therefore, the control unit 100 can obtain output data that has a significant impact on the user from the large-scale language model 22. Furthermore, even when the content-based information has little impact on the user, the control unit 100 can generate prompts based on interrupt information if interrupt information is available.

[0189] Furthermore, Avatar A320 will respond to content related to user attributes based on output data by outputting response messages, etc., but will not respond to content unrelated to user attributes. In addition, Avatar A340 can respond to important information based on interrupt information, even when the content has little impact on the user.

[0190] In this way, Avatar A320 reacts to things that are most influential to the user based on the user's attributes. As a result, users feel as if Avatar A320 is responding specifically to them, as it reacts to their preferences and things related to their place of residence, thus increasing the sense of personalization. Furthermore, by having Avatar A320 react to important information such as interruptions, users can feel as if Avatar A320 is concerned about them, thus fostering a greater sense of closeness to Avatar A320.

[0191] [4. Variant] This disclosure is not limited to the embodiments described above, and various modifications are possible. In other words, embodiments obtained by combining technical means that are appropriately modified within the scope of this disclosure are also included in the technical scope.

[0192] In this disclosure, prompt generation was performed after determining information that would have a significant impact on the user. Here, for example, when the output data acquisition unit 1034 acquires output data from the large-scale language model 22, it may determine whether or not the information would have a significant impact on the user, and if it is determined to be such information, it may output it to the avatar output control unit 1040.

[0193] Furthermore, although the embodiments described above are explained separately for the sake of explanation, they can be combined and implemented to the extent possible. In addition, we intend to obtain rights to any of the technologies described in this specification through amendments or divisional applications.

[0194] Furthermore, although the embodiments described above used HDMI as an example of the device connection format, other connection formats may also be used. For example, at the time of filing, DVI, VGA, DisplayPort, etc., may be used as connection formats that can utilize EDID, which indicates display capability. Also, although the embodiments described above use EDID as information indicating display capability, other information may be used as long as it indicates display capability. In addition, although two HDMI standards, the first standard and the second standard, have been used as examples in the explanation, three or more standards may be mixed. Even in this case, for example, the device that outputs video will be able to use the highest resolution among the resolutions that the device that displays video can output.

[0195] Furthermore, in each embodiment, the program that operates in each device is a program that controls the CPU and other components (a program that makes the computer function) in order to realize the functions of the embodiments described above. The information handled by these devices is temporarily stored in a temporary storage device (for example, RAM) during processing, and then stored in various ROMs or HDDs, and read, modified, and written by the CPU as needed.

[0196] Here, the recording medium for storing the program may be any of the following: semiconductor media (e.g., ROM or non-volatile memory card), optical recording medium or magneto-optical recording medium (e.g., DVD (Digital Versatile Disc), CD (Compact Disc), BD (Blu-ray® Disc)), magnetic recording medium (e.g., magnetic tape, flexible disk), etc.

[0197] Furthermore, when distributing the program to the market, it can be stored on a portable recording medium and distributed, or transferred to a server computer connected via a network such as the Internet. In this case, the storage device of the server device is, of course, also included in this disclosure.

[0198] Furthermore, the data mentioned above may not be stored within the device itself, but rather stored on an external device and retrieved as needed. For example, the data may be stored on a NAS (Network Attached Storage) or on the cloud.

[0199] Furthermore, the scope of this disclosure is not limited to the configurations explicitly described in the specification, but also includes combinations of the technologies disclosed herein. While the configurations for which patent protection is sought are described in the attached claims, there is no intention to exclude them from the technical scope simply because they are not described in the claims.

[0200] Furthermore, the phrases "in the case of..." and "when..." in the above-mentioned specification are explained as examples only, and do not represent a configuration limited to those described. Even for configurations other than those described, we disclose information that would be obvious to a person skilled in the art, and we intend to acquire rights to such information.

[0201] Furthermore, the descriptions of the processes and data flows described in the specification are not limited to the order in which they are described. For example, configurations in which parts of the process are deleted or the order is rearranged are also disclosed, and the company intends to acquire rights to them.

[0202] Furthermore, although the functions described in the embodiments are described as being performed by each device, they may also be implemented by a single device or by utilizing an external server. For example, they may be implemented as a standalone device, such as an electronic device including a set-top box. In this case, the electronic device may have at least some of the functions shown in Figures 1 to 4, with the other functions provided elsewhere. Alternatively, the electronic device may have all of these functions. In particular, if there are multiple language models, all of the multiple language models may be provided by the electronic device, or at least some may be provided by the electronic device, with the other language models provided elsewhere. Furthermore, the display of messages in the embodiments described above may be selectable by the user. Similarly, the display of avatar speech may be selectable by the user.

[0203] Furthermore, the display screen described in the embodiments and drawings above is merely an example, and does not limit the displayed items, content, or arrangement. For example, the items displayed on the display screen can be rearranged, or the number of items can be changed. Also, the example display screen only shows a portion of the functions that can actually be implemented for illustrative purposes, and additional necessary information and items may be displayed, or conversely, items may be omitted as needed.

[0204] Furthermore, each functional block or feature of the apparatus used in the embodiments described above may be implemented or executed by an electrical circuit, such as an integrated circuit or a plurality of integrated circuits. An electrical circuit designed to perform the functions described herein may include a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic device, discrete gates or transistor logic, discrete hardware components, or a combination thereof. The general-purpose processor may be a microprocessor, a conventional processor, controller, microcontroller, or state machine. The aforementioned electrical circuit may consist of digital circuits or analog circuits. Also, if advances in semiconductor technology lead to the emergence of integrated circuit technologies that replace current integrated circuits, one or more aspects of this disclosure may use new integrated circuits based on such technologies. [Explanation of Symbols]

[0205] 1 System 10 Display device 100 Control Unit 110 Storage section 112 ROM 114 RAM 116 storage 120 Broadcast Control Unit 130 Operation section 140 Display section 150 Voice input section 160 Audio output section 170 Photography Department 180 Communications Department 20 Server Devices 22 Large-scale language models 30 broadcasting stations 40 Distribution device

Claims

1. An input unit for inputting content including at least images or audio, A recognition unit that recognizes the content of the aforementioned content, A prompt generation unit that generates prompts to be input to the language model, An output data acquisition unit that acquires the output from the language model as output data, An output unit that generates and outputs a response statement from the aforementioned output data, Equipped with, The prompt generation unit, The recognition unit calculates the degree of impact from the content it recognizes, When the impact level exceeds a threshold, the system determines that the impact on the user is significant and generates the prompt containing the content recognized by the recognition unit. Information processing device.

2. The information processing apparatus according to claim 1, wherein the prompt generation unit determines from the content that the content contains new information for the user, and determines that the impact is high.

3. The information processing apparatus according to claim 1, wherein the prompt generation unit determines that the content contains specific words or phrases, if such words or phrases are included in the content, the degree of impact is high.

4. The information processing apparatus according to claim 3, wherein the prompt generation unit determines that the content has a high degree of impact if it contains either a proper noun or a numerical value.

5. The system further includes a content attribute acquisition unit that acquires attributes related to the aforementioned content, The prompt generation unit generates the prompt when the content attribute matches a specific content type. The information processing apparatus according to claim 1.

6. The voice input unit that receives voice input from the user, The system further includes a history storage unit that stores a combination of the input sentence recognized by the recognition unit from the voice input unit and the response sentence as a conversation history. The information processing apparatus according to claim 1, wherein the prompt generation unit determines the degree of impact on the user based on the words and phrases included in the content and the conversation history.

7. The system further includes a user attribute acquisition unit that acquires user attributes from the conversation history, The prompt generation unit, The prompt is generated, including the content of the aforementioned content and the user attributes. The information processing apparatus according to claim 6.

8. A display control unit that controls the display of a display screen having a first display area capable of displaying the aforementioned content and a second display area capable of displaying an avatar corresponding to a character, The information processing apparatus according to claim 1, further comprising an output unit that outputs the response sentence from the avatar.

9. It further includes a setting unit that allows setting a first mode and a second mode, The prompt generation unit, When the setting unit is set to the first mode, the prompt is generated when the influence level is greater than or equal to the threshold. When the setting unit is set to the second mode, the prompt is generated regardless of the degree of influence. The information processing apparatus according to claim 1.

10. An input step that includes inputting content, at least including images or audio, A recognition step of recognizing the content of the aforementioned content, A prompt generation step that generates prompts to be input to the language model, An output data acquisition step which acquires the output from the language model as output data, The output step involves generating and outputting a response statement from the aforementioned output data. Equipped with, The prompt generation step is, The recognition unit calculates the degree of impact from the content it recognizes, When the impact level exceeds a threshold, the system determines that the impact on the user is significant and generates the prompt containing the content recognized by the recognition unit. Processing method.

11. On the computer, An input function that allows input of content including at least images or audio, A recognition function that recognizes the content of the aforementioned content, A prompt generation function that generates prompts to be input to the language model, An output data acquisition function that acquires the output from the language model as output data, The output function generates and outputs a response statement from the aforementioned output data. Equipped with, The aforementioned preset generation function is The recognition unit calculates the degree of impact from the content it recognizes, When the impact level exceeds a threshold, the system determines that the impact on the user is significant and generates the prompt containing the content recognized by the recognition unit. program.

Citation Information

Patent Citations

  • Information processing system and information processing method

    JP7281241B1