System

The system addresses the limitations of traditional news distribution by providing visual and auditory content with real-time interaction, allowing users to engage more deeply with news through avatars.

JP2026014847APending Publication Date: 2026-01-29SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024116321
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-19
Publication Date
2026-01-29

AI Technical Summary

Technical Problem

Traditional news distribution methods lack visual and auditory elements, making it difficult for users to intuitively understand news content and maintain interest, and they fail to provide real-time interactive features.

Method used

A system that acquires data from external sources, selects relevant information, generates audio and video, distributes it, processes real-time comments and questions, and uses avatars for intuitive information delivery.

Benefits of technology

Enables users to engage with news visually and audibly, enhancing understanding and interest through real-time interaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026014847000001_ABST
    Figure 2026014847000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system comprising: means for obtaining data from an external source; means for storing the obtained data; means for filtering information from the stored data based on certain parameters; means for generating the filtered information as audio and video; means for delivering the generated audio and video; and means for processing comments and questions received in real-time during delivery and generating responses.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] In today's information society, many people require access to fast and accurate information. However, traditional news distribution methods rely solely on text information and lack visual and auditory information, making it difficult for users to intuitively understand the news content. Furthermore, the lack of real-time interactive elements makes it difficult to increase users' interest and understanding of the news. It is necessary to solve these issues and provide an environment where users can gain a deeper understanding of the news and actively participate in it. [Means for solving the problem]

[0005] The present invention provides a system including a means for acquiring data from external information sources and storing the acquired data, a means for selecting information from the stored data based on specific parameters, a means for generating the selected information as audio and video, a means for distributing the generated audio and video, and a means for processing comments and questions received in real time during distribution and generating responses. This system allows users to easily receive information visually and audibly, deepening their understanding of the news, and adding real-time interactive elements to increase interest in and understanding of the news. Furthermore, visual display using avatars allows for the provision of information in a user-friendly and intuitive manner.

[0006] "External information sources" are information services such as databases and APIs that provide data and information that exists outside the system.

[0007] A "means for acquiring data" is a software module or hardware device that collects the required data from an external source.

[0008] The "means for storing data" refers to a storage device such as a database or file system for storing acquired data temporarily or long-term.

[0009] "Information filtering means" refers to algorithms or filtering techniques that extract required information from stored data based on specific parameters.

[0010] The "means for generating audio and video" refers to a voice synthesis engine or video generation engine for converting text data into audio and video.

[0011] The "means of delivery" refers to the streaming server and delivery protocol used to deliver the generated audio and video content to users.

[0012] "Means for processing comments and questions and generating responses" refers to chatbots or natural language processing engines that analyze comments and questions submitted by users in real time and generate appropriate responses to them.

[0013] A "visual avatar" is a virtual character used to visually convey information to a user.

[0014] "Means for providing an interface" refers to a user interface that includes input forms and buttons that allow a user to interact with the system. [Brief explanation of the drawings]

[0015] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11]FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0016] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0017] First, the terms used in the following description will be explained.

[0018] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0019] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0020] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0021] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0022] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0023] [First embodiment]

[0024] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0025] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0026] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0027] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0028] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0029] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0030] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0031] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0032] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0033] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0034] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0035] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0036] To implement the present invention, a system including a plurality of components is required, which will be described below with specific examples.

[0037] Data Acquisition Phase

[0038] The server periodically sends requests to external information sources (such as the API of a news service) to obtain the latest news data. The server securely sends the requests using API keys and authentication information. The obtained news data is often received in JSON format or similar.

[0039] Data storage phase

[0040] The server stores the acquired news data in a database. The news data includes information such as the title, text, number of clicks, and publication date and time. This allows for efficient management of data required for subsequent processing.

[0041] News selection phase

[0042] The server sifts through the news data from the database, selecting the latest and most interesting news articles based on certain parameters such as click count, publication date, etc. The sift data is temporarily stored in memory for further processing.

[0043] Avatar generation phase

[0044] The server generates a text script for reading based on the selected news data. The generated script is sent to the avatar engine, which then generates a video of an avatar reading the script. The avatar engine uses speech synthesis and animation technologies to generate natural-looking movements and voices.

[0045] Video distribution phase

[0046] The server uploads the generated avatar video to the streaming server, encodes it, and prepares it for distribution. Once preparation is complete, a stream URL is generated that users can access.

[0047] Interactive Phase

[0048] The terminal provides the user with a user interface, which includes a video player, a comment input field, a question button, etc. When the user uses this interface to input comments or questions in real time, they are sent to the server.

[0049] The server receives and processes comments and questions from users in real time, sends the processed information to the avatar engine, and the avatar generates an appropriate response, which is then re-integrated into the video stream and delivered to the user.

[0050] Specific examples

[0051] For example, suppose a news site experiences a sudden spike in clicks on a particular article. The server periodically retrieves data from the news site and detects this popular article. Next, it generates a text script from the article's title and text, and an avatar reads it aloud using audio and video. The generated video is delivered to the user in real time, and the user can enter "I'd like to know more about the background to this news story" in the comment input field. The server receives the comment and sends it to the avatar engine, where the avatar generates and delivers a response such as "The background to this news story is..."

[0052] This system allows users to receive news visually and audibly, allowing them to enjoy real-time interactive communication.

[0053] The processing flow will be explained below.

[0054] Step 1:

[0055] The server periodically sends requests to an external source (such as a news service API) to obtain the latest news data. The request is sent securely using an API key or authentication information, and the news data is received in JSON format or other formats.

[0056] Step 2:

[0057] The server stores the acquired news data in a database. The news data includes information such as the title, text, number of clicks, and publication date and time, and is stored in the database using the INSERT statement.

[0058] Step 3:

[0059] The server sifts through the news data from the database, sorts the news articles by most clicked based on certain parameters (e.g., number of clicks, publication date), and selects the most recent and likely most interesting news articles.

[0060] Step 4:

[0061] The server generates a text script for reading out the selected news data, formats it into a format such as "Title: ●●, Body: ●●", and sends the generated text script to the avatar engine.

[0062] Step 5:

[0063] The server receives the avatar video generated by the avatar engine and uploads it to the streaming server, where it encodes the video file and prepares it for distribution.

[0064] Step 6:

[0065] The device provides the user with a user interface that includes a video player, a comment input field, a question button, etc. The user uses this interface to watch videos and input comments and questions in real time.

[0066] Step 7:

[0067] The server receives comments and questions sent from the device in real time and processes them. It receives comments using real-time communication technologies such as WebSocket and adds them to a queue.

[0068] Step 8:

[0069] The server takes each comment from the queue and sends it to the avatar engine, which then generates an appropriate response, which is generated as audio and video and added to the live stream.

[0070] Step 9:

[0071] The user receives real-time responses from the avatar via a live stream, providing a visual and auditory experience, and can enter further comments or questions as needed.

[0072] Example 1

[0073] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0074] There is a need to efficiently manage data obtained from external information sources and quickly select important information based on specific parameters. It is also important to deliver this information in a way that users can perceive visually and audibly in a natural way. Furthermore, there is a need for a system that can respond immediately to comments and questions received in real time during the broadcast, enabling interactive communication. However, current systems have difficulty fully meeting these requirements, and technical challenges remain, particularly in real-time processing and audio-video integration.

[0075] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0076] In this invention, the server includes: means for acquiring data from external information sources; means for saving the acquired data; means for selecting information from the saved data based on specific parameters; means for generating audio and video from the selected information; means for distributing the generated audio and video; means for processing comments and questions received in real time during distribution and generating responses; means for selecting news data from data stored in a database and generating videos read by avatars using speech synthesis and animation technologies; means for encoding and distributing avatar videos uploaded to a streaming service; and means for processing user comments and questions on the encoded videos in real time and generating and redistributing avatar responses. This allows users to naturally receive important information visually and audibly, enabling further interactive communication in real time.

[0077] "External information sources" are resources for obtaining data provided by the Internet or specific service providers.

[0078] "Data acquisition means" are the methods and techniques used to collect information from external sources.

[0079] "Means for storing data" refers to the methods and techniques used to hold acquired information in a database or other storage device.

[0080] "Specific parameters" are numerical values ​​or conditions that serve as criteria when selecting or processing data.

[0081] "Means for selecting information" are methods and techniques for extracting necessary information from stored data.

[0082] "Means for generating audio and video" refers to methods and technologies for synthesizing audio and producing video based on selected information.

[0083] "Means of distribution" refers to the methods and technologies used to transmit the generated audio and video to users over the Internet.

[0084] "Means for processing comments and questions received in real time and generating responses" refers to methods and technologies for receiving immediate feedback from users during a broadcast and generating responses based on that feedback.

[0085] A "database" is a system or software for efficiently managing and manipulating large amounts of data.

[0086] "News data" refers to the latest information obtained from news services.

[0087] "Speech synthesis technology" is a technology for artificially generating speech based on text data.

[0088] "Animation technology" is a technique that makes still images appear to be moving by displaying them in succession.

[0089] An "avatar" is a computer-generated virtual representation of a person or character.

[0090] A "streaming service" is a service for delivering digital content in real time over the Internet.

[0091] "Encoding" is the process of converting digital data into a particular format.

[0092] A "user interface" refers to the screen and operation method that allows a user to interact with a system.

[0093] MODE FOR CARRYING OUT THE INVENTION

[0094] To implement this invention, a system including multiple components is required. This system mainly consists of three elements: a server, a terminal, and a user.

[0095] First, the server periodically sends requests to an external information source (for example, the API of a news service) to obtain the latest news data. This request uses HTTPS communication, and an API key or OAuth token is used to ensure security. The obtained news data is typically returned in JSON format. An example request is "GET / latest-news?apiKey=YOUR_API_KEY".

[0096] Next, the server stores the retrieved news data in a database (e.g., PostgreSQL). The database contains the title, content, click count, and publication date / time of each news article. The server efficiently organizes this information and stores it in a structured format. For example, the following query is used: "INSERT INTO news_articles (title, content, clicks, published_at) VALUES ('Example News', 'Example News Content', 120, '2023-10-05 14:48:00');"

[0097] The server selects news data from the database based on certain parameters, such as the articles with the most clicks or the most recent publication date. This selection is performed using an SQL query, such as "SELECT FROM news_articles ORDER BY clicks DESC LIMIT 1;".

[0098] Based on the selected news data, the server generates a text script for reading aloud. This is then used with speech synthesis and animation technology to generate a video in which an avatar reads the script aloud. An avatar engine (e.g., Unity or Unreal Engine) is used for this process. For example, the generated script might read, "We will report the next news item. The title is 'Example News'."

[0099] The generated avatar video is uploaded to a streaming service (e.g., Wowza Streaming Engine) by the server. The server encodes the video and prepares it for distribution. Once preparation for distribution is complete, a stream URL is generated that users can access. For example, the server sends a request called "POST / upload" and receives the stream URL as a response.

[0100] The device provides the user with a user interface (UI). This UI includes a video player, a comment input field, a question button, and so on. The user uses these interfaces to input comments and questions in real time. For example, the user might input a comment such as "I'd like to know more about the background of this news," and the device sends the comment to the server. A request called "POST / comments" is sent, and the comment content is included in the request.

[0101] The server receives comments and questions from users in real time and processes them. The processed information is sent back to the avatar engine, where the avatar generates an appropriate response. The generated response is incorporated into the video stream and delivered to the user. For example, the server passes a script such as "Some background on this news..." to the avatar engine to generate a new video.

[0102] As an example of a prompt sentence, the text "Tell me about the latest news" is input to the generative AI model. Based on this prompt sentence, the entire system operates, carrying out a series of processes from news information collection to distribution and interactive response.

[0103] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0104] Step 1:

[0105] The server periodically sends requests to external news sources to retrieve the latest news data. The input is the API endpoint and authentication information (e.g., API key) of the external news source. The server sends the request "GET / latest-news?apiKey=YOUR_API_KEY" and receives the news data in JSON format. This data includes the title, body, number of clicks, publication date, etc. of the news article.

[0106] Step 2:

[0107] The server saves the acquired news data in a database. The input is news data in JSON format. The server organizes this data and stores it in a database (e.g., PostgreSQL). Specifically, it executes the query "INSERT INTO news_articles (title, content, clicks, published_at) VALUES ('Example news', 'Example news content', 120, '2023-10-05 14:48:00');". The output is the news data saved in the database.

[0108] Step 3:

[0109] The server selects news data from a database based on specific parameters. The input is the news data stored in the database. The server selects the most relevant news articles based on parameters such as the number of clicks and publication date, and temporarily stores them in memory. For example, it executes the query "SELECT FROM news_articles ORDER BY clicks DESC LIMIT 1;". The output is the selected news data.

[0110] Step 4:

[0111] The server generates a text script for reading aloud based on the selected news data. The input is the selected news data. The server extracts the title and content from the news data and generates a text script that says, "We will report the next news item. The title is 'Example News'." The output is a text script for reading aloud.

[0112] Step 5:

[0113] The server sends this text script to the avatar engine, which uses speech synthesis and animation technologies to generate a video in which an avatar reads the script. The input is the text script. The avatar engine (e.g., Unity or Unreal Engine) generates a video based on the passed script. The output is the generated avatar video.

[0114] Step 6:

[0115] The server uploads the generated avatar video to the streaming service, encodes it, and prepares it for distribution. The input is the generated avatar video. Specifically, it sends a "POST / upload" request to the streaming service (e.g., Wowza Streaming Engine). The output is the stream URL that users can access.

[0116] Step 7:

[0117] The device provides a user interface (UI) to the user. The input is the stream URL. This UI includes a video player, a comment input field, a question button, and so on. The user inputs comments and questions in real time through this interface. Specifically, the user inputs "I'd like to know more about the background of this news," and the device sends a "POST / comments" request to the server. The output is the user's comments and questions.

[0118] Step 8:

[0119] The server receives and processes comments and questions from users in real time. The input is the user's comment or question. The server sends the received comment to the avatar engine, and the avatar generates an appropriate response. For example, a script such as "About the background of this news..." is passed to the avatar engine to generate a new video. The output is the generated avatar's response video.

[0120] Step 9:

[0121] The server uploads the generated response video back to the streaming service and delivers it to the user. The input is the avatar's response video. The server uploads new videos to the streaming service, updating the delivery in real time. The output is the updated video stream.

[0122] (Application example 1)

[0123] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0124] In recent years, with the digitalization of information and the spread of the Internet, methods of news distribution have become more diverse. However, systems that allow users to obtain information interactively in real time are insufficient, and there is a need for more information to be provided visually and audibly. In addition to allowing users to efficiently browse a wide range of news, there is also a need for systems that can provide information based on the user's interests.

[0125] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0126] In this invention, the server includes means for acquiring data from external information sources, means for storing the acquired data, means for selecting information from the stored data based on specific parameters, means for generating the selected information as audio and video, means for distributing the generated audio and video, means for processing comments and questions received in real time during distribution and generating responses, and means for providing the data on a smartphone application, thereby enabling users to watch the news interactively in real time and receive information visually and audibly.

[0127] "External information sources" refer to information sources that exist outside the system, such as online news services or APIs.

[0128] "Means for Obtaining Data" refers to methods or techniques for periodically requesting and receiving data from external sources.

[0129] "Means for storing data" refers to the methods or techniques for storing acquired data in a database or other storage medium.

[0130] "Information screening means" refers to methods or technologies for extracting selected information from stored data based on specific criteria (e.g., number of clicks or publication date).

[0131] "Means for generating audio and video" refers to technologies for creating audio and video from selected information, in particular methods for using speech synthesis and animation technologies to read aloud using an avatar.

[0132] "Delivery Means" means the method or technology by which the generated audio and video is made available to users in streaming format.

[0133] "Means for processing comments and questions and generating responses" refers to a method or technology for receiving comments and questions sent in real time by users during a broadcast and generating and returning appropriate responses to them.

[0134] "Smartphone Application" means software that runs on a smartphone device and that provides users with the ability to interactively browse news.

[0135] "Real-time comments and questions" refers to comments and questions submitted by users during a video broadcast, and refers to information processed to respond to such comments and questions in a timely manner.

[0136] "Means for processing comments and questions received in real time during the broadcast and generating responses" refers to a method for instantly analyzing feedback from users during a live broadcast and generating appropriate responses.

[0137]

[0138] To implement the present invention, a system including a plurality of components is required, which will be described below with specific examples.

[0139] Data Acquisition Phase

[0140] The server periodically sends requests to external information sources (such as the API of a news service) to obtain the latest news data. The server securely sends the requests using API keys and authentication information. The obtained news data is often received in JSON format or similar.

[0141] Data storage phase

[0142] The server stores the acquired news data in a database. The news data includes information such as the title, text, number of clicks, and publication date and time. This allows for efficient management of data required for subsequent processing.

[0143] News selection phase

[0144] The server sifts through the news data from the database, selecting the latest and most interesting news articles based on certain parameters such as click count, publication date, etc. The sift data is temporarily stored in memory for further processing.

[0145] Avatar generation phase

[0146] The server generates a text script for reading based on the selected news data. The generated script is sent to the avatar engine, which then generates a video of an avatar reading the script. The avatar engine uses speech synthesis and animation technologies to generate natural-looking movements and voices.

[0147] Video distribution phase

[0148] The server uploads the generated avatar video to the streaming server, encodes it, and prepares it for distribution. Once preparation is complete, a stream URL is generated that users can access. Users can watch the video through this URL on their smartphone application.

[0149] Interactive Phase

[0150] Users can enter comments and questions in real time using the smartphone application, including the video player, comment input field, and question button. The server receives and processes comments and questions from users in real time. The processed information is sent to the avatar engine, where the avatar generates an appropriate response. The generated response is then incorporated back into the video stream and delivered to the user.

[0151] For example, suppose a news site experiences a sudden spike in clicks on a particular article. The server periodically retrieves data from the news site and detects this popular article. Next, it generates a text script from the article's title and text, and an avatar reads it aloud using audio and video. The generated video is then delivered to the user in real time. The user can then enter "I'd like to know more about the background to this news story" in the comment input field. The server receives the comment, sends it to the avatar engine, and the avatar generates and delivers a response such as "The background to this news story is..."

[0152] Example prompt sentence:

[0153] "News title: {title}

[0154] News text: {text}

[0155] Please use this as a basis to generate a video of the avatar reading the text."

[0156] This system allows users to receive news visually and audibly, allowing them to enjoy real-time interactive communication.

[0157] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0158] Step 1: Data Acquisition Phase

[0159] The server sends an HTTP request to the news service's API to obtain the latest news data. The request header, including the API key and authentication information, is used as input. The response from the API receives news data in JSON format. The data includes the article title, text, number of clicks, publication date, and so on.

[0160] Step 2: Data storage phase

[0161] The server stores the retrieved news data in a local database. As input, it uses each field of the received news data (title, body, number of clicks, publication date, etc.). As output, each news article is stored in the database. Specifically, it adds data to the database using the SQL INSERT statement.

[0162] Step 3: News selection phase

[0163] The server selects news data from the database. As input, it uses the news data stored in the database. The selection criteria are specific parameters such as the number of clicks or publication date. As output, it selects the most interesting news articles. Specifically, it uses a SQL SELECT statement to filter the articles that match the criteria.

[0164] Step 4: Avatar generation phase

[0165] The server generates a text script for reading aloud based on the selected news data. The selected news data (title, body text) is used as input. The text script is generated as output. The generated script is then sent to the avatar engine, which generates audio and video for the avatar to read aloud. Specifically, the server calls the avatar engine's API to convert the text into audio and animation.

[0166] Step 5: Video distribution phase

[0167] The server uploads the generated avatar video to the streaming server, encodes it, and prepares it for distribution. The generated avatar video is used as input. As output, a stream URL that can be accessed by users is generated. Specifically, the server uses a file transfer protocol to upload the video to the streaming server and generate the URL.

[0168] Step 6: User Interaction Phase

[0169] Users enter comments and questions in real time via a smartphone application using a video player, comment input field, question button, etc. The comments and questions entered by the user on the application are used as input. The input content is sent to the server as output. Specifically, the system operates by using a UI component that accepts user input and a network function that sends it to the server.

[0170] Step 7: Comment and question handling phase

[0171] The server receives and processes comments and questions submitted by users in real time. It uses the comments and questions submitted by users as input. As output, it generates an appropriate response. Specifically, it analyzes the content of the comments and questions and sends them to the avatar engine to generate a response.

[0172] Step 8: Response Delivery Phase

[0173] The server then incorporates the generated response back into the video stream and delivers it to the user. It uses the generated response as input, and delivers an updated video stream to the user as output. Specifically, it uses the streaming server's API to incorporate the response into the video and delivers the updated stream.

[0174] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0175] To implement this invention, a system including a number of components is required, which will be described below with specific examples.

[0176] Data Acquisition Phase

[0177] The server periodically sends requests to external sources (such as the API of a news service) to obtain the latest news data. The requests are sent securely using API keys and authentication information, and the news data is received in JSON format or other formats.

[0178] Data storage phase

[0179] The server stores the acquired news data in a database. The news data includes information such as the title, text, number of clicks, and publication date and time. This allows for efficient management of data required for subsequent processing.

[0180] News selection phase

[0181] The server selects news data from the database, sorts the news articles by the number of clicks based on certain parameters (such as the number of clicks or the publication date), and selects the latest and most interesting news articles. The selected data is temporarily stored in memory for further processing.

[0182] Avatar generation phase

[0183] The server generates a text script for reading based on the selected news data. The generated script is sent to the avatar engine, which then generates a video of an avatar reading the script. The avatar engine uses speech synthesis and animation technologies to generate natural-looking movements and voices.

[0184] Video distribution phase

[0185] The server uploads the generated avatar video to the streaming server, encodes it, and prepares it for distribution. Once preparation is complete, a stream URL is generated that users can access.

[0186] Interactive Phase

[0187] The device provides the user with a user interface that includes a video player, a comment input field, a question button, etc. The user uses this interface to watch videos and input comments and questions in real time.

[0188] The server receives comments and questions sent from the device in real time and processes them. It receives comments using real-time communication technologies such as WebSocket and adds them to a queue.

[0189] The server takes each comment from the queue and sends it to the avatar engine, which then generates an appropriate response, which is generated as audio and video and added to the live stream.

[0190] Emotion Recognition Phase

[0191] The server uses an emotion engine to analyze comments and questions entered by users. The emotion engine recognizes the user's emotions (e.g., joy, sadness, anger, etc.) from the text contained in the comments and questions.

[0192] The server adjusts the response of the generated avatar appropriately based on the user's emotion recognized by the emotion engine. For example, if the user expresses anger, the avatar's response will be calm.

[0193] Display and Feedback Phase

[0194] The terminal provides a method for visually displaying the user's emotional information recognized by the emotion engine, allowing the user to confirm that their own emotional state is reflected.

[0195] Specific examples

[0196] For example, suppose a news site experiences a sudden spike in clicks on a particular article. The server periodically retrieves data from the news site to identify popular articles. It then generates a text script from the article's title and text, and an avatar reads it aloud with audio and video. The generated video is then streamed to the user in real time, who can then type in a comment field, "I'd like to know more about the background to this news story."

[0197] The server receives the comment and recognizes the "interest" in the emotion engine. The avatar's response generated by the avatar engine provides detailed information in a way that attracts the user's interest.

[0198] This system not only allows users to receive news visually and audibly, but also allows for real-time interactive communication, and further enhances the user experience by recognizing users' emotions and providing responses accordingly.

[0199] The processing flow will be explained below.

[0200] Step 1:

[0201] The server periodically sends requests to an external source (such as a news service API) to retrieve the latest news data. The request is sent securely using an API key or authentication information, and the news data is received in JSON format.

[0202] Step 2:

[0203] The server stores the acquired news data in a database. The news data includes information such as the title, text, number of clicks, and publication date and time, and is stored in the database using the INSERT statement.

[0204] Step 3:

[0205] The server sorts the news data from the database, executes a SELECT statement based on specific parameters (number of clicks and publication date), and sorts the news articles by most clicked.

[0206] Step 4:

[0207] The server generates a text script for reading out the selected news data, formats it as "Title: ●●, Body: ●●", and sends the generated text script to the avatar engine.

[0208] Step 5:

[0209] The server receives the avatar video generated by the avatar engine, uploads it to the streaming server, encodes the video file, and generates a stream URL that users can access.

[0210] Step 6:

[0211] The terminal provides a user interface, including a video player, a comment input field, and a question button, which users use to watch videos and input comments and questions.

[0212] Step 7:

[0213] The server receives comments and questions sent from the device in real time and adds the comments to a queue using real-time communication technologies such as WebSocket.

[0214] Step 8:

[0215] The server takes each comment from the queue and sends it to the avatar engine, which then generates an appropriate response, which is then generated as audio and video and added to the live stream.

[0216] Step 9:

[0217] The server recognizes emotions from comments and questions submitted by users using an emotion engine, which extracts the user's emotions (e.g., joy, sadness, anger) from the text.

[0218] Step 10:

[0219] The server adjusts the response generated by the avatar engine based on the user's emotion recognized by the emotion engine, for example, if the user expresses anger, the avatar's response will be in a calm tone.

[0220] Step 11:

[0221] The device visually displays the user's emotional information recognized by the emotion engine, allowing the user to confirm that their own emotional state is reflected.

[0222] Step 12:

[0223] The user receives real-time responses from the avatar and can continue the interactive communication by entering further comments or questions, a process that is repeated to improve the user experience.

[0224] Example 2

[0225] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0226] Conventional news delivery systems are limited to viewing news and have limited user interaction. Furthermore, they do not provide responses that take into account the user's emotions, resulting in a uniform user experience and low satisfaction for some users. Furthermore, they often lack interactivity because responses to real-time comments and questions are not always prompt.

[0227] The specific processing by the specific processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for acquiring data from an external information source, means for saving the acquired data, means for selecting information from the saved data based on specific parameters, means for generating audio and video from the selected information, means for distributing the generated audio and video, means for processing comments and questions received in real time during distribution and generating responses, and means for performing sentiment analysis on the received comments and questions and adjusting responses based on the results. This increases interactivity in news distribution and enables responses that correspond to the user's emotions. It also enhances real-time interaction with users and improves user satisfaction.

[0228] "External sources" refers to sources for obtaining data from outside the system, such as news service APIs and websites.

[0229] "Means for obtaining data" refers to a combination of programming and hardware for sending requests to external sources and receiving the required data.

[0230] "Means for storing data" refers to a database management system or storage device for efficiently storing acquired data.

[0231] "Means for filtering information from data based on specific parameters" refers to programs that filter stored data using criteria such as the number of clicks or publication date and time to extract the desired information.

[0232] "Audio and video generation means" means a program that converts selected information into a text script and uses speech synthesis and animation techniques to make it visually and audibly reproducible.

[0233] "Means for delivering audio and video" refers to the streaming server and encoding technology used to deliver the generated audio and video to users.

[0234] "Means for processing comments and questions received in real time and generating responses" refers to programs and communication technologies for collecting and analyzing user input in real time and generating appropriate responses based on the results.

[0235] "Means for analyzing emotions and adjusting responses based on the results" refers to a program that analyzes the emotions contained in comments and questions from users and generates an appropriate response based on those emotions.

[0236] To implement the present invention, several hardware and software components are required, which will be described below with specific examples.

[0237] Data Acquisition Phase

[0238] The server is responsible for obtaining data from external sources. Specifically, it periodically sends requests to the news service's API to obtain the latest news data. This is done by securely sending requests using API keys and authentication information, and receiving news data in JSON format or similar. For example, the server uses the NewsAPI to obtain the latest 50 articles at a time.

[0239] Data storage phase

[0240] The server stores the acquired news data in a database system (e.g., MySQL, PostgreSQL). News data includes information such as the title, text, number of clicks, and publication date and time. The server inserts this data into a database to efficiently manage it. For example, the news title "Breaking News," its text, number of clicks "325," and publication date and time "2023-10-01" are stored in the database.

[0241] News selection phase

[0242] When filtering news data from the database, the server selects news articles based on certain parameters, sorts the news articles based on the number of clicks or publication date, or sorts the articles by most clicked or most recently published, and stores the filtered data in memory. For example, it selects the top 10 articles with the most clicks.

[0243] Avatar generation phase

[0244] The server generates a text script based on the selected news data and sends the script to an avatar engine (e.g., FaceRig, Voki). The avatar engine uses speech synthesis technology (e.g., Google Text-to-Speech API) and animation technology to generate avatar videos with natural voices and movements. For example, it generates a video of an avatar reading a news article based on the text script.

[0245] Video distribution phase

[0246] The server uploads the generated avatar video to a streaming server (e.g., Wowza Streaming Engine), encodes it, and prepares it for distribution. Once preparation for distribution is complete, a stream URL is generated that users can access. For example, "https: / / streamingserver.com / newsXstream" is provided to users.

[0247] Interactive Phase

[0248] The device provides the user with a user interface that includes a video player, a comment input field, and a question button. The user uses these interfaces to watch videos and input comments and questions in real time. The server receives comments and questions sent from the device using real-time communication technology such as WebSocket and adds them to a queue. For example, a user might input a comment such as, "I'd like to know more about the background of this news."

[0249] Emotion Recognition Phase

[0250] The server uses an emotion engine (e.g., IBM Watson Tone Analyzer) to analyze the user's comments and questions and recognize the user's emotions. Based on this analysis, the server adjusts the tone and content of the response appropriately. For example, if the user's comments are analyzed as indicating "interest," the avatar's response will be adjusted to be more detailed and interesting.

[0251] Display and Feedback Steps

[0252] The device provides an interface that visually displays the user's emotional information recognized by the emotion engine, allowing the user to confirm that their emotional state is accurately reflected. For example, an icon displayed next to the comment field represents the user's emotion.

[0253] Examples:

[0254] For example, if a news site experiences a sudden spike in clicks on a particular article, the server periodically retrieves data from that news site to identify popular articles. Next, a text script is generated from the article's title and text, and an avatar reads it aloud using audio and video. The generated video is then streamed to the user in real time, who can then type in a comment field, "I'd like to know more about the background to this news story."

[0255] The server receives the comments and recognizes "interest" using the emotion engine. The avatar engine generates responses from avatars, providing detailed information in a way that piques the user's interest. This system allows users to not only receive news visually and audibly, but also enjoy real-time interactive communication.

[0256] Example prompt sentence:

[0257] "Generate text scripts from the titles and text of breaking news articles and create videos of avatars reading them. Also, generate responses to user comments and respond in the right tone based on sentiment analysis."

[0258] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0259] Step 1: Data Acquisition Step

[0260] The server retrieves data from external information sources. Specifically, it periodically sends requests to the news service's API. The input is a request that includes the API key and authentication information. The server sends the request, and the data it retrieves is news data in JSON format. Specifically, the server accesses the News API and sends a request saying, "Please retrieve the latest 50 news articles." In response, it receives the JSON data of the news articles.

[0261] Step 2: Data storage step

[0262] The server saves the acquired news data in a database. The input is the JSON-formatted news data acquired in step 1. The server parses the JSON data, extracts fields such as the news article title, text, number of clicks, and publication date and time, and inserts them into a database (e.g., MySQL, PostgreSQL). The output is the news data saved in the database. Specifically, the news title "Breaking News", text "Detailed Content", number of clicks "325", and publication date and time "2023-10-01" are saved in the database.

[0263] Step 3: News selection step

[0264] The server selects news data from the database. The input is the news data stored in the database. The server sorts the news articles based on certain parameters (e.g., number of clicks or publication date) and selects the latest and most clicked news articles. The output is the selected news data. Specifically, it selects the top 10 articles with the most clicks. The data is temporarily stored in memory.

[0265] Step 4: Avatar generation step

[0266] The server generates a text script based on the selected news data. The input is the news data selected in step 3. The server combines the news title and text to generate a text script and sends it to the avatar engine. The output is an avatar video that reads the text script aloud. Specifically, it generates a text script that reads the news title "Breaking News" and its text aloud and sends it to the avatar engine. The avatar engine generates the video using speech synthesis and animation technologies.

[0267] Step 5: Video distribution step

[0268] The server uploads the generated avatar video to the streaming server. The input is the avatar video generated in step 4. The server uploads the video to the streaming server (e.g., Wowza Streaming Engine), encodes it, and prepares it for distribution. The output is a stream URL that users can access. Specifically, the generated avatar video is uploaded to "https: / / streamingserver.com / newsXstream" so that users can access it.

[0269] Step 6: Interactive Step

[0270] The terminal provides an interface to the user. The input is the stream URL from the streaming server. The terminal displays a video player, a comment input field, a question button, etc. to the user. The user uses these interfaces to watch the video and input comments and questions in real time. The output is comments and questions from the user. Specifically, the user inputs a comment such as "I'd like to know more about the background of this news."

[0271] Step 7: Emotion Recognition Step

[0272] The server analyzes the user's comments and questions using an emotion engine. The input is the comment or question received from the user in step 6. The server uses an emotion engine (e.g., IBM Watson Tone Analyzer) to recognize the user's emotions (e.g., interest, joy, sadness, anger). The output is the emotion analysis result and an adjusted response. Specifically, if the user's comment is recognized as indicating "interest," the server adjusts the avatar's response to be more detailed and interesting.

[0273] Step 8: Display and Feedback Step

[0274] The device visually displays the user's emotional information. The input is the emotion analysis result obtained in step 7. The device displays an icon that expresses the user's emotion next to the comment field. The output is the visual display of the user's emotional information. Specifically, emotion icons (e.g., smiley, tearful, angry expressions) are displayed in the comment field, allowing the user to confirm that their emotions are accurately reflected.

[0275] (Application example 2)

[0276] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0277] In conventional news delivery systems, news acquisition, selection, and delivery are one-way, limiting user interaction. Furthermore, it is difficult to generate appropriate responses based on user emotions, making it difficult to improve the user experience. In particular, the lack of real-time emotion recognition and its reflection can lead to a decline in user engagement.

[0278] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for acquiring data from an external information source, means for saving the acquired data, means for selecting information from the saved data based on specific parameters, means for generating the selected information as audio and video, means for distributing the generated audio and video, means for processing comments and questions received in real time during distribution and generating a response, means for recognizing the user's emotions from the comments and questions and adjusting the response, and means for visually displaying the results of the emotion recognition. This enables interactive communication in real time while watching the news, and makes it possible to provide an appropriate response according to the user's emotions.

[0279] "External sources" are data sources such as news services, APIs, and data feeds available on the Internet.

[0280] "Data Acquisition Methods" are the systems and processes used to periodically request and acquire required information from external sources.

[0281] "Means for storing data" refers to a database or storage device for temporarily or permanently storing acquired information.

[0282] "Information filtering means" refers to algorithms or programs that extract important data from stored information based on specific conditions.

[0283] The "means for generating audio and video" is a system that creates visual and auditory content based on text information using speech synthesis technology and video generation technology.

[0284] The "means for delivering audio and video" refers to a server or network configuration for streaming the generated media content to users.

[0285] A "means for processing comments and questions received in real time and generating responses" is a program or process that receives input from users in real time and generates appropriate responses based on that input.

[0286] "Means for recognizing user emotions from comments and questions and adjusting responses" refers to emotion recognition engines and algorithms that analyze emotions from input text and adjust response content based on the results.

[0287] The "means for visually displaying the results of emotion recognition" refers to a function or program for visually displaying the analyzed emotion information on a user interface.

[0288] To implement this invention, a system including a server, a user terminal, a generative AI model, etc. Specifically, the system is realized by the following steps.

[0289] First, the server periodically retrieves data from external sources, using API keys and authentication information to securely retrieve data, such as from a news service API. The retrieved data is received in JSON format.

[0290] The acquired data is then stored in a database on the server, which stores information such as the news article title, text, number of clicks, and publication date and time, allowing for efficient management of the data required for subsequent processing.

[0291] The server then sorts through the vast amount of news articles stored in the database, using criteria such as click counts and publication dates to select the most interesting news articles.

[0292] Based on the selected news data, the server generates a text script for reading. This script is sent to an avatar engine, which uses speech synthesis and animation technologies to create an avatar video that generates natural-looking movements and voices. The avatar engine can be Amazon Polly or Google Text-to-Speech.

[0293] The generated avatar video is uploaded to a streaming server, where it is encoded and prepared for distribution. Once preparation is complete, a stream URL is generated that users can access.

[0294] The user terminal provides a user interface including a video player, a comment input field, a question button, etc. The user can use this interface to watch the video and input comments and questions in real time.

[0295] The server receives comments and questions sent from the device in real time and processes them using WebSocket technology, etc. Comments and questions are added to a queue, then sequentially retrieved and sent to the avatar engine, where the avatar generates an appropriate response.

[0296] In the emotion recognition phase, the server uses an emotion engine to analyze the comments and questions entered by the user. The emotion engine uses IBM Watson Natural Language Understanding and other technologies. It recognizes emotions from the user's input text and adjusts the response accordingly. For example, if the user expresses anger, the avatar's response will be calmer.

[0297] The results of emotion recognition are visually displayed on the user's device, allowing them to see responses that reflect their own emotional state, enabling more personalized, real-time communication.

[0298] Specific examples

[0299] For example, a "news avatar" reads the latest news article, and a viewer enters a comment such as, "I'd like to know more about the background to this news." The server receives this comment in real time and uses an emotion recognition engine to identify "interest." The avatar engine then generates a response such as, "Let me explain the details of this news story," providing information in a way that piques the viewer's interest. This process allows users to enjoy real-time interactive communication while watching the news broadcast.

[0300] Prompt Sentence Examples

[0301] "A news avatar will read the latest news to viewers, and users can enter comments to instantly respond. Emotion recognition will be performed for each comment, and the avatar will respond appropriately. For example, create a system that responds to a comment such as 'I'd like to know more about the background of this news story,' with 'I'll explain the details of this news story.'"

[0302] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0303] Step 1:

[0304] The server accesses an external information source (the API of a news service) to obtain the latest news data. At this time, it securely requests the data using an API key and authentication information, and receives the news data in JSON format. The input is the news data received from the API, and the output is raw news data that is temporarily stored inside the server.

[0305] Step 2:

[0306] The server stores the acquired news data in a database, which stores information such as the title, text, number of clicks, and publication date and time of news articles. The input is raw news data, and the output is structured data stored in the database. This allows for efficient management of the data required for subsequent processing.

[0307] Step 3:

[0308] The server selects news data from the database based on specific parameters (e.g., number of clicks or publication date). The selected news data is temporarily stored in memory. The input is the news data stored in the database, and the output is the selected news data.

[0309] Step 4:

[0310] The server generates a text script for reading based on the selected news data. This script is sent to the avatar engine, which uses speech synthesis and animation technologies to make the avatar move and speak naturally. The input is the selected news data, and the output is the generated avatar video.

[0311] Step 5:

[0312] The server uploads the generated avatar video to the streaming server, encodes it, and prepares it for distribution. Once it is ready for distribution, a stream URL that users can access is generated. The input is the avatar video, and the output is the stream URL.

[0313] Step 6:

[0314] The user terminal provides a user interface that includes a video player, a comment input field, a question button, etc. The user uses the interface to watch videos and input comments and questions in real time. The input is the user's comments and questions, and the output is the corresponding interface operation.

[0315] Step 7:

[0316] The server receives comments and questions sent from the terminal in real time and adds the comments to a queue using WebSocket technology. The input is the comments and questions received from the user, and the output is the data added to the comment queue.

[0317] Step 8:

[0318] The server sequentially retrieves comments added to the queue and sends them to the avatar engine, where the avatar generates an appropriate response. The input is the comment retrieved from the comment queue, and the output is the generated avatar response.

[0319] Step 9:

[0320] The server analyzes the user's comments and questions using an emotion engine and adjusts the avatar's response based on the emotion. The input is the user's comments and questions, and the output is the adjusted avatar response. The emotion engine uses IBM Watson Natural Language Understanding and other technologies.

[0321] Step 10:

[0322] The user device visually displays the emotion recognition results, allowing the user to confirm the response that reflects their own emotional state. The input is the analysis result of the emotion engine, and the output is the visually displayed emotional information.

[0323] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0324] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0325] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0326] [Second embodiment]

[0327] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0328] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0329] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0330] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0331] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0332] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0333] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0334] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0335] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0336] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0337] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0338] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0339] To implement the present invention, a system including a plurality of components is required, which will be described below with specific examples.

[0340] Data Acquisition Phase

[0341] The server periodically sends requests to external information sources (such as the API of a news service) to obtain the latest news data. The server securely sends the requests using API keys and authentication information. The obtained news data is often received in JSON format or similar.

[0342] Data storage phase

[0343] The server stores the acquired news data in a database. The news data includes information such as the title, text, number of clicks, and publication date and time. This allows for efficient management of data required for subsequent processing.

[0344] News selection phase

[0345] The server sifts through the news data from the database, selecting the latest and most interesting news articles based on certain parameters such as click count, publication date, etc. The sift data is temporarily stored in memory for further processing.

[0346] Avatar generation phase

[0347] The server generates a text script for reading based on the selected news data. The generated script is sent to the avatar engine, which then generates a video of an avatar reading the script. The avatar engine uses speech synthesis and animation technologies to generate natural-looking movements and voices.

[0348] Video distribution phase

[0349] The server uploads the generated avatar video to the streaming server, encodes it, and prepares it for distribution. Once preparation is complete, a stream URL is generated that users can access.

[0350] Interactive Phase

[0351] The terminal provides the user with a user interface, which includes a video player, a comment input field, a question button, etc. When the user uses this interface to input comments or questions in real time, they are sent to the server.

[0352] The server receives and processes comments and questions from users in real time, sends the processed information to the avatar engine, and the avatar generates an appropriate response, which is then re-integrated into the video stream and delivered to the user.

[0353] Specific examples

[0354] For example, suppose a news site experiences a sudden spike in clicks on a particular article. The server periodically retrieves data from the news site and detects this popular article. Next, it generates a text script from the article's title and text, and an avatar reads it aloud using audio and video. The generated video is delivered to the user in real time, and the user can enter "I'd like to know more about the background to this news story" in the comment input field. The server receives the comment and sends it to the avatar engine, where the avatar generates and delivers a response such as "The background to this news story is..."

[0355] This system allows users to receive news visually and audibly, allowing them to enjoy real-time interactive communication.

[0356] The processing flow will be explained below.

[0357] Step 1:

[0358] The server periodically sends requests to an external source (such as a news service API) to obtain the latest news data. The request is sent securely using an API key or authentication information, and the news data is received in JSON format or other formats.

[0359] Step 2:

[0360] The server stores the acquired news data in a database. The news data includes information such as the title, text, number of clicks, and publication date and time, and is stored in the database using the INSERT statement.

[0361] Step 3:

[0362] The server sifts through the news data from the database, sorts the news articles by most clicked based on certain parameters (e.g., number of clicks, publication date), and selects the most recent and likely most interesting news articles.

[0363] Step 4:

[0364] The server generates a text script for reading out the selected news data, formats it into a format such as "Title: ●●, Body: ●●", and sends the generated text script to the avatar engine.

[0365] Step 5:

[0366] The server receives the avatar video generated by the avatar engine and uploads it to the streaming server, where it encodes the video file and prepares it for distribution.

[0367] Step 6:

[0368] The device provides the user with a user interface that includes a video player, a comment input field, a question button, etc. The user uses this interface to watch videos and input comments and questions in real time.

[0369] Step 7:

[0370] The server receives comments and questions sent from the device in real time and processes them. It receives comments using real-time communication technologies such as WebSocket and adds them to a queue.

[0371] Step 8:

[0372] The server takes each comment from the queue and sends it to the avatar engine, which then generates an appropriate response, which is generated as audio and video and added to the live stream.

[0373] Step 9:

[0374] The user receives real-time responses from the avatar via a live stream, providing a visual and auditory experience, and can enter further comments or questions as needed.

[0375] Example 1

[0376] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0377] There is a need to efficiently manage data obtained from external information sources and quickly select important information based on specific parameters. It is also important to deliver this information in a way that users can perceive visually and audibly in a natural way. Furthermore, there is a need for a system that can respond immediately to comments and questions received in real time during the broadcast, enabling interactive communication. However, current systems have difficulty fully meeting these requirements, and technical challenges remain, particularly in real-time processing and audio-video integration.

[0378] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0379] In this invention, the server includes: means for acquiring data from external information sources; means for saving the acquired data; means for selecting information from the saved data based on specific parameters; means for generating audio and video from the selected information; means for distributing the generated audio and video; means for processing comments and questions received in real time during distribution and generating responses; means for selecting news data from data stored in a database and generating videos read by avatars using speech synthesis and animation technologies; means for encoding and distributing avatar videos uploaded to a streaming service; and means for processing user comments and questions on the encoded videos in real time and generating and redistributing avatar responses. This allows users to naturally receive important information visually and audibly, enabling further interactive communication in real time.

[0380] "External information sources" are resources for obtaining data provided by the Internet or specific service providers.

[0381] "Data acquisition means" are the methods and techniques used to collect information from external sources.

[0382] "Means for storing data" refers to the methods and techniques used to hold acquired information in a database or other storage device.

[0383] "Specific parameters" are numerical values ​​or conditions that serve as criteria when selecting or processing data.

[0384] "Means for selecting information" are methods and techniques for extracting necessary information from stored data.

[0385] "Means for generating audio and video" refers to methods and technologies for synthesizing audio and producing video based on selected information.

[0386] "Means of distribution" refers to the methods and technologies used to transmit the generated audio and video to users over the Internet.

[0387] "Means for processing comments and questions received in real time and generating responses" refers to methods and technologies for receiving immediate feedback from users during a broadcast and generating responses based on that feedback.

[0388] A "database" is a system or software for efficiently managing and manipulating large amounts of data.

[0389] "News data" refers to the latest information obtained from news services.

[0390] "Speech synthesis technology" is a technology for artificially generating speech based on text data.

[0391] "Animation technology" is a technique that makes still images appear to be moving by displaying them in succession.

[0392] An "avatar" is a computer-generated virtual representation of a person or character.

[0393] A "streaming service" is a service for delivering digital content in real time over the Internet.

[0394] "Encoding" is the process of converting digital data into a particular format.

[0395] A "user interface" refers to the screen and operation method that allows a user to interact with a system.

[0396] MODE FOR CARRYING OUT THE INVENTION

[0397] To implement this invention, a system including multiple components is required. This system mainly consists of three elements: a server, a terminal, and a user.

[0398] First, the server periodically sends requests to an external information source (for example, the API of a news service) to obtain the latest news data. This request uses HTTPS communication, and an API key or OAuth token is used to ensure security. The obtained news data is typically returned in JSON format. An example request is "GET / latest-news?apiKey=YOUR_API_KEY".

[0399] Next, the server stores the retrieved news data in a database (e.g., PostgreSQL). The database contains the title, content, click count, and publication date / time of each news article. The server efficiently organizes this information and stores it in a structured format. For example, the following query is used: "INSERT INTO news_articles (title, content, clicks, published_at) VALUES ('Example News', 'Example News Content', 120, '2023-10-05 14:48:00');"

[0400] The server selects news data from the database based on certain parameters, such as the articles with the most clicks or the most recent publication date. This selection is performed using an SQL query, such as "SELECT FROM news_articles ORDER BY clicks DESC LIMIT 1;".

[0401] Based on the selected news data, the server generates a text script for reading aloud. This is then used with speech synthesis and animation technology to generate a video in which an avatar reads the script aloud. An avatar engine (e.g., Unity or Unreal Engine) is used for this process. For example, the generated script might read, "We will report the next news item. The title is 'Example News'."

[0402] The generated avatar video is uploaded to a streaming service (e.g., Wowza Streaming Engine) by the server. The server encodes the video and prepares it for distribution. Once preparation for distribution is complete, a stream URL is generated that users can access. For example, the server sends a request called "POST / upload" and receives the stream URL as a response.

[0403] The device provides the user with a user interface (UI). This UI includes a video player, a comment input field, a question button, and so on. The user uses these interfaces to input comments and questions in real time. For example, the user might input a comment such as "I'd like to know more about the background of this news," and the device sends the comment to the server. A request called "POST / comments" is sent, and the comment content is included in the request.

[0404] The server receives comments and questions from users in real time and processes them. The processed information is sent back to the avatar engine, where the avatar generates an appropriate response. The generated response is incorporated into the video stream and delivered to the user. For example, the server passes a script such as "Some background on this news..." to the avatar engine to generate a new video.

[0405] As an example of a prompt sentence, the text "Tell me about the latest news" is input to the generative AI model. Based on this prompt sentence, the entire system operates, carrying out a series of processes from news information collection to distribution and interactive response.

[0406] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0407] Step 1:

[0408] The server periodically sends requests to external news sources to retrieve the latest news data. The input is the API endpoint and authentication information (e.g., API key) of the external news source. The server sends the request "GET / latest-news?apiKey=YOUR_API_KEY" and receives the news data in JSON format. This data includes the title, body, number of clicks, publication date, etc. of the news article.

[0409] Step 2:

[0410] The server saves the acquired news data in a database. The input is news data in JSON format. The server organizes this data and stores it in a database (e.g., PostgreSQL). Specifically, it executes the query "INSERT INTO news_articles (title, content, clicks, published_at) VALUES ('Example news', 'Example news content', 120, '2023-10-05 14:48:00');". The output is the news data saved in the database.

[0411] Step 3:

[0412] The server selects news data from a database based on specific parameters. The input is the news data stored in the database. The server selects the most relevant news articles based on parameters such as the number of clicks and publication date, and temporarily stores them in memory. For example, it executes the query "SELECT FROM news_articles ORDER BY clicks DESC LIMIT 1;". The output is the selected news data.

[0413] Step 4:

[0414] The server generates a text script for reading aloud based on the selected news data. The input is the selected news data. The server extracts the title and content from the news data and generates a text script that says, "We will report the next news item. The title is 'Example News'." The output is a text script for reading aloud.

[0415] Step 5:

[0416] The server sends this text script to the avatar engine, which uses speech synthesis and animation technologies to generate a video in which an avatar reads the script. The input is the text script. The avatar engine (e.g., Unity or Unreal Engine) generates a video based on the passed script. The output is the generated avatar video.

[0417] Step 6:

[0418] The server uploads the generated avatar video to the streaming service, encodes it, and prepares it for distribution. The input is the generated avatar video. Specifically, it sends a "POST / upload" request to the streaming service (e.g., Wowza Streaming Engine). The output is the stream URL that users can access.

[0419] Step 7:

[0420] The device provides a user interface (UI) to the user. The input is the stream URL. This UI includes a video player, a comment input field, a question button, and so on. The user inputs comments and questions in real time through this interface. Specifically, the user inputs "I'd like to know more about the background of this news," and the device sends a "POST / comments" request to the server. The output is the user's comments and questions.

[0421] Step 8:

[0422] The server receives and processes comments and questions from users in real time. The input is the user's comment or question. The server sends the received comment to the avatar engine, and the avatar generates an appropriate response. For example, a script such as "About the background of this news..." is passed to the avatar engine to generate a new video. The output is the generated avatar's response video.

[0423] Step 9:

[0424] The server uploads the generated response video back to the streaming service and delivers it to the user. The input is the avatar's response video. The server uploads new videos to the streaming service, updating the delivery in real time. The output is the updated video stream.

[0425] (Application example 1)

[0426] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0427] In recent years, with the digitalization of information and the spread of the Internet, methods of news distribution have become more diverse. However, systems that allow users to obtain information interactively in real time are insufficient, and there is a need for more information to be provided visually and audibly. In addition to allowing users to efficiently browse a wide range of news, there is also a need for systems that can provide information based on the user's interests.

[0428] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0429] In this invention, the server includes means for acquiring data from external information sources, means for storing the acquired data, means for selecting information from the stored data based on specific parameters, means for generating the selected information as audio and video, means for distributing the generated audio and video, means for processing comments and questions received in real time during distribution and generating responses, and means for providing the data on a smartphone application, thereby enabling users to watch the news interactively in real time and receive information visually and audibly.

[0430] "External information sources" refer to information sources that exist outside the system, such as online news services or APIs.

[0431] "Means for Obtaining Data" refers to methods or techniques for periodically requesting and receiving data from external sources.

[0432] "Means for storing data" refers to the methods or techniques for storing acquired data in a database or other storage medium.

[0433] "Information screening means" refers to methods or technologies for extracting selected information from stored data based on specific criteria (e.g., number of clicks or publication date).

[0434] "Means for generating audio and video" refers to technologies for creating audio and video from selected information, in particular methods for using speech synthesis and animation technologies to read aloud using an avatar.

[0435] "Delivery Means" means the method or technology by which the generated audio and video is made available to users in streaming format.

[0436] "Means for processing comments and questions and generating responses" refers to a method or technology for receiving comments and questions sent in real time by users during a broadcast and generating and returning appropriate responses to them.

[0437] "Smartphone Application" means software that runs on a smartphone device and that provides users with the ability to interactively browse news.

[0438] "Real-time comments and questions" refers to comments and questions submitted by users during a video broadcast, and refers to information processed to respond to such comments and questions in a timely manner.

[0439] "Means for processing comments and questions received in real time during the broadcast and generating responses" refers to a method for instantly analyzing feedback from users during a live broadcast and generating appropriate responses.

[0440]

[0441] To implement the present invention, a system including a plurality of components is required, which will be described below with specific examples.

[0442] Data Acquisition Phase

[0443] The server periodically sends requests to external information sources (such as the API of a news service) to obtain the latest news data. The server securely sends the requests using API keys and authentication information. The obtained news data is often received in JSON format or similar.

[0444] Data storage phase

[0445] The server stores the acquired news data in a database. The news data includes information such as the title, text, number of clicks, and publication date and time. This allows for efficient management of data required for subsequent processing.

[0446] News selection phase

[0447] The server sifts through the news data from the database, selecting the latest and most interesting news articles based on certain parameters such as click count, publication date, etc. The sift data is temporarily stored in memory for further processing.

[0448] Avatar generation phase

[0449] The server generates a text script for reading based on the selected news data. The generated script is sent to the avatar engine, which then generates a video of an avatar reading the script. The avatar engine uses speech synthesis and animation technologies to generate natural-looking movements and voices.

[0450] Video distribution phase

[0451] The server uploads the generated avatar video to the streaming server, encodes it, and prepares it for distribution. Once preparation is complete, a stream URL is generated that users can access. Users can watch the video through this URL on their smartphone application.

[0452] Interactive Phase

[0453] Users can enter comments and questions in real time using the smartphone application, including the video player, comment input field, and question button. The server receives and processes comments and questions from users in real time. The processed information is sent to the avatar engine, where the avatar generates an appropriate response. The generated response is then incorporated back into the video stream and delivered to the user.

[0454] For example, suppose a news site experiences a sudden spike in clicks on a particular article. The server periodically retrieves data from the news site and detects this popular article. Next, it generates a text script from the article's title and text, and an avatar reads it aloud using audio and video. The generated video is then delivered to the user in real time. The user can then enter "I'd like to know more about the background to this news story" in the comment input field. The server receives the comment, sends it to the avatar engine, and the avatar generates and delivers a response such as "The background to this news story is..."

[0455] Example prompt sentence:

[0456] "News title: {title}

[0457] News text: {text}

[0458] Please use this as a basis to generate a video of the avatar reading the text."

[0459] This system allows users to receive news visually and audibly, allowing them to enjoy real-time interactive communication.

[0460] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0461] Step 1: Data Acquisition Phase

[0462] The server sends an HTTP request to the news service's API to obtain the latest news data. The request header, including the API key and authentication information, is used as input. The response from the API receives news data in JSON format. The data includes the article title, text, number of clicks, publication date, and so on.

[0463] Step 2: Data storage phase

[0464] The server stores the retrieved news data in a local database. As input, it uses each field of the received news data (title, body, number of clicks, publication date, etc.). As output, each news article is stored in the database. Specifically, it adds data to the database using the SQL INSERT statement.

[0465] Step 3: News selection phase

[0466] The server selects news data from the database. As input, it uses the news data stored in the database. The selection criteria are specific parameters such as the number of clicks or publication date. As output, it selects the most interesting news articles. Specifically, it uses a SQL SELECT statement to filter the articles that match the criteria.

[0467] Step 4: Avatar generation phase

[0468] The server generates a text script for reading aloud based on the selected news data. The selected news data (title, body text) is used as input. The text script is generated as output. The generated script is then sent to the avatar engine, which generates audio and video for the avatar to read aloud. Specifically, the server calls the avatar engine's API to convert the text into audio and animation.

[0469] Step 5: Video distribution phase

[0470] The server uploads the generated avatar video to the streaming server, encodes it, and prepares it for distribution. The generated avatar video is used as input. As output, a stream URL that can be accessed by users is generated. Specifically, the server uses a file transfer protocol to upload the video to the streaming server and generate the URL.

[0471] Step 6: User Interaction Phase

[0472] Users enter comments and questions in real time via a smartphone application using a video player, comment input field, question button, etc. The comments and questions entered by the user on the application are used as input. The input content is sent to the server as output. Specifically, the system operates by using a UI component that accepts user input and a network function that sends it to the server.

[0473] Step 7: Comment and question handling phase

[0474] The server receives and processes comments and questions submitted by users in real time. It uses the comments and questions submitted by users as input. As output, it generates an appropriate response. Specifically, it analyzes the content of the comments and questions and sends them to the avatar engine to generate a response.

[0475] Step 8: Response Delivery Phase

[0476] The server then incorporates the generated response back into the video stream and delivers it to the user. It uses the generated response as input, and delivers an updated video stream to the user as output. Specifically, it uses the streaming server's API to incorporate the response into the video and delivers the updated stream.

[0477] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0478] To implement this invention, a system including a number of components is required, which will be described below with specific examples.

[0479] Data Acquisition Phase

[0480] The server periodically sends requests to external sources (such as the API of a news service) to obtain the latest news data. The requests are sent securely using API keys and authentication information, and the news data is received in JSON format or other formats.

[0481] Data storage phase

[0482] The server stores the acquired news data in a database. The news data includes information such as the title, text, number of clicks, and publication date and time. This allows for efficient management of data required for subsequent processing.

[0483] News selection phase

[0484] The server selects news data from the database, sorts the news articles by the number of clicks based on certain parameters (such as the number of clicks or the publication date), and selects the latest and most interesting news articles. The selected data is temporarily stored in memory for further processing.

[0485] Avatar generation phase

[0486] The server generates a text script for reading based on the selected news data. The generated script is sent to the avatar engine, which then generates a video of an avatar reading the script. The avatar engine uses speech synthesis and animation technologies to generate natural-looking movements and voices.

[0487] Video distribution phase

[0488] The server uploads the generated avatar video to the streaming server, encodes it, and prepares it for distribution. Once preparation is complete, a stream URL is generated that users can access.

[0489] Interactive Phase

[0490] The device provides the user with a user interface that includes a video player, a comment input field, a question button, etc. The user uses this interface to watch videos and input comments and questions in real time.

[0491] The server receives comments and questions sent from the device in real time and processes them. It receives comments using real-time communication technologies such as WebSocket and adds them to a queue.

[0492] The server takes each comment from the queue and sends it to the avatar engine, which then generates an appropriate response, which is generated as audio and video and added to the live stream.

[0493] Emotion Recognition Phase

[0494] The server uses an emotion engine to analyze comments and questions entered by users. The emotion engine recognizes the user's emotions (e.g., joy, sadness, anger, etc.) from the text contained in the comments and questions.

[0495] The server adjusts the response of the generated avatar appropriately based on the user's emotion recognized by the emotion engine. For example, if the user expresses anger, the avatar's response will be calm.

[0496] Display and Feedback Phase

[0497] The terminal provides a method for visually displaying the user's emotional information recognized by the emotion engine, allowing the user to confirm that their own emotional state is reflected.

[0498] Specific examples

[0499] For example, suppose a news site experiences a sudden spike in clicks on a particular article. The server periodically retrieves data from the news site to identify popular articles. It then generates a text script from the article's title and text, and an avatar reads it aloud with audio and video. The generated video is then streamed to the user in real time, who can then type in a comment field, "I'd like to know more about the background to this news story."

[0500] The server receives the comment and recognizes the "interest" in the emotion engine. The avatar's response generated by the avatar engine provides detailed information in a way that attracts the user's interest.

[0501] This system not only allows users to receive news visually and audibly, but also allows for real-time interactive communication, and further enhances the user experience by recognizing users' emotions and providing responses accordingly.

[0502] The processing flow will be explained below.

[0503] Step 1:

[0504] The server periodically sends requests to an external source (such as a news service API) to retrieve the latest news data. The request is sent securely using an API key or authentication information, and the news data is received in JSON format.

[0505] Step 2:

[0506] The server stores the acquired news data in a database. The news data includes information such as the title, text, number of clicks, and publication date and time, and is stored in the database using the INSERT statement.

[0507] Step 3:

[0508] The server sorts the news data from the database, executes a SELECT statement based on specific parameters (number of clicks and publication date), and sorts the news articles by most clicked.

[0509] Step 4:

[0510] The server generates a text script for reading out the selected news data, formats it as "Title: ●●, Body: ●●", and sends the generated text script to the avatar engine.

[0511] Step 5:

[0512] The server receives the avatar video generated by the avatar engine, uploads it to the streaming server, encodes the video file, and generates a stream URL that users can access.

[0513] Step 6:

[0514] The terminal provides a user interface, including a video player, a comment input field, and a question button, which users use to watch videos and input comments and questions.

[0515] Step 7:

[0516] The server receives comments and questions sent from the device in real time and adds the comments to a queue using real-time communication technologies such as WebSocket.

[0517] Step 8:

[0518] The server takes each comment from the queue and sends it to the avatar engine, which then generates an appropriate response, which is then generated as audio and video and added to the live stream.

[0519] Step 9:

[0520] The server recognizes emotions from comments and questions submitted by users using an emotion engine, which extracts the user's emotions (e.g., joy, sadness, anger) from the text.

[0521] Step 10:

[0522] The server adjusts the response generated by the avatar engine based on the user's emotion recognized by the emotion engine, for example, if the user expresses anger, the avatar's response will be in a calm tone.

[0523] Step 11:

[0524] The device visually displays the user's emotional information recognized by the emotion engine, allowing the user to confirm that their own emotional state is reflected.

[0525] Step 12:

[0526] The user receives real-time responses from the avatar and can continue the interactive communication by entering further comments or questions, a process that is repeated to improve the user experience.

[0527] Example 2

[0528] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0529] Conventional news delivery systems are limited to viewing news and have limited user interaction. Furthermore, they do not provide responses that take into account the user's emotions, resulting in a uniform user experience and low satisfaction for some users. Furthermore, they often lack interactivity because responses to real-time comments and questions are not always prompt.

[0530] The specific processing by the specific processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for acquiring data from an external information source, means for saving the acquired data, means for selecting information from the saved data based on specific parameters, means for generating audio and video from the selected information, means for distributing the generated audio and video, means for processing comments and questions received in real time during distribution and generating responses, and means for performing sentiment analysis on the received comments and questions and adjusting responses based on the results. This increases interactivity in news distribution and enables responses that correspond to the user's emotions. It also enhances real-time interaction with users and improves user satisfaction.

[0531] "External sources" refers to sources for obtaining data from outside the system, such as news service APIs and websites.

[0532] "Means for obtaining data" refers to a combination of programming and hardware for sending requests to external sources and receiving the required data.

[0533] "Means for storing data" refers to a database management system or storage device for efficiently storing acquired data.

[0534] "Means for filtering information from data based on specific parameters" refers to programs that filter stored data using criteria such as the number of clicks or publication date and time to extract the desired information.

[0535] "Audio and video generation means" means a program that converts selected information into a text script and uses speech synthesis and animation techniques to make it visually and audibly reproducible.

[0536] "Means for delivering audio and video" refers to the streaming server and encoding technology used to deliver the generated audio and video to users.

[0537] "Means for processing comments and questions received in real time and generating responses" refers to programs and communication technologies for collecting and analyzing user input in real time and generating appropriate responses based on the results.

[0538] "Means for analyzing emotions and adjusting responses based on the results" refers to a program that analyzes the emotions contained in comments and questions from users and generates an appropriate response based on those emotions.

[0539] To implement the present invention, several hardware and software components are required, which will be described below with specific examples.

[0540] Data Acquisition Phase

[0541] The server is responsible for obtaining data from external sources. Specifically, it periodically sends requests to the news service's API to obtain the latest news data. This is done by securely sending requests using API keys and authentication information, and receiving news data in JSON format or similar. For example, the server uses the NewsAPI to obtain the latest 50 articles at a time.

[0542] Data storage phase

[0543] The server stores the acquired news data in a database system (e.g., MySQL, PostgreSQL). News data includes information such as the title, text, number of clicks, and publication date and time. The server inserts this data into a database to efficiently manage it. For example, the news title "Breaking News," its text, number of clicks "325," and publication date and time "2023-10-01" are stored in the database.

[0544] News selection phase

[0545] When filtering news data from the database, the server selects news articles based on certain parameters, sorts the news articles based on the number of clicks or publication date, or sorts the articles by most clicked or most recently published, and stores the filtered data in memory. For example, it selects the top 10 articles with the most clicks.

[0546] Avatar generation phase

[0547] The server generates a text script based on the selected news data and sends the script to an avatar engine (e.g., FaceRig, Voki). The avatar engine uses speech synthesis technology (e.g., Google Text-to-Speech API) and animation technology to generate avatar videos with natural voices and movements. For example, it generates a video of an avatar reading a news article based on the text script.

[0548] Video distribution phase

[0549] The server uploads the generated avatar video to a streaming server (e.g., Wowza Streaming Engine), encodes it, and prepares it for distribution. Once preparation for distribution is complete, a stream URL is generated that users can access. For example, "https: / / streamingserver.com / newsXstream" is provided to users.

[0550] Interactive Phase

[0551] The device provides the user with a user interface that includes a video player, a comment input field, and a question button. The user uses these interfaces to watch videos and input comments and questions in real time. The server receives comments and questions sent from the device using real-time communication technology such as WebSocket and adds them to a queue. For example, a user might input a comment such as, "I'd like to know more about the background of this news."

[0552] Emotion Recognition Phase

[0553] The server uses an emotion engine (e.g., IBM Watson Tone Analyzer) to analyze the user's comments and questions and recognize the user's emotions. Based on this analysis, the server adjusts the tone and content of the response appropriately. For example, if the user's comments are analyzed as indicating "interest," the avatar's response will be adjusted to be more detailed and interesting.

[0554] Display and Feedback Steps

[0555] The device provides an interface that visually displays the user's emotional information recognized by the emotion engine, allowing the user to confirm that their emotional state is accurately reflected. For example, an icon displayed next to the comment field represents the user's emotion.

[0556] Examples:

[0557] For example, if a news site experiences a sudden spike in clicks on a particular article, the server periodically retrieves data from that news site to identify popular articles. Next, a text script is generated from the article's title and text, and an avatar reads it aloud using audio and video. The generated video is then streamed to the user in real time, who can then type in a comment field, "I'd like to know more about the background to this news story."

[0558] The server receives the comments and recognizes "interest" using the emotion engine. The avatar engine generates responses from avatars, providing detailed information in a way that piques the user's interest. This system allows users to not only receive news visually and audibly, but also enjoy real-time interactive communication.

[0559] Example prompt sentence:

[0560] "Generate text scripts from the titles and text of breaking news articles and create videos of avatars reading them. Also, generate responses to user comments and respond in the right tone based on sentiment analysis."

[0561] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0562] Step 1: Data Acquisition Step

[0563] The server retrieves data from external information sources. Specifically, it periodically sends requests to the news service's API. The input is a request that includes the API key and authentication information. The server sends the request, and the data it retrieves is news data in JSON format. Specifically, the server accesses the News API and sends a request saying, "Please retrieve the latest 50 news articles." In response, it receives the JSON data of the news articles.

[0564] Step 2: Data storage step

[0565] The server saves the acquired news data in a database. The input is the JSON-formatted news data acquired in step 1. The server parses the JSON data, extracts fields such as the news article title, text, number of clicks, and publication date and time, and inserts them into a database (e.g., MySQL, PostgreSQL). The output is the news data saved in the database. Specifically, the news title "Breaking News", text "Detailed Content", number of clicks "325", and publication date and time "2023-10-01" are saved in the database.

[0566] Step 3: News selection step

[0567] The server selects news data from the database. The input is the news data stored in the database. The server sorts the news articles based on certain parameters (e.g., number of clicks or publication date) and selects the latest and most clicked news articles. The output is the selected news data. Specifically, it selects the top 10 articles with the most clicks. The data is temporarily stored in memory.

[0568] Step 4: Avatar generation step

[0569] The server generates a text script based on the selected news data. The input is the news data selected in step 3. The server combines the news title and text to generate a text script and sends it to the avatar engine. The output is an avatar video that reads the text script aloud. Specifically, it generates a text script that reads the news title "Breaking News" and its text aloud and sends it to the avatar engine. The avatar engine generates the video using speech synthesis and animation technologies.

[0570] Step 5: Video distribution step

[0571] The server uploads the generated avatar video to the streaming server. The input is the avatar video generated in step 4. The server uploads the video to the streaming server (e.g., Wowza Streaming Engine), encodes it, and prepares it for distribution. The output is a stream URL that users can access. Specifically, the generated avatar video is uploaded to "https: / / streamingserver.com / newsXstream" so that users can access it.

[0572] Step 6: Interactive Step

[0573] The terminal provides an interface to the user. The input is the stream URL from the streaming server. The terminal displays a video player, a comment input field, a question button, etc. to the user. The user uses these interfaces to watch the video and input comments and questions in real time. The output is comments and questions from the user. Specifically, the user inputs a comment such as "I'd like to know more about the background of this news."

[0574] Step 7: Emotion Recognition Step

[0575] The server analyzes the user's comments and questions using an emotion engine. The input is the comment or question received from the user in step 6. The server uses an emotion engine (e.g., IBM Watson Tone Analyzer) to recognize the user's emotions (e.g., interest, joy, sadness, anger). The output is the emotion analysis result and an adjusted response. Specifically, if the user's comment is recognized as indicating "interest," the server adjusts the avatar's response to be more detailed and interesting.

[0576] Step 8: Display and Feedback Step

[0577] The device visually displays the user's emotional information. The input is the emotion analysis result obtained in step 7. The device displays an icon that expresses the user's emotion next to the comment field. The output is the visual display of the user's emotional information. Specifically, emotion icons (e.g., smiley, tearful, angry expressions) are displayed in the comment field, allowing the user to confirm that their emotions are accurately reflected.

[0578] (Application example 2)

[0579] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0580] In conventional news delivery systems, news acquisition, selection, and delivery are one-way, limiting user interaction. Furthermore, it is difficult to generate appropriate responses based on user emotions, making it difficult to improve the user experience. In particular, the lack of real-time emotion recognition and its reflection can lead to a decline in user engagement.

[0581] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for acquiring data from an external information source, means for saving the acquired data, means for selecting information from the saved data based on specific parameters, means for generating the selected information as audio and video, means for distributing the generated audio and video, means for processing comments and questions received in real time during distribution and generating a response, means for recognizing the user's emotions from the comments and questions and adjusting the response, and means for visually displaying the results of the emotion recognition. This enables interactive communication in real time while watching the news, and makes it possible to provide an appropriate response according to the user's emotions.

[0582] "External sources" are data sources such as news services, APIs, and data feeds available on the Internet.

[0583] "Data Acquisition Methods" are the systems and processes used to periodically request and acquire required information from external sources.

[0584] "Means for storing data" refers to a database or storage device for temporarily or permanently storing acquired information.

[0585] "Information filtering means" refers to algorithms or programs that extract important data from stored information based on specific conditions.

[0586] The "means for generating audio and video" is a system that creates visual and auditory content based on text information using speech synthesis technology and video generation technology.

[0587] The "means for delivering audio and video" refers to a server or network configuration for streaming the generated media content to users.

[0588] A "means for processing comments and questions received in real time and generating responses" is a program or process that receives input from users in real time and generates appropriate responses based on that input.

[0589] "Means for recognizing user emotions from comments and questions and adjusting responses" refers to emotion recognition engines and algorithms that analyze emotions from input text and adjust response content based on the results.

[0590] The "means for visually displaying the results of emotion recognition" refers to a function or program for visually displaying the analyzed emotion information on a user interface.

[0591] To implement this invention, a system including a server, a user terminal, a generative AI model, etc. Specifically, the system is realized by the following steps.

[0592] First, the server periodically retrieves data from external sources, using API keys and authentication information to securely retrieve data, such as from a news service API. The retrieved data is received in JSON format.

[0593] The acquired data is then stored in a database on the server, which stores information such as the news article title, text, number of clicks, and publication date and time, allowing for efficient management of the data required for subsequent processing.

[0594] The server then sorts through the vast amount of news articles stored in the database, using criteria such as click counts and publication dates to select the most interesting news articles.

[0595] Based on the selected news data, the server generates a text script for reading. This script is sent to an avatar engine, which uses speech synthesis and animation technologies to create an avatar video that generates natural-looking movements and voices. The avatar engine can be Amazon Polly or Google Text-to-Speech.

[0596] The generated avatar video is uploaded to a streaming server, where it is encoded and prepared for distribution. Once preparation is complete, a stream URL is generated that users can access.

[0597] The user terminal provides a user interface including a video player, a comment input field, a question button, etc. The user can use this interface to watch the video and input comments and questions in real time.

[0598] The server receives comments and questions sent from the device in real time and processes them using WebSocket technology, etc. Comments and questions are added to a queue, then sequentially retrieved and sent to the avatar engine, where the avatar generates an appropriate response.

[0599] In the emotion recognition phase, the server uses an emotion engine to analyze the comments and questions entered by the user. The emotion engine uses IBM Watson Natural Language Understanding and other technologies. It recognizes emotions from the user's input text and adjusts the response accordingly. For example, if the user expresses anger, the avatar's response will be calmer.

[0600] The results of emotion recognition are visually displayed on the user's device, allowing them to see responses that reflect their own emotional state, enabling more personalized, real-time communication.

[0601] Specific examples

[0602] For example, a "news avatar" reads the latest news article, and a viewer enters a comment such as, "I'd like to know more about the background to this news." The server receives this comment in real time and uses an emotion recognition engine to identify "interest." The avatar engine then generates a response such as, "Let me explain the details of this news story," providing information in a way that piques the viewer's interest. This process allows users to enjoy real-time interactive communication while watching the news broadcast.

[0603] Prompt Sentence Examples

[0604] "A news avatar will read the latest news to viewers, and users can enter comments to instantly respond. Emotion recognition will be performed for each comment, and the avatar will respond appropriately. For example, create a system that responds to a comment such as 'I'd like to know more about the background of this news story,' with 'I'll explain the details of this news story.'"

[0605] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0606] Step 1:

[0607] The server accesses an external information source (the API of a news service) to obtain the latest news data. At this time, it securely requests the data using an API key and authentication information, and receives the news data in JSON format. The input is the news data received from the API, and the output is raw news data that is temporarily stored inside the server.

[0608] Step 2:

[0609] The server stores the acquired news data in a database, which stores information such as the title, text, number of clicks, and publication date and time of news articles. The input is raw news data, and the output is structured data stored in the database. This allows for efficient management of the data required for subsequent processing.

[0610] Step 3:

[0611] The server selects news data from the database based on specific parameters (e.g., number of clicks or publication date). The selected news data is temporarily stored in memory. The input is the news data stored in the database, and the output is the selected news data.

[0612] Step 4:

[0613] The server generates a text script for reading based on the selected news data. This script is sent to the avatar engine, which uses speech synthesis and animation technologies to make the avatar move and speak naturally. The input is the selected news data, and the output is the generated avatar video.

[0614] Step 5:

[0615] The server uploads the generated avatar video to the streaming server, encodes it, and prepares it for distribution. Once it is ready for distribution, a stream URL that users can access is generated. The input is the avatar video, and the output is the stream URL.

[0616] Step 6:

[0617] The user terminal provides a user interface that includes a video player, a comment input field, a question button, etc. The user uses the interface to watch videos and input comments and questions in real time. The input is the user's comments and questions, and the output is the corresponding interface operation.

[0618] Step 7:

[0619] The server receives comments and questions sent from the terminal in real time and adds the comments to a queue using WebSocket technology. The input is the comments and questions received from the user, and the output is the data added to the comment queue.

[0620] Step 8:

[0621] The server sequentially retrieves comments added to the queue and sends them to the avatar engine, where the avatar generates an appropriate response. The input is the comment retrieved from the comment queue, and the output is the generated avatar response.

[0622] Step 9:

[0623] The server analyzes the user's comments and questions using an emotion engine and adjusts the avatar's response based on the emotion. The input is the user's comments and questions, and the output is the adjusted avatar response. The emotion engine uses IBM Watson Natural Language Understanding and other technologies.

[0624] Step 10:

[0625] The user device visually displays the emotion recognition results, allowing the user to confirm the response that reflects their own emotional state. The input is the analysis result of the emotion engine, and the output is the visually displayed emotional information.

[0626] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0627] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0628] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0629] [Third embodiment]

[0630] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0631] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0632] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0633] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0634] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0635] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0636] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0637] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0638] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0639] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0640] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0641] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0642] To implement the present invention, a system including a plurality of components is required, which will be described below with specific examples.

[0643] Data Acquisition Phase

[0644] The server periodically sends requests to external information sources (such as the API of a news service) to obtain the latest news data. The server securely sends the requests using API keys and authentication information. The obtained news data is often received in JSON format or similar.

[0645] Data storage phase

[0646] The server stores the acquired news data in a database. The news data includes information such as the title, text, number of clicks, and publication date and time. This allows for efficient management of data required for subsequent processing.

[0647] News selection phase

[0648] The server sifts through the news data from the database, selecting the latest and most interesting news articles based on certain parameters such as click count, publication date, etc. The sift data is temporarily stored in memory for further processing.

[0649] Avatar generation phase

[0650] The server generates a text script for reading based on the selected news data. The generated script is sent to the avatar engine, which then generates a video of an avatar reading the script. The avatar engine uses speech synthesis and animation technologies to generate natural-looking movements and voices.

[0651] Video distribution phase

[0652] The server uploads the generated avatar video to the streaming server, encodes it, and prepares it for distribution. Once preparation is complete, a stream URL is generated that users can access.

[0653] Interactive Phase

[0654] The terminal provides the user with a user interface, which includes a video player, a comment input field, a question button, etc. When the user uses this interface to input comments or questions in real time, they are sent to the server.

[0655] The server receives and processes comments and questions from users in real time, sends the processed information to the avatar engine, and the avatar generates an appropriate response, which is then re-integrated into the video stream and delivered to the user.

[0656] Specific examples

[0657] For example, suppose a news site experiences a sudden spike in clicks on a particular article. The server periodically retrieves data from the news site and detects this popular article. Next, it generates a text script from the article's title and text, and an avatar reads it aloud using audio and video. The generated video is delivered to the user in real time, and the user can enter "I'd like to know more about the background to this news story" in the comment input field. The server receives the comment and sends it to the avatar engine, where the avatar generates and delivers a response such as "The background to this news story is..."

[0658] This system allows users to receive news visually and audibly, allowing them to enjoy real-time interactive communication.

[0659] The processing flow will be explained below.

[0660] Step 1:

[0661] The server periodically sends requests to an external source (such as a news service API) to obtain the latest news data. The request is sent securely using an API key or authentication information, and the news data is received in JSON format or other formats.

[0662] Step 2:

[0663] The server stores the acquired news data in a database. The news data includes information such as the title, text, number of clicks, and publication date and time, and is stored in the database using the INSERT statement.

[0664] Step 3:

[0665] The server sifts through the news data from the database, sorts the news articles by most clicked based on certain parameters (e.g., number of clicks, publication date), and selects the most recent and likely most interesting news articles.

[0666] Step 4:

[0667] The server generates a text script for reading out the selected news data, formats it into a format such as "Title: ●●, Body: ●●", and sends the generated text script to the avatar engine.

[0668] Step 5:

[0669] The server receives the avatar video generated by the avatar engine and uploads it to the streaming server, where it encodes the video file and prepares it for distribution.

[0670] Step 6:

[0671] The device provides the user with a user interface that includes a video player, a comment input field, a question button, etc. The user uses this interface to watch videos and input comments and questions in real time.

[0672] Step 7:

[0673] The server receives comments and questions sent from the device in real time and processes them. It receives comments using real-time communication technologies such as WebSocket and adds them to a queue.

[0674] Step 8:

[0675] The server takes each comment from the queue and sends it to the avatar engine, which then generates an appropriate response, which is generated as audio and video and added to the live stream.

[0676] Step 9:

[0677] The user receives real-time responses from the avatar via a live stream, providing a visual and auditory experience, and can enter further comments or questions as needed.

[0678] Example 1

[0679] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0680] There is a need to efficiently manage data obtained from external information sources and quickly select important information based on specific parameters. It is also important to deliver this information in a way that users can perceive visually and audibly in a natural way. Furthermore, there is a need for a system that can respond immediately to comments and questions received in real time during the broadcast, enabling interactive communication. However, current systems have difficulty fully meeting these requirements, and technical challenges remain, particularly in real-time processing and audio-video integration.

[0681] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0682] In this invention, the server includes: means for acquiring data from external information sources; means for saving the acquired data; means for selecting information from the saved data based on specific parameters; means for generating audio and video from the selected information; means for distributing the generated audio and video; means for processing comments and questions received in real time during distribution and generating responses; means for selecting news data from data stored in a database and generating videos read by avatars using speech synthesis and animation technologies; means for encoding and distributing avatar videos uploaded to a streaming service; and means for processing user comments and questions on the encoded videos in real time and generating and redistributing avatar responses. This allows users to naturally receive important information visually and audibly, enabling further interactive communication in real time.

[0683] "External information sources" are resources for obtaining data provided by the Internet or specific service providers.

[0684] "Data acquisition means" are the methods and techniques used to collect information from external sources.

[0685] "Means for storing data" refers to the methods and techniques used to hold acquired information in a database or other storage device.

[0686] "Specific parameters" are numerical values ​​or conditions that serve as criteria when selecting or processing data.

[0687] "Means for selecting information" are methods and techniques for extracting necessary information from stored data.

[0688] "Means for generating audio and video" refers to methods and technologies for synthesizing audio and producing video based on selected information.

[0689] "Means of distribution" refers to the methods and technologies used to transmit the generated audio and video to users over the Internet.

[0690] "Means for processing comments and questions received in real time and generating responses" refers to methods and technologies for receiving immediate feedback from users during a broadcast and generating responses based on that feedback.

[0691] A "database" is a system or software for efficiently managing and manipulating large amounts of data.

[0692] "News data" refers to the latest information obtained from news services.

[0693] "Speech synthesis technology" is a technology for artificially generating speech based on text data.

[0694] "Animation technology" is a technique that makes still images appear to be moving by displaying them in succession.

[0695] An "avatar" is a computer-generated virtual representation of a person or character.

[0696] A "streaming service" is a service for delivering digital content in real time over the Internet.

[0697] "Encoding" is the process of converting digital data into a particular format.

[0698] A "user interface" refers to the screen and operation method that allows a user to interact with a system.

[0699] MODE FOR CARRYING OUT THE INVENTION

[0700] To implement this invention, a system including multiple components is required. This system mainly consists of three elements: a server, a terminal, and a user.

[0701] First, the server periodically sends requests to an external information source (for example, the API of a news service) to obtain the latest news data. This request uses HTTPS communication, and an API key or OAuth token is used to ensure security. The obtained news data is typically returned in JSON format. An example request is "GET / latest-news?apiKey=YOUR_API_KEY".

[0702] Next, the server stores the retrieved news data in a database (e.g., PostgreSQL). The database contains the title, content, click count, and publication date / time of each news article. The server efficiently organizes this information and stores it in a structured format. For example, the following query is used: "INSERT INTO news_articles (title, content, clicks, published_at) VALUES ('Example News', 'Example News Content', 120, '2023-10-05 14:48:00');"

[0703] The server selects news data from the database based on certain parameters, such as the articles with the most clicks or the most recent publication date. This selection is performed using an SQL query, such as "SELECT FROM news_articles ORDER BY clicks DESC LIMIT 1;".

[0704] Based on the selected news data, the server generates a text script for reading aloud. This is then used with speech synthesis and animation technology to generate a video in which an avatar reads the script aloud. An avatar engine (e.g., Unity or Unreal Engine) is used for this process. For example, the generated script might read, "We will report the next news item. The title is 'Example News'."

[0705] The generated avatar video is uploaded to a streaming service (e.g., Wowza Streaming Engine) by the server. The server encodes the video and prepares it for distribution. Once preparation for distribution is complete, a stream URL is generated that users can access. For example, the server sends a request called "POST / upload" and receives the stream URL as a response.

[0706] The device provides the user with a user interface (UI). This UI includes a video player, a comment input field, a question button, and so on. The user uses these interfaces to input comments and questions in real time. For example, the user might input a comment such as "I'd like to know more about the background of this news," and the device sends the comment to the server. A request called "POST / comments" is sent, and the comment content is included in the request.

[0707] The server receives comments and questions from users in real time and processes them. The processed information is sent back to the avatar engine, where the avatar generates an appropriate response. The generated response is incorporated into the video stream and delivered to the user. For example, the server passes a script such as "Some background on this news..." to the avatar engine to generate a new video.

[0708] As an example of a prompt sentence, the text "Tell me about the latest news" is input to the generative AI model. Based on this prompt sentence, the entire system operates, carrying out a series of processes from news information collection to distribution and interactive response.

[0709] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0710] Step 1:

[0711] The server periodically sends requests to external news sources to retrieve the latest news data. The input is the API endpoint and authentication information (e.g., API key) of the external news source. The server sends the request "GET / latest-news?apiKey=YOUR_API_KEY" and receives the news data in JSON format. This data includes the title, body, number of clicks, publication date, etc. of the news article.

[0712] Step 2:

[0713] The server saves the acquired news data in a database. The input is news data in JSON format. The server organizes this data and stores it in a database (e.g., PostgreSQL). Specifically, it executes the query "INSERT INTO news_articles (title, content, clicks, published_at) VALUES ('Example news', 'Example news content', 120, '2023-10-05 14:48:00');". The output is the news data saved in the database.

[0714] Step 3:

[0715] The server selects news data from a database based on specific parameters. The input is the news data stored in the database. The server selects the most relevant news articles based on parameters such as the number of clicks and publication date, and temporarily stores them in memory. For example, it executes the query "SELECT FROM news_articles ORDER BY clicks DESC LIMIT 1;". The output is the selected news data.

[0716] Step 4:

[0717] The server generates a text script for reading aloud based on the selected news data. The input is the selected news data. The server extracts the title and content from the news data and generates a text script that says, "We will report the next news item. The title is 'Example News'." The output is a text script for reading aloud.

[0718] Step 5:

[0719] The server sends this text script to the avatar engine, which uses speech synthesis and animation technologies to generate a video in which an avatar reads the script. The input is the text script. The avatar engine (e.g., Unity or Unreal Engine) generates a video based on the passed script. The output is the generated avatar video.

[0720] Step 6:

[0721] The server uploads the generated avatar video to the streaming service, encodes it, and prepares it for distribution. The input is the generated avatar video. Specifically, it sends a "POST / upload" request to the streaming service (e.g., Wowza Streaming Engine). The output is the stream URL that users can access.

[0722] Step 7:

[0723] The device provides a user interface (UI) to the user. The input is the stream URL. This UI includes a video player, a comment input field, a question button, and so on. The user inputs comments and questions in real time through this interface. Specifically, the user inputs "I'd like to know more about the background of this news," and the device sends a "POST / comments" request to the server. The output is the user's comments and questions.

[0724] Step 8:

[0725] The server receives and processes comments and questions from users in real time. The input is the user's comment or question. The server sends the received comment to the avatar engine, and the avatar generates an appropriate response. For example, a script such as "About the background of this news..." is passed to the avatar engine to generate a new video. The output is the generated avatar's response video.

[0726] Step 9:

[0727] The server uploads the generated response video back to the streaming service and delivers it to the user. The input is the avatar's response video. The server uploads new videos to the streaming service, updating the delivery in real time. The output is the updated video stream.

[0728] (Application example 1)

[0729] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0730] In recent years, with the digitalization of information and the spread of the Internet, methods of news distribution have become more diverse. However, systems that allow users to obtain information interactively in real time are insufficient, and there is a need for more information to be provided visually and audibly. In addition to allowing users to efficiently browse a wide range of news, there is also a need for systems that can provide information based on the user's interests.

[0731] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0732] In this invention, the server includes means for acquiring data from external information sources, means for storing the acquired data, means for selecting information from the stored data based on specific parameters, means for generating the selected information as audio and video, means for distributing the generated audio and video, means for processing comments and questions received in real time during distribution and generating responses, and means for providing the data on a smartphone application, thereby enabling users to watch the news interactively in real time and receive information visually and audibly.

[0733] "External information sources" refer to information sources that exist outside the system, such as online news services or APIs.

[0734] "Means for Obtaining Data" refers to methods or techniques for periodically requesting and receiving data from external sources.

[0735] "Means for storing data" refers to the methods or techniques for storing acquired data in a database or other storage medium.

[0736] "Information screening means" refers to methods or technologies for extracting selected information from stored data based on specific criteria (e.g., number of clicks or publication date).

[0737] "Means for generating audio and video" refers to technologies for creating audio and video from selected information, in particular methods for using speech synthesis and animation technologies to read aloud using an avatar.

[0738] "Delivery Means" means the method or technology by which the generated audio and video is made available to users in streaming format.

[0739] "Means for processing comments and questions and generating responses" refers to a method or technology for receiving comments and questions sent in real time by users during a broadcast and generating and returning appropriate responses to them.

[0740] "Smartphone Application" means software that runs on a smartphone device and that provides users with the ability to interactively browse news.

[0741] "Real-time comments and questions" refers to comments and questions submitted by users during a video broadcast, and refers to information processed to respond to such comments and questions in a timely manner.

[0742] "Means for processing comments and questions received in real time during the broadcast and generating responses" refers to a method for instantly analyzing feedback from users during a live broadcast and generating appropriate responses.

[0743]

[0744] To implement the present invention, a system including a plurality of components is required, which will be described below with specific examples.

[0745] Data Acquisition Phase

[0746] The server periodically sends requests to external information sources (such as the API of a news service) to obtain the latest news data. The server securely sends the requests using API keys and authentication information. The obtained news data is often received in JSON format or similar.

[0747] Data storage phase

[0748] The server stores the acquired news data in a database. The news data includes information such as the title, text, number of clicks, and publication date and time. This allows for efficient management of data required for subsequent processing.

[0749] News selection phase

[0750] The server sifts through the news data from the database, selecting the latest and most interesting news articles based on certain parameters such as click count, publication date, etc. The sift data is temporarily stored in memory for further processing.

[0751] Avatar generation phase

[0752] The server generates a text script for reading based on the selected news data. The generated script is sent to the avatar engine, which then generates a video of an avatar reading the script. The avatar engine uses speech synthesis and animation technologies to generate natural-looking movements and voices.

[0753] Video distribution phase

[0754] The server uploads the generated avatar video to the streaming server, encodes it, and prepares it for distribution. Once preparation is complete, a stream URL is generated that users can access. Users can watch the video through this URL on their smartphone application.

[0755] Interactive Phase

[0756] Users can enter comments and questions in real time using the smartphone application, including the video player, comment input field, and question button. The server receives and processes comments and questions from users in real time. The processed information is sent to the avatar engine, where the avatar generates an appropriate response. The generated response is then incorporated back into the video stream and delivered to the user.

[0757] For example, suppose a news site experiences a sudden spike in clicks on a particular article. The server periodically retrieves data from the news site and detects this popular article. Next, it generates a text script from the article's title and text, and an avatar reads it aloud using audio and video. The generated video is then delivered to the user in real time. The user can then enter "I'd like to know more about the background to this news story" in the comment input field. The server receives the comment, sends it to the avatar engine, and the avatar generates and delivers a response such as "The background to this news story is..."

[0758] Example prompt sentence:

[0759] "News title: {title}

[0760] News text: {text}

[0761] Please use this as a basis to generate a video of the avatar reading the text."

[0762] This system allows users to receive news visually and audibly, allowing them to enjoy real-time interactive communication.

[0763] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0764] Step 1: Data Acquisition Phase

[0765] The server sends an HTTP request to the news service's API to obtain the latest news data. The request header, including the API key and authentication information, is used as input. The response from the API receives news data in JSON format. The data includes the article title, text, number of clicks, publication date, and so on.

[0766] Step 2: Data storage phase

[0767] The server stores the retrieved news data in a local database. As input, it uses each field of the received news data (title, body, number of clicks, publication date, etc.). As output, each news article is stored in the database. Specifically, it adds data to the database using the SQL INSERT statement.

[0768] Step 3: News selection phase

[0769] The server selects news data from the database. As input, it uses the news data stored in the database. The selection criteria are specific parameters such as the number of clicks or publication date. As output, it selects the most interesting news articles. Specifically, it uses a SQL SELECT statement to filter the articles that match the criteria.

[0770] Step 4: Avatar generation phase

[0771] The server generates a text script for reading aloud based on the selected news data. The selected news data (title, body text) is used as input. The text script is generated as output. The generated script is then sent to the avatar engine, which generates audio and video for the avatar to read aloud. Specifically, the server calls the avatar engine's API to convert the text into audio and animation.

[0772] Step 5: Video distribution phase

[0773] The server uploads the generated avatar video to the streaming server, encodes it, and prepares it for distribution. The generated avatar video is used as input. As output, a stream URL that can be accessed by users is generated. Specifically, the server uses a file transfer protocol to upload the video to the streaming server and generate the URL.

[0774] Step 6: User Interaction Phase

[0775] Users enter comments and questions in real time via a smartphone application using a video player, comment input field, question button, etc. The comments and questions entered by the user on the application are used as input. The input content is sent to the server as output. Specifically, the system operates by using a UI component that accepts user input and a network function that sends it to the server.

[0776] Step 7: Comment and question handling phase

[0777] The server receives and processes comments and questions submitted by users in real time. It uses the comments and questions submitted by users as input. As output, it generates an appropriate response. Specifically, it analyzes the content of the comments and questions and sends them to the avatar engine to generate a response.

[0778] Step 8: Response Delivery Phase

[0779] The server then incorporates the generated response back into the video stream and delivers it to the user. It uses the generated response as input, and delivers an updated video stream to the user as output. Specifically, it uses the streaming server's API to incorporate the response into the video and delivers the updated stream.

[0780] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0781] To implement this invention, a system including a number of components is required, which will be described below with specific examples.

[0782] Data Acquisition Phase

[0783] The server periodically sends requests to external sources (such as the API of a news service) to obtain the latest news data. The requests are sent securely using API keys and authentication information, and the news data is received in JSON format or other formats.

[0784] Data storage phase

[0785] The server stores the acquired news data in a database. The news data includes information such as the title, text, number of clicks, and publication date and time. This allows for efficient management of data required for subsequent processing.

[0786] News selection phase

[0787] The server selects news data from the database, sorts the news articles by the number of clicks based on certain parameters (such as the number of clicks or the publication date), and selects the latest and most interesting news articles. The selected data is temporarily stored in memory for further processing.

[0788] Avatar generation phase

[0789] The server generates a text script for reading based on the selected news data. The generated script is sent to the avatar engine, which then generates a video of an avatar reading the script. The avatar engine uses speech synthesis and animation technologies to generate natural-looking movements and voices.

[0790] Video distribution phase

[0791] The server uploads the generated avatar video to the streaming server, encodes it, and prepares it for distribution. Once preparation is complete, a stream URL is generated that users can access.

[0792] Interactive Phase

[0793] The device provides the user with a user interface that includes a video player, a comment input field, a question button, etc. The user uses this interface to watch videos and input comments and questions in real time.

[0794] The server receives comments and questions sent from the device in real time and processes them. It receives comments using real-time communication technologies such as WebSocket and adds them to a queue.

[0795] The server takes each comment from the queue and sends it to the avatar engine, which then generates an appropriate response, which is generated as audio and video and added to the live stream.

[0796] Emotion Recognition Phase

[0797] The server uses an emotion engine to analyze comments and questions entered by users. The emotion engine recognizes the user's emotions (e.g., joy, sadness, anger, etc.) from the text contained in the comments and questions.

[0798] The server adjusts the response of the generated avatar appropriately based on the user's emotion recognized by the emotion engine. For example, if the user expresses anger, the avatar's response will be calm.

[0799] Display and Feedback Phase

[0800] The terminal provides a method for visually displaying the user's emotional information recognized by the emotion engine, allowing the user to confirm that their own emotional state is reflected.

[0801] Specific examples

[0802] For example, suppose a news site experiences a sudden spike in clicks on a particular article. The server periodically retrieves data from the news site to identify popular articles. It then generates a text script from the article's title and text, and an avatar reads it aloud with audio and video. The generated video is then streamed to the user in real time, who can then type in a comment field, "I'd like to know more about the background to this news story."

[0803] The server receives the comment and recognizes the "interest" in the emotion engine. The avatar's response generated by the avatar engine provides detailed information in a way that attracts the user's interest.

[0804] This system not only allows users to receive news visually and audibly, but also allows for real-time interactive communication, and further enhances the user experience by recognizing users' emotions and providing responses accordingly.

[0805] The processing flow will be explained below.

[0806] Step 1:

[0807] The server periodically sends requests to an external source (such as a news service API) to retrieve the latest news data. The request is sent securely using an API key or authentication information, and the news data is received in JSON format.

[0808] Step 2:

[0809] The server stores the acquired news data in a database. The news data includes information such as the title, text, number of clicks, and publication date and time, and is stored in the database using the INSERT statement.

[0810] Step 3:

[0811] The server sorts the news data from the database, executes a SELECT statement based on specific parameters (number of clicks and publication date), and sorts the news articles by most clicked.

[0812] Step 4:

[0813] The server generates a text script for reading out the selected news data, formats it as "Title: ●●, Body: ●●", and sends the generated text script to the avatar engine.

[0814] Step 5:

[0815] The server receives the avatar video generated by the avatar engine, uploads it to the streaming server, encodes the video file, and generates a stream URL that users can access.

[0816] Step 6:

[0817] The terminal provides a user interface, including a video player, a comment input field, and a question button, which users use to watch videos and input comments and questions.

[0818] Step 7:

[0819] The server receives comments and questions sent from the device in real time and adds the comments to a queue using real-time communication technologies such as WebSocket.

[0820] Step 8:

[0821] The server takes each comment from the queue and sends it to the avatar engine, which then generates an appropriate response, which is then generated as audio and video and added to the live stream.

[0822] Step 9:

[0823] The server recognizes emotions from comments and questions submitted by users using an emotion engine, which extracts the user's emotions (e.g., joy, sadness, anger) from the text.

[0824] Step 10:

[0825] The server adjusts the response generated by the avatar engine based on the user's emotion recognized by the emotion engine, for example, if the user expresses anger, the avatar's response will be in a calm tone.

[0826] Step 11:

[0827] The device visually displays the user's emotional information recognized by the emotion engine, allowing the user to confirm that their own emotional state is reflected.

[0828] Step 12:

[0829] The user receives real-time responses from the avatar and can continue the interactive communication by entering further comments or questions, a process that is repeated to improve the user experience.

[0830] Example 2

[0831] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0832] Conventional news delivery systems are limited to viewing news and have limited user interaction. Furthermore, they do not provide responses that take into account the user's emotions, resulting in a uniform user experience and low satisfaction for some users. Furthermore, they often lack interactivity because responses to real-time comments and questions are not always prompt.

[0833] The specific processing by the specific processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for acquiring data from an external information source, means for saving the acquired data, means for selecting information from the saved data based on specific parameters, means for generating audio and video from the selected information, means for distributing the generated audio and video, means for processing comments and questions received in real time during distribution and generating responses, and means for performing sentiment analysis on the received comments and questions and adjusting responses based on the results. This increases interactivity in news distribution and enables responses that correspond to the user's emotions. It also enhances real-time interaction with users and improves user satisfaction.

[0834] "External sources" refers to sources for obtaining data from outside the system, such as news service APIs and websites.

[0835] "Means for obtaining data" refers to a combination of programming and hardware for sending requests to external sources and receiving the required data.

[0836] "Means for storing data" refers to a database management system or storage device for efficiently storing acquired data.

[0837] "Means for filtering information from data based on specific parameters" refers to programs that filter stored data using criteria such as the number of clicks or publication date and time to extract the desired information.

[0838] "Audio and video generation means" means a program that converts selected information into a text script and uses speech synthesis and animation techniques to make it visually and audibly reproducible.

[0839] "Means for delivering audio and video" refers to the streaming server and encoding technology used to deliver the generated audio and video to users.

[0840] "Means for processing comments and questions received in real time and generating responses" refers to programs and communication technologies for collecting and analyzing user input in real time and generating appropriate responses based on the results.

[0841] "Means for analyzing emotions and adjusting responses based on the results" refers to a program that analyzes the emotions contained in comments and questions from users and generates an appropriate response based on those emotions.

[0842] To implement the present invention, several hardware and software components are required, which will be described below with specific examples.

[0843] Data Acquisition Phase

[0844] The server is responsible for obtaining data from external sources. Specifically, it periodically sends requests to the news service's API to obtain the latest news data. This is done by securely sending requests using API keys and authentication information, and receiving news data in JSON format or similar. For example, the server uses the NewsAPI to obtain the latest 50 articles at a time.

[0845] Data storage phase

[0846] The server stores the acquired news data in a database system (e.g., MySQL, PostgreSQL). News data includes information such as the title, text, number of clicks, and publication date and time. The server inserts this data into a database to efficiently manage it. For example, the news title "Breaking News," its text, number of clicks "325," and publication date and time "2023-10-01" are stored in the database.

[0847] News selection phase

[0848] When filtering news data from the database, the server selects news articles based on certain parameters, sorts the news articles based on the number of clicks or publication date, or sorts the articles by most clicked or most recently published, and stores the filtered data in memory. For example, it selects the top 10 articles with the most clicks.

[0849] Avatar generation phase

[0850] The server generates a text script based on the selected news data and sends the script to an avatar engine (e.g., FaceRig, Voki). The avatar engine uses speech synthesis technology (e.g., Google Text-to-Speech API) and animation technology to generate avatar videos with natural voices and movements. For example, it generates a video of an avatar reading a news article based on the text script.

[0851] Video distribution phase

[0852] The server uploads the generated avatar video to a streaming server (e.g., Wowza Streaming Engine), encodes it, and prepares it for distribution. Once preparation for distribution is complete, a stream URL is generated that users can access. For example, "https: / / streamingserver.com / newsXstream" is provided to users.

[0853] Interactive Phase

[0854] The device provides the user with a user interface that includes a video player, a comment input field, and a question button. The user uses these interfaces to watch videos and input comments and questions in real time. The server receives comments and questions sent from the device using real-time communication technology such as WebSocket and adds them to a queue. For example, a user might input a comment such as, "I'd like to know more about the background of this news."

[0855] Emotion Recognition Phase

[0856] The server uses an emotion engine (e.g., IBM Watson Tone Analyzer) to analyze the user's comments and questions and recognize the user's emotions. Based on this analysis, the server adjusts the tone and content of the response appropriately. For example, if the user's comments are analyzed as indicating "interest," the avatar's response will be adjusted to be more detailed and interesting.

[0857] Display and Feedback Steps

[0858] The device provides an interface that visually displays the user's emotional information recognized by the emotion engine, allowing the user to confirm that their emotional state is accurately reflected. For example, an icon displayed next to the comment field represents the user's emotion.

[0859] Examples:

[0860] For example, if a news site experiences a sudden spike in clicks on a particular article, the server periodically retrieves data from that news site to identify popular articles. Next, a text script is generated from the article's title and text, and an avatar reads it aloud using audio and video. The generated video is then streamed to the user in real time, who can then type in a comment field, "I'd like to know more about the background to this news story."

[0861] The server receives the comments and recognizes "interest" using the emotion engine. The avatar engine generates responses from avatars, providing detailed information in a way that piques the user's interest. This system allows users to not only receive news visually and audibly, but also enjoy real-time interactive communication.

[0862] Example prompt sentence:

[0863] "Generate text scripts from the titles and text of breaking news articles and create videos of avatars reading them. Also, generate responses to user comments and respond in the right tone based on sentiment analysis."

[0864] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0865] Step 1: Data Acquisition Step

[0866] The server retrieves data from external information sources. Specifically, it periodically sends requests to the news service's API. The input is a request that includes the API key and authentication information. The server sends the request, and the data it retrieves is news data in JSON format. Specifically, the server accesses the News API and sends a request saying, "Please retrieve the latest 50 news articles." In response, it receives the JSON data of the news articles.

[0867] Step 2: Data storage step

[0868] The server saves the acquired news data in a database. The input is the JSON-formatted news data acquired in step 1. The server parses the JSON data, extracts fields such as the news article title, text, number of clicks, and publication date and time, and inserts them into a database (e.g., MySQL, PostgreSQL). The output is the news data saved in the database. Specifically, the news title "Breaking News", text "Detailed Content", number of clicks "325", and publication date and time "2023-10-01" are saved in the database.

[0869] Step 3: News selection step

[0870] The server selects news data from the database. The input is the news data stored in the database. The server sorts the news articles based on certain parameters (e.g., number of clicks or publication date) and selects the latest and most clicked news articles. The output is the selected news data. Specifically, it selects the top 10 articles with the most clicks. The data is temporarily stored in memory.

[0871] Step 4: Avatar generation step

[0872] The server generates a text script based on the selected news data. The input is the news data selected in step 3. The server combines the news title and text to generate a text script and sends it to the avatar engine. The output is an avatar video that reads the text script aloud. Specifically, it generates a text script that reads the news title "Breaking News" and its text aloud and sends it to the avatar engine. The avatar engine generates the video using speech synthesis and animation technologies.

[0873] Step 5: Video distribution step

[0874] The server uploads the generated avatar video to the streaming server. The input is the avatar video generated in step 4. The server uploads the video to the streaming server (e.g., Wowza Streaming Engine), encodes it, and prepares it for distribution. The output is a stream URL that users can access. Specifically, the generated avatar video is uploaded to "https: / / streamingserver.com / newsXstream" so that users can access it.

[0875] Step 6: Interactive Step

[0876] The terminal provides an interface to the user. The input is the stream URL from the streaming server. The terminal displays a video player, a comment input field, a question button, etc. to the user. The user uses these interfaces to watch the video and input comments and questions in real time. The output is comments and questions from the user. Specifically, the user inputs a comment such as "I'd like to know more about the background of this news."

[0877] Step 7: Emotion Recognition Step

[0878] The server analyzes the user's comments and questions using an emotion engine. The input is the comment or question received from the user in step 6. The server uses an emotion engine (e.g., IBM Watson Tone Analyzer) to recognize the user's emotions (e.g., interest, joy, sadness, anger). The output is the emotion analysis result and an adjusted response. Specifically, if the user's comment is recognized as indicating "interest," the server adjusts the avatar's response to be more detailed and interesting.

[0879] Step 8: Display and Feedback Step

[0880] The device visually displays the user's emotional information. The input is the emotion analysis result obtained in step 7. The device displays an icon that expresses the user's emotion next to the comment field. The output is the visual display of the user's emotional information. Specifically, emotion icons (e.g., smiley, tearful, angry expressions) are displayed in the comment field, allowing the user to confirm that their emotions are accurately reflected.

[0881] (Application example 2)

[0882] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0883] In conventional news delivery systems, news acquisition, selection, and delivery are one-way, limiting user interaction. Furthermore, it is difficult to generate appropriate responses based on user emotions, making it difficult to improve the user experience. In particular, the lack of real-time emotion recognition and its reflection can lead to a decline in user engagement.

[0884] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for acquiring data from an external information source, means for saving the acquired data, means for selecting information from the saved data based on specific parameters, means for generating the selected information as audio and video, means for distributing the generated audio and video, means for processing comments and questions received in real time during distribution and generating a response, means for recognizing the user's emotions from the comments and questions and adjusting the response, and means for visually displaying the results of the emotion recognition. This enables interactive communication in real time while watching the news, and makes it possible to provide an appropriate response according to the user's emotions.

[0885] "External sources" are data sources such as news services, APIs, and data feeds available on the Internet.

[0886] "Data Acquisition Methods" are the systems and processes used to periodically request and acquire required information from external sources.

[0887] "Means for storing data" refers to a database or storage device for temporarily or permanently storing acquired information.

[0888] "Information filtering means" refers to algorithms or programs that extract important data from stored information based on specific conditions.

[0889] The "means for generating audio and video" is a system that creates visual and auditory content based on text information using speech synthesis technology and video generation technology.

[0890] The "means for delivering audio and video" refers to a server or network configuration for streaming the generated media content to users.

[0891] A "means for processing comments and questions received in real time and generating responses" is a program or process that receives input from users in real time and generates appropriate responses based on that input.

[0892] "Means for recognizing user emotions from comments and questions and adjusting responses" refers to emotion recognition engines and algorithms that analyze emotions from input text and adjust response content based on the results.

[0893] The "means for visually displaying the results of emotion recognition" refers to a function or program for visually displaying the analyzed emotion information on a user interface.

[0894] To implement this invention, a system including a server, a user terminal, a generative AI model, etc. Specifically, the system is realized by the following steps.

[0895] First, the server periodically retrieves data from external sources, using API keys and authentication information to securely retrieve data, such as from a news service API. The retrieved data is received in JSON format.

[0896] The acquired data is then stored in a database on the server, which stores information such as the news article title, text, number of clicks, and publication date and time, allowing for efficient management of the data required for subsequent processing.

[0897] The server then sorts through the vast amount of news articles stored in the database, using criteria such as click counts and publication dates to select the most interesting news articles.

[0898] Based on the selected news data, the server generates a text script for reading. This script is sent to an avatar engine, which uses speech synthesis and animation technologies to create an avatar video that generates natural-looking movements and voices. The avatar engine can be Amazon Polly or Google Text-to-Speech.

[0899] The generated avatar video is uploaded to a streaming server, where it is encoded and prepared for distribution. Once preparation is complete, a stream URL is generated that users can access.

[0900] The user terminal provides a user interface including a video player, a comment input field, a question button, etc. The user can use this interface to watch the video and input comments and questions in real time.

[0901] The server receives comments and questions sent from the device in real time and processes them using WebSocket technology, etc. Comments and questions are added to a queue, then sequentially retrieved and sent to the avatar engine, where the avatar generates an appropriate response.

[0902] In the emotion recognition phase, the server uses an emotion engine to analyze the comments and questions entered by the user. The emotion engine uses IBM Watson Natural Language Understanding and other technologies. It recognizes emotions from the user's input text and adjusts the response accordingly. For example, if the user expresses anger, the avatar's response will be calmer.

[0903] The results of emotion recognition are visually displayed on the user's device, allowing them to see responses that reflect their own emotional state, enabling more personalized, real-time communication.

[0904] Specific examples

[0905] For example, a "news avatar" reads the latest news article, and a viewer enters a comment such as, "I'd like to know more about the background to this news." The server receives this comment in real time and uses an emotion recognition engine to identify "interest." The avatar engine then generates a response such as, "Let me explain the details of this news story," providing information in a way that piques the viewer's interest. This process allows users to enjoy real-time interactive communication while watching the news broadcast.

[0906] Prompt Sentence Examples

[0907] "A news avatar will read the latest news to viewers, and users can enter comments to instantly respond. Emotion recognition will be performed for each comment, and the avatar will respond appropriately. For example, create a system that responds to a comment such as 'I'd like to know more about the background of this news story,' with 'I'll explain the details of this news story.'"

[0908] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0909] Step 1:

[0910] The server accesses an external information source (the API of a news service) to obtain the latest news data. At this time, it securely requests the data using an API key and authentication information, and receives the news data in JSON format. The input is the news data received from the API, and the output is raw news data that is temporarily stored inside the server.

[0911] Step 2:

[0912] The server stores the acquired news data in a database, which stores information such as the title, text, number of clicks, and publication date and time of news articles. The input is raw news data, and the output is structured data stored in the database. This allows for efficient management of the data required for subsequent processing.

[0913] Step 3:

[0914] The server selects news data from the database based on specific parameters (e.g., number of clicks or publication date). The selected news data is temporarily stored in memory. The input is the news data stored in the database, and the output is the selected news data.

[0915] Step 4:

[0916] The server generates a text script for reading based on the selected news data. This script is sent to the avatar engine, which uses speech synthesis and animation technologies to make the avatar move and speak naturally. The input is the selected news data, and the output is the generated avatar video.

[0917] Step 5:

[0918] The server uploads the generated avatar video to the streaming server, encodes it, and prepares it for distribution. Once it is ready for distribution, a stream URL that users can access is generated. The input is the avatar video, and the output is the stream URL.

[0919] Step 6:

[0920] The user terminal provides a user interface that includes a video player, a comment input field, a question button, etc. The user uses the interface to watch videos and input comments and questions in real time. The input is the user's comments and questions, and the output is the corresponding interface operation.

[0921] Step 7:

[0922] The server receives comments and questions sent from the terminal in real time and adds the comments to a queue using WebSocket technology. The input is the comments and questions received from the user, and the output is the data added to the comment queue.

[0923] Step 8:

[0924] The server sequentially retrieves comments added to the queue and sends them to the avatar engine, where the avatar generates an appropriate response. The input is the comment retrieved from the comment queue, and the output is the generated avatar response.

[0925] Step 9:

[0926] The server analyzes the user's comments and questions using an emotion engine and adjusts the avatar's response based on the emotion. The input is the user's comments and questions, and the output is the adjusted avatar response. The emotion engine uses IBM Watson Natural Language Understanding and other technologies.

[0927] Step 10:

[0928] The user device visually displays the emotion recognition results, allowing the user to confirm the response that reflects their own emotional state. The input is the analysis result of the emotion engine, and the output is the visually displayed emotional information.

[0929] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0930] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0931] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[0932] [Fourth embodiment]

[0933] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[0934] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0935] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0936] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[0937] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0938] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0939] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0940] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[0941] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0942] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0943] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0944] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0945] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[0946] To implement the present invention, a system including a plurality of components is required, which will be described below with specific examples.

[0947] Data Acquisition Phase

[0948] The server periodically sends requests to external information sources (such as the API of a news service) to obtain the latest news data. The server securely sends the requests using API keys and authentication information. The obtained news data is often received in JSON format or similar.

[0949] Data storage phase

[0950] The server stores the acquired news data in a database. The news data includes information such as the title, text, number of clicks, and publication date and time. This allows for efficient management of data required for subsequent processing.

[0951] News selection phase

[0952] The server sifts through the news data from the database, selecting the latest and most interesting news articles based on certain parameters such as click count, publication date, etc. The sift data is temporarily stored in memory for further processing.

[0953] Avatar generation phase

[0954] The server generates a text script for reading based on the selected news data. The generated script is sent to the avatar engine, which then generates a video of an avatar reading the script. The avatar engine uses speech synthesis and animation technologies to generate natural-looking movements and voices.

[0955] Video distribution phase

[0956] The server uploads the generated avatar video to the streaming server, encodes it, and prepares it for distribution. Once preparation is complete, a stream URL is generated that users can access.

[0957] Interactive Phase

[0958] The terminal provides the user with a user interface, which includes a video player, a comment input field, a question button, etc. When the user uses this interface to input comments or questions in real time, they are sent to the server.

[0959] The server receives and processes comments and questions from users in real time, sends the processed information to the avatar engine, and the avatar generates an appropriate response, which is then re-integrated into the video stream and delivered to the user.

[0960] Specific examples

[0961] For example, suppose a news site experiences a sudden spike in clicks on a particular article. The server periodically retrieves data from the news site and detects this popular article. Next, it generates a text script from the article's title and text, and an avatar reads it aloud using audio and video. The generated video is delivered to the user in real time, and the user can enter "I'd like to know more about the background to this news story" in the comment input field. The server receives the comment and sends it to the avatar engine, where the avatar generates and delivers a response such as "The background to this news story is..."

[0962] This system allows users to receive news visually and audibly, allowing them to enjoy real-time interactive communication.

[0963] The processing flow will be explained below.

[0964] Step 1:

[0965] The server periodically sends requests to an external source (such as a news service API) to obtain the latest news data. The request is sent securely using an API key or authentication information, and the news data is received in JSON format or other formats.

[0966] Step 2:

[0967] The server stores the acquired news data in a database. The news data includes information such as the title, text, number of clicks, and publication date and time, and is stored in the database using the INSERT statement.

[0968] Step 3:

[0969] The server sifts through the news data from the database, sorts the news articles by most clicked based on certain parameters (e.g., number of clicks, publication date), and selects the most recent and likely most interesting news articles.

[0970] Step 4:

[0971] The server generates a text script for reading out the selected news data, formats it into a format such as "Title: ●●, Body: ●●", and sends the generated text script to the avatar engine.

[0972] Step 5:

[0973] The server receives the avatar video generated by the avatar engine and uploads it to the streaming server, where it encodes the video file and prepares it for distribution.

[0974] Step 6:

[0975] The device provides the user with a user interface that includes a video player, a comment input field, a question button, etc. The user uses this interface to watch videos and input comments and questions in real time.

[0976] Step 7:

[0977] The server receives comments and questions sent from the device in real time and processes them. It receives comments using real-time communication technologies such as WebSocket and adds them to a queue.

[0978] Step 8:

[0979] The server takes each comment from the queue and sends it to the avatar engine, which then generates an appropriate response, which is generated as audio and video and added to the live stream.

[0980] Step 9:

[0981] The user receives real-time responses from the avatar via a live stream, providing a visual and auditory experience, and can enter further comments or questions as needed.

[0982] Example 1

[0983] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[0984] There is a need to efficiently manage data obtained from external information sources and quickly select important information based on specific parameters. It is also important to deliver this information in a way that users can perceive visually and audibly in a natural way. Furthermore, there is a need for a system that can respond immediately to comments and questions received in real time during the broadcast, enabling interactive communication. However, current systems have difficulty fully meeting these requirements, and technical challenges remain, particularly in real-time processing and audio-video integration.

[0985] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0986] In this invention, the server includes: means for acquiring data from external information sources; means for saving the acquired data; means for selecting information from the saved data based on specific parameters; means for generating audio and video from the selected information; means for distributing the generated audio and video; means for processing comments and questions received in real time during distribution and generating responses; means for selecting news data from data stored in a database and generating videos read by avatars using speech synthesis and animation technologies; means for encoding and distributing avatar videos uploaded to a streaming service; and means for processing user comments and questions on the encoded videos in real time and generating and redistributing avatar responses. This allows users to naturally receive important information visually and audibly, enabling further interactive communication in real time.

[0987] "External information sources" are resources for obtaining data provided by the Internet or specific service providers.

[0988] "Data acquisition means" are the methods and techniques used to collect information from external sources.

[0989] "Means for storing data" refers to the methods and techniques used to hold acquired information in a database or other storage device.

[0990] "Specific parameters" are numerical values ​​or conditions that serve as criteria when selecting or processing data.

[0991] "Means for selecting information" are methods and techniques for extracting necessary information from stored data.

[0992] "Means for generating audio and video" refers to methods and technologies for synthesizing audio and producing video based on selected information.

[0993] "Means of distribution" refers to the methods and technologies used to transmit the generated audio and video to users over the Internet.

[0994] "Means for processing comments and questions received in real time and generating responses" refers to methods and technologies for receiving immediate feedback from users during a broadcast and generating responses based on that feedback.

[0995] A "database" is a system or software for efficiently managing and manipulating large amounts of data.

[0996] "News data" refers to the latest information obtained from news services.

[0997] "Speech synthesis technology" is a technology for artificially generating speech based on text data.

[0998] "Animation technology" is a technique that makes still images appear to be moving by displaying them in succession.

[0999] An "avatar" is a computer-generated virtual representation of a person or character.

[1000] A "streaming service" is a service for delivering digital content in real time over the Internet.

[1001] "Encoding" is the process of converting digital data into a particular format.

[1002] A "user interface" refers to the screen and operation method that allows a user to interact with a system.

[1003] MODE FOR CARRYING OUT THE INVENTION

[1004] To implement this invention, a system including multiple components is required. This system mainly consists of three elements: a server, a terminal, and a user.

[1005] First, the server periodically sends requests to an external information source (for example, the API of a news service) to obtain the latest news data. This request uses HTTPS communication, and an API key or OAuth token is used to ensure security. The obtained news data is typically returned in JSON format. An example request is "GET / latest-news?apiKey=YOUR_API_KEY".

[1006] Next, the server stores the retrieved news data in a database (e.g., PostgreSQL). The database contains the title, content, click count, and publication date / time of each news article. The server efficiently organizes this information and stores it in a structured format. For example, the following query is used: "INSERT INTO news_articles (title, content, clicks, published_at) VALUES ('Example News', 'Example News Content', 120, '2023-10-05 14:48:00');"

[1007] The server selects news data from the database based on certain parameters, such as the articles with the most clicks or the most recent publication date. This selection is performed using an SQL query, such as "SELECT FROM news_articles ORDER BY clicks DESC LIMIT 1;".

[1008] Based on the selected news data, the server generates a text script for reading aloud. This is then used with speech synthesis and animation technology to generate a video in which an avatar reads the script aloud. An avatar engine (e.g., Unity or Unreal Engine) is used for this process. For example, the generated script might read, "We will report the next news item. The title is 'Example News'."

[1009] The generated avatar video is uploaded to a streaming service (e.g., Wowza Streaming Engine) by the server. The server encodes the video and prepares it for distribution. Once preparation for distribution is complete, a stream URL is generated that users can access. For example, the server sends a request called "POST / upload" and receives the stream URL as a response.

[1010] The device provides the user with a user interface (UI). This UI includes a video player, a comment input field, a question button, and so on. The user uses these interfaces to input comments and questions in real time. For example, the user might input a comment such as "I'd like to know more about the background of this news," and the device sends the comment to the server. A request called "POST / comments" is sent, and the comment content is included in the request.

[1011] The server receives comments and questions from users in real time and processes them. The processed information is sent back to the avatar engine, where the avatar generates an appropriate response. The generated response is incorporated into the video stream and delivered to the user. For example, the server passes a script such as "Some background on this news..." to the avatar engine to generate a new video.

[1012] As an example of a prompt sentence, the text "Tell me about the latest news" is input to the generative AI model. Based on this prompt sentence, the entire system operates, carrying out a series of processes from news information collection to distribution and interactive response.

[1013] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1014] Step 1:

[1015] The server periodically sends requests to external news sources to retrieve the latest news data. The input is the API endpoint and authentication information (e.g., API key) of the external news source. The server sends the request "GET / latest-news?apiKey=YOUR_API_KEY" and receives the news data in JSON format. This data includes the title, body, number of clicks, publication date, etc. of the news article.

[1016] Step 2:

[1017] The server saves the acquired news data in a database. The input is news data in JSON format. The server organizes this data and stores it in a database (e.g., PostgreSQL). Specifically, it executes the query "INSERT INTO news_articles (title, content, clicks, published_at) VALUES ('Example news', 'Example news content', 120, '2023-10-05 14:48:00');". The output is the news data saved in the database.

[1018] Step 3:

[1019] The server selects news data from a database based on specific parameters. The input is the news data stored in the database. The server selects the most relevant news articles based on parameters such as the number of clicks and publication date, and temporarily stores them in memory. For example, it executes the query "SELECT FROM news_articles ORDER BY clicks DESC LIMIT 1;". The output is the selected news data.

[1020] Step 4:

[1021] The server generates a text script for reading aloud based on the selected news data. The input is the selected news data. The server extracts the title and content from the news data and generates a text script that says, "We will report the next news item. The title is 'Example News'." The output is a text script for reading aloud.

[1022] Step 5:

[1023] The server sends this text script to the avatar engine, which uses speech synthesis and animation technologies to generate a video in which an avatar reads the script. The input is the text script. The avatar engine (e.g., Unity or Unreal Engine) generates a video based on the passed script. The output is the generated avatar video.

[1024] Step 6:

[1025] The server uploads the generated avatar video to the streaming service, encodes it, and prepares it for distribution. The input is the generated avatar video. Specifically, it sends a "POST / upload" request to the streaming service (e.g., Wowza Streaming Engine). The output is the stream URL that users can access.

[1026] Step 7:

[1027] The device provides a user interface (UI) to the user. The input is the stream URL. This UI includes a video player, a comment input field, a question button, and so on. The user inputs comments and questions in real time through this interface. Specifically, the user inputs "I'd like to know more about the background of this news," and the device sends a "POST / comments" request to the server. The output is the user's comments and questions.

[1028] Step 8:

[1029] The server receives and processes comments and questions from users in real time. The input is the user's comment or question. The server sends the received comment to the avatar engine, and the avatar generates an appropriate response. For example, a script such as "About the background of this news..." is passed to the avatar engine to generate a new video. The output is the generated avatar's response video.

[1030] Step 9:

[1031] The server uploads the generated response video back to the streaming service and delivers it to the user. The input is the avatar's response video. The server uploads new videos to the streaming service, updating the delivery in real time. The output is the updated video stream.

[1032] (Application example 1)

[1033] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1034] In recent years, with the digitalization of information and the spread of the Internet, methods of news distribution have become more diverse. However, systems that allow users to obtain information interactively in real time are insufficient, and there is a need for more information to be provided visually and audibly. In addition to allowing users to efficiently browse a wide range of news, there is also a need for systems that can provide information based on the user's interests.

[1035] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1036] In this invention, the server includes means for acquiring data from external information sources, means for storing the acquired data, means for selecting information from the stored data based on specific parameters, means for generating the selected information as audio and video, means for distributing the generated audio and video, means for processing comments and questions received in real time during distribution and generating responses, and means for providing the data on a smartphone application, thereby enabling users to watch the news interactively in real time and receive information visually and audibly.

[1037] "External information sources" refer to information sources that exist outside the system, such as online news services or APIs.

[1038] "Means for Obtaining Data" refers to methods or techniques for periodically requesting and receiving data from external sources.

[1039] "Means for storing data" refers to the methods or techniques for storing acquired data in a database or other storage medium.

[1040] "Information screening means" refers to methods or technologies for extracting selected information from stored data based on specific criteria (e.g., number of clicks or publication date).

[1041] "Means for generating audio and video" refers to technologies for creating audio and video from selected information, in particular methods for using speech synthesis and animation technologies to read aloud using an avatar.

[1042] "Delivery Means" means the method or technology by which the generated audio and video is made available to users in streaming format.

[1043] "Means for processing comments and questions and generating responses" refers to a method or technology for receiving comments and questions sent in real time by users during a broadcast and generating and returning appropriate responses to them.

[1044] "Smartphone Application" means software that runs on a smartphone device and that provides users with the ability to interactively browse news.

[1045] "Real-time comments and questions" refers to comments and questions submitted by users during a video broadcast, and refers to information processed to respond to such comments and questions in a timely manner.

[1046] "Means for processing comments and questions received in real time during the broadcast and generating responses" refers to a method for instantly analyzing feedback from users during a live broadcast and generating appropriate responses.

[1047]

[1048] To implement the present invention, a system including a plurality of components is required, which will be described below with specific examples.

[1049] Data Acquisition Phase

[1050] The server periodically sends requests to external information sources (such as the API of a news service) to obtain the latest news data. The server securely sends the requests using API keys and authentication information. The obtained news data is often received in JSON format or similar.

[1051] Data storage phase

[1052] The server stores the acquired news data in a database. The news data includes information such as the title, text, number of clicks, and publication date and time. This allows for efficient management of data required for subsequent processing.

[1053] News selection phase

[1054] The server sifts through the news data from the database, selecting the latest and most interesting news articles based on certain parameters such as click count, publication date, etc. The sift data is temporarily stored in memory for further processing.

[1055] Avatar generation phase

[1056] The server generates a text script for reading based on the selected news data. The generated script is sent to the avatar engine, which then generates a video of an avatar reading the script. The avatar engine uses speech synthesis and animation technologies to generate natural-looking movements and voices.

[1057] Video distribution phase

[1058] The server uploads the generated avatar video to the streaming server, encodes it, and prepares it for distribution. Once preparation is complete, a stream URL is generated that users can access. Users can watch the video through this URL on their smartphone application.

[1059] Interactive Phase

[1060] Users can enter comments and questions in real time using the smartphone application, including the video player, comment input field, and question button. The server receives and processes comments and questions from users in real time. The processed information is sent to the avatar engine, where the avatar generates an appropriate response. The generated response is then incorporated back into the video stream and delivered to the user.

[1061] For example, suppose a news site experiences a sudden spike in clicks on a particular article. The server periodically retrieves data from the news site and detects this popular article. Next, it generates a text script from the article's title and text, and an avatar reads it aloud using audio and video. The generated video is then delivered to the user in real time. The user can then enter "I'd like to know more about the background to this news story" in the comment input field. The server receives the comment, sends it to the avatar engine, and the avatar generates and delivers a response such as "The background to this news story is..."

[1062] Example prompt sentence:

[1063] "News title: {title}

[1064] News text: {text}

[1065] Please use this as a basis to generate a video of the avatar reading the text."

[1066] This system allows users to receive news visually and audibly, allowing them to enjoy real-time interactive communication.

[1067] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1068] Step 1: Data Acquisition Phase

[1069] The server sends an HTTP request to the news service's API to obtain the latest news data. The request header, including the API key and authentication information, is used as input. The response from the API receives news data in JSON format. The data includes the article title, text, number of clicks, publication date, and so on.

[1070] Step 2: Data storage phase

[1071] The server stores the retrieved news data in a local database. As input, it uses each field of the received news data (title, body, number of clicks, publication date, etc.). As output, each news article is stored in the database. Specifically, it adds data to the database using the SQL INSERT statement.

[1072] Step 3: News selection phase

[1073] The server selects news data from the database. As input, it uses the news data stored in the database. The selection criteria are specific parameters such as the number of clicks or publication date. As output, it selects the most interesting news articles. Specifically, it uses a SQL SELECT statement to filter the articles that match the criteria.

[1074] Step 4: Avatar generation phase

[1075] The server generates a text script for reading aloud based on the selected news data. The selected news data (title, body text) is used as input. The text script is generated as output. The generated script is then sent to the avatar engine, which generates audio and video for the avatar to read aloud. Specifically, the server calls the avatar engine's API to convert the text into audio and animation.

[1076] Step 5: Video distribution phase

[1077] The server uploads the generated avatar video to the streaming server, encodes it, and prepares it for distribution. The generated avatar video is used as input. As output, a stream URL that can be accessed by users is generated. Specifically, the server uses a file transfer protocol to upload the video to the streaming server and generate the URL.

[1078] Step 6: User Interaction Phase

[1079] Users enter comments and questions in real time via a smartphone application using a video player, comment input field, question button, etc. The comments and questions entered by the user on the application are used as input. The input content is sent to the server as output. Specifically, the system operates by using a UI component that accepts user input and a network function that sends it to the server.

[1080] Step 7: Comment and question handling phase

[1081] The server receives and processes comments and questions submitted by users in real time. It uses the comments and questions submitted by users as input. As output, it generates an appropriate response. Specifically, it analyzes the content of the comments and questions and sends them to the avatar engine to generate a response.

[1082] Step 8: Response Delivery Phase

[1083] The server then incorporates the generated response back into the video stream and delivers it to the user. It uses the generated response as input, and delivers an updated video stream to the user as output. Specifically, it uses the streaming server's API to incorporate the response into the video and delivers the updated stream.

[1084] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1085] To implement this invention, a system including a number of components is required, which will be described below with specific examples.

[1086] Data Acquisition Phase

[1087] The server periodically sends requests to external sources (such as the API of a news service) to obtain the latest news data. The requests are sent securely using API keys and authentication information, and the news data is received in JSON format or other formats.

[1088] Data storage phase

[1089] The server stores the acquired news data in a database. The news data includes information such as the title, text, number of clicks, and publication date and time. This allows for efficient management of data required for subsequent processing.

[1090] News selection phase

[1091] The server selects news data from the database, sorts the news articles by the number of clicks based on certain parameters (such as the number of clicks or the publication date), and selects the latest and most interesting news articles. The selected data is temporarily stored in memory for further processing.

[1092] Avatar generation phase

[1093] The server generates a text script for reading based on the selected news data. The generated script is sent to the avatar engine, which then generates a video of an avatar reading the script. The avatar engine uses speech synthesis and animation technologies to generate natural-looking movements and voices.

[1094] Video distribution phase

[1095] The server uploads the generated avatar video to the streaming server, encodes it, and prepares it for distribution. Once preparation is complete, a stream URL is generated that users can access.

[1096] Interactive Phase

[1097] The device provides the user with a user interface that includes a video player, a comment input field, a question button, etc. The user uses this interface to watch videos and input comments and questions in real time.

[1098] The server receives comments and questions sent from the device in real time and processes them. It receives comments using real-time communication technologies such as WebSocket and adds them to a queue.

[1099] The server takes each comment from the queue and sends it to the avatar engine, which then generates an appropriate response, which is generated as audio and video and added to the live stream.

[1100] Emotion Recognition Phase

[1101] The server uses an emotion engine to analyze comments and questions entered by users. The emotion engine recognizes the user's emotions (e.g., joy, sadness, anger, etc.) from the text contained in the comments and questions.

[1102] The server adjusts the response of the generated avatar appropriately based on the user's emotion recognized by the emotion engine. For example, if the user expresses anger, the avatar's response will be calm.

[1103] Display and Feedback Phase

[1104] The terminal provides a method for visually displaying the user's emotional information recognized by the emotion engine, allowing the user to confirm that their own emotional state is reflected.

[1105] Specific examples

[1106] For example, suppose a news site experiences a sudden spike in clicks on a particular article. The server periodically retrieves data from the news site to identify popular articles. It then generates a text script from the article's title and text, and an avatar reads it aloud with audio and video. The generated video is then streamed to the user in real time, who can then type in a comment field, "I'd like to know more about the background to this news story."

[1107] The server receives the comment and recognizes the "interest" in the emotion engine. The avatar's response generated by the avatar engine provides detailed information in a way that attracts the user's interest.

[1108] This system not only allows users to receive news visually and audibly, but also allows for real-time interactive communication, and further enhances the user experience by recognizing users' emotions and providing responses accordingly.

[1109] The processing flow will be explained below.

[1110] Step 1:

[1111] The server periodically sends requests to an external source (such as a news service API) to retrieve the latest news data. The request is sent securely using an API key or authentication information, and the news data is received in JSON format.

[1112] Step 2:

[1113] The server stores the acquired news data in a database. The news data includes information such as the title, text, number of clicks, and publication date and time, and is stored in the database using the INSERT statement.

[1114] Step 3:

[1115] The server sorts the news data from the database, executes a SELECT statement based on specific parameters (number of clicks and publication date), and sorts the news articles by most clicked.

[1116] Step 4:

[1117] The server generates a text script for reading out the selected news data, formats it as "Title: ●●, Body: ●●", and sends the generated text script to the avatar engine.

[1118] Step 5:

[1119] The server receives the avatar video generated by the avatar engine, uploads it to the streaming server, encodes the video file, and generates a stream URL that users can access.

[1120] Step 6:

[1121] The terminal provides a user interface, including a video player, a comment input field, and a question button, which users use to watch videos and input comments and questions.

[1122] Step 7:

[1123] The server receives comments and questions sent from the device in real time and adds the comments to a queue using real-time communication technologies such as WebSocket.

[1124] Step 8:

[1125] The server takes each comment from the queue and sends it to the avatar engine, which then generates an appropriate response, which is then generated as audio and video and added to the live stream.

[1126] Step 9:

[1127] The server recognizes emotions from comments and questions submitted by users using an emotion engine, which extracts the user's emotions (e.g., joy, sadness, anger) from the text.

[1128] Step 10:

[1129] The server adjusts the response generated by the avatar engine based on the user's emotion recognized by the emotion engine, for example, if the user expresses anger, the avatar's response will be in a calm tone.

[1130] Step 11:

[1131] The device visually displays the user's emotional information recognized by the emotion engine, allowing the user to confirm that their own emotional state is reflected.

[1132] Step 12:

[1133] The user receives real-time responses from the avatar and can continue the interactive communication by entering further comments or questions, a process that is repeated to improve the user experience.

[1134] Example 2

[1135] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1136] Conventional news delivery systems are limited to viewing news and have limited user interaction. Furthermore, they do not provide responses that take into account the user's emotions, resulting in a uniform user experience and low satisfaction for some users. Furthermore, they often lack interactivity because responses to real-time comments and questions are not always prompt.

[1137] The specific processing by the specific processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for acquiring data from an external information source, means for saving the acquired data, means for selecting information from the saved data based on specific parameters, means for generating audio and video from the selected information, means for distributing the generated audio and video, means for processing comments and questions received in real time during distribution and generating responses, and means for performing sentiment analysis on the received comments and questions and adjusting responses based on the results. This increases interactivity in news distribution and enables responses that correspond to the user's emotions. It also enhances real-time interaction with users and improves user satisfaction.

[1138] "External sources" refers to sources for obtaining data from outside the system, such as news service APIs and websites.

[1139] "Means for obtaining data" refers to a combination of programming and hardware for sending requests to external sources and receiving the required data.

[1140] "Means for storing data" refers to a database management system or storage device for efficiently storing acquired data.

[1141] "Means for filtering information from data based on specific parameters" refers to programs that filter stored data using criteria such as the number of clicks or publication date and time to extract the desired information.

[1142] "Audio and video generation means" means a program that converts selected information into a text script and uses speech synthesis and animation techniques to make it visually and audibly reproducible.

[1143] "Means for delivering audio and video" refers to the streaming server and encoding technology used to deliver the generated audio and video to users.

[1144] "Means for processing comments and questions received in real time and generating responses" refers to programs and communication technologies for collecting and analyzing user input in real time and generating appropriate responses based on the results.

[1145] "Means for analyzing emotions and adjusting responses based on the results" refers to a program that analyzes the emotions contained in comments and questions from users and generates an appropriate response based on those emotions.

[1146] To implement the present invention, several hardware and software components are required, which will be described below with specific examples.

[1147] Data Acquisition Phase

[1148] The server is responsible for obtaining data from external sources. Specifically, it periodically sends requests to the news service's API to obtain the latest news data. This is done by securely sending requests using API keys and authentication information, and receiving news data in JSON format or similar. For example, the server uses the NewsAPI to obtain the latest 50 articles at a time.

[1149] Data storage phase

[1150] The server stores the acquired news data in a database system (e.g., MySQL, PostgreSQL). News data includes information such as the title, text, number of clicks, and publication date and time. The server inserts this data into a database to efficiently manage it. For example, the news title "Breaking News," its text, number of clicks "325," and publication date and time "2023-10-01" are stored in the database.

[1151] News selection phase

[1152] When filtering news data from the database, the server selects news articles based on certain parameters, sorts the news articles based on the number of clicks or publication date, or sorts the articles by most clicked or most recently published, and stores the filtered data in memory. For example, it selects the top 10 articles with the most clicks.

[1153] Avatar generation phase

[1154] The server generates a text script based on the selected news data and sends the script to an avatar engine (e.g., FaceRig, Voki). The avatar engine uses speech synthesis technology (e.g., Google Text-to-Speech API) and animation technology to generate avatar videos with natural voices and movements. For example, it generates a video of an avatar reading a news article based on the text script.

[1155] Video distribution phase

[1156] The server uploads the generated avatar video to a streaming server (e.g., Wowza Streaming Engine), encodes it, and prepares it for distribution. Once preparation for distribution is complete, a stream URL is generated that users can access. For example, "https: / / streamingserver.com / newsXstream" is provided to users.

[1157] Interactive Phase

[1158] The device provides the user with a user interface that includes a video player, a comment input field, and a question button. The user uses these interfaces to watch videos and input comments and questions in real time. The server receives comments and questions sent from the device using real-time communication technology such as WebSocket and adds them to a queue. For example, a user might input a comment such as, "I'd like to know more about the background of this news."

[1159] Emotion Recognition Phase

[1160] The server uses an emotion engine (e.g., IBM Watson Tone Analyzer) to analyze the user's comments and questions and recognize the user's emotions. Based on this analysis, the server adjusts the tone and content of the response appropriately. For example, if the user's comments are analyzed as indicating "interest," the avatar's response will be adjusted to be more detailed and interesting.

[1161] Display and Feedback Steps

[1162] The device provides an interface that visually displays the user's emotional information recognized by the emotion engine, allowing the user to confirm that their emotional state is accurately reflected. For example, an icon displayed next to the comment field represents the user's emotion.

[1163] Examples:

[1164] For example, if a news site experiences a sudden spike in clicks on a particular article, the server periodically retrieves data from that news site to identify popular articles. Next, a text script is generated from the article's title and text, and an avatar reads it aloud using audio and video. The generated video is then streamed to the user in real time, who can then type in a comment field, "I'd like to know more about the background to this news story."

[1165] The server receives the comments and recognizes "interest" using the emotion engine. The avatar engine generates responses from avatars, providing detailed information in a way that piques the user's interest. This system allows users to not only receive news visually and audibly, but also enjoy real-time interactive communication.

[1166] Example prompt sentence:

[1167] "Generate text scripts from the titles and text of breaking news articles and create videos of avatars reading them. Also, generate responses to user comments and respond in the right tone based on sentiment analysis."

[1168] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1169] Step 1: Data Acquisition Step

[1170] The server retrieves data from external information sources. Specifically, it periodically sends requests to the news service's API. The input is a request that includes the API key and authentication information. The server sends the request, and the data it retrieves is news data in JSON format. Specifically, the server accesses the News API and sends a request saying, "Please retrieve the latest 50 news articles." In response, it receives the JSON data of the news articles.

[1171] Step 2: Data storage step

[1172] The server saves the acquired news data in a database. The input is the JSON-formatted news data acquired in step 1. The server parses the JSON data, extracts fields such as the news article title, text, number of clicks, and publication date and time, and inserts them into a database (e.g., MySQL, PostgreSQL). The output is the news data saved in the database. Specifically, the news title "Breaking News", text "Detailed Content", number of clicks "325", and publication date and time "2023-10-01" are saved in the database.

[1173] Step 3: News selection step

[1174] The server selects news data from the database. The input is the news data stored in the database. The server sorts the news articles based on certain parameters (e.g., number of clicks or publication date) and selects the latest and most clicked news articles. The output is the selected news data. Specifically, it selects the top 10 articles with the most clicks. The data is temporarily stored in memory.

[1175] Step 4: Avatar generation step

[1176] The server generates a text script based on the selected news data. The input is the news data selected in step 3. The server combines the news title and text to generate a text script and sends it to the avatar engine. The output is an avatar video that reads the text script aloud. Specifically, it generates a text script that reads the news title "Breaking News" and its text aloud and sends it to the avatar engine. The avatar engine generates the video using speech synthesis and animation technologies.

[1177] Step 5: Video distribution step

[1178] The server uploads the generated avatar video to the streaming server. The input is the avatar video generated in step 4. The server uploads the video to the streaming server (e.g., Wowza Streaming Engine), encodes it, and prepares it for distribution. The output is a stream URL that users can access. Specifically, the generated avatar video is uploaded to "https: / / streamingserver.com / newsXstream" so that users can access it.

[1179] Step 6: Interactive Step

[1180] The terminal provides an interface to the user. The input is the stream URL from the streaming server. The terminal displays a video player, a comment input field, a question button, etc. to the user. The user uses these interfaces to watch the video and input comments and questions in real time. The output is comments and questions from the user. Specifically, the user inputs a comment such as "I'd like to know more about the background of this news."

[1181] Step 7: Emotion Recognition Step

[1182] The server analyzes the user's comments and questions using an emotion engine. The input is the comment or question received from the user in step 6. The server uses an emotion engine (e.g., IBM Watson Tone Analyzer) to recognize the user's emotions (e.g., interest, joy, sadness, anger). The output is the emotion analysis result and an adjusted response. Specifically, if the user's comment is recognized as indicating "interest," the server adjusts the avatar's response to be more detailed and interesting.

[1183] Step 8: Display and Feedback Step

[1184] The device visually displays the user's emotional information. The input is the emotion analysis result obtained in step 7. The device displays an icon that expresses the user's emotion next to the comment field. The output is the visual display of the user's emotional information. Specifically, emotion icons (e.g., smiley, tearful, angry expressions) are displayed in the comment field, allowing the user to confirm that their emotions are accurately reflected.

[1185] (Application example 2)

[1186] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1187] In conventional news delivery systems, news acquisition, selection, and delivery are one-way, limiting user interaction. Furthermore, it is difficult to generate appropriate responses based on user emotions, making it difficult to improve the user experience. In particular, the lack of real-time emotion recognition and its reflection can lead to a decline in user engagement.

[1188] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for acquiring data from an external information source, means for saving the acquired data, means for selecting information from the saved data based on specific parameters, means for generating the selected information as audio and video, means for distributing the generated audio and video, means for processing comments and questions received in real time during distribution and generating a response, means for recognizing the user's emotions from the comments and questions and adjusting the response, and means for visually displaying the results of the emotion recognition. This enables interactive communication in real time while watching the news, and makes it possible to provide an appropriate response according to the user's emotions.

[1189] "External sources" are data sources such as news services, APIs, and data feeds available on the Internet.

[1190] "Data Acquisition Methods" are the systems and processes used to periodically request and acquire required information from external sources.

[1191] "Means for storing data" refers to a database or storage device for temporarily or permanently storing acquired information.

[1192] "Information filtering means" refers to algorithms or programs that extract important data from stored information based on specific conditions.

[1193] The "means for generating audio and video" is a system that creates visual and auditory content based on text information using speech synthesis technology and video generation technology.

[1194] The "means for delivering audio and video" refers to a server or network configuration for streaming the generated media content to users.

[1195] A "means for processing comments and questions received in real time and generating responses" is a program or process that receives input from users in real time and generates appropriate responses based on that input.

[1196] "Means for recognizing user emotions from comments and questions and adjusting responses" refers to emotion recognition engines and algorithms that analyze emotions from input text and adjust response content based on the results.

[1197] The "means for visually displaying the results of emotion recognition" refers to a function or program for visually displaying the analyzed emotion information on a user interface.

[1198] To implement this invention, a system including a server, a user terminal, a generative AI model, etc. Specifically, the system is realized by the following steps.

[1199] First, the server periodically retrieves data from external sources, using API keys and authentication information to securely retrieve data, such as from a news service API. The retrieved data is received in JSON format.

[1200] The acquired data is then stored in a database on the server, which stores information such as the news article title, text, number of clicks, and publication date and time, allowing for efficient management of the data required for subsequent processing.

[1201] The server then sorts through the vast amount of news articles stored in the database, using criteria such as click counts and publication dates to select the most interesting news articles.

[1202] Based on the selected news data, the server generates a text script for reading. This script is sent to an avatar engine, which uses speech synthesis and animation technologies to create an avatar video that generates natural-looking movements and voices. The avatar engine can be Amazon Polly or Google Text-to-Speech.

[1203] The generated avatar video is uploaded to a streaming server, where it is encoded and prepared for distribution. Once preparation is complete, a stream URL is generated that users can access.

[1204] The user terminal provides a user interface including a video player, a comment input field, a question button, etc. The user can use this interface to watch the video and input comments and questions in real time.

[1205] The server receives comments and questions sent from the device in real time and processes them using WebSocket technology, etc. Comments and questions are added to a queue, then sequentially retrieved and sent to the avatar engine, where the avatar generates an appropriate response.

[1206] In the emotion recognition phase, the server uses an emotion engine to analyze the comments and questions entered by the user. The emotion engine uses IBM Watson Natural Language Understanding and other technologies. It recognizes emotions from the user's input text and adjusts the response accordingly. For example, if the user expresses anger, the avatar's response will be calmer.

[1207] The results of emotion recognition are visually displayed on the user's device, allowing them to see responses that reflect their own emotional state, enabling more personalized, real-time communication.

[1208] Specific examples

[1209] For example, a "news avatar" reads the latest news article, and a viewer enters a comment such as, "I'd like to know more about the background to this news." The server receives this comment in real time and uses an emotion recognition engine to identify "interest." The avatar engine then generates a response such as, "Let me explain the details of this news story," providing information in a way that piques the viewer's interest. This process allows users to enjoy real-time interactive communication while watching the news broadcast.

[1210] Prompt Sentence Examples

[1211] "A news avatar will read the latest news to viewers, and users can enter comments to instantly respond. Emotion recognition will be performed for each comment, and the avatar will respond appropriately. For example, create a system that responds to a comment such as 'I'd like to know more about the background of this news story,' with 'I'll explain the details of this news story.'"

[1212] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1213] Step 1:

[1214] The server accesses an external information source (the API of a news service) to obtain the latest news data. At this time, it securely requests the data using an API key and authentication information, and receives the news data in JSON format. The input is the news data received from the API, and the output is raw news data that is temporarily stored inside the server.

[1215] Step 2:

[1216] The server stores the acquired news data in a database, which stores information such as the title, text, number of clicks, and publication date and time of news articles. The input is raw news data, and the output is structured data stored in the database. This allows for efficient management of the data required for subsequent processing.

[1217] Step 3:

[1218] The server selects news data from the database based on specific parameters (e.g., number of clicks or publication date). The selected news data is temporarily stored in memory. The input is the news data stored in the database, and the output is the selected news data.

[1219] Step 4:

[1220] The server generates a text script for reading based on the selected news data. This script is sent to the avatar engine, which uses speech synthesis and animation technologies to make the avatar move and speak naturally. The input is the selected news data, and the output is the generated avatar video.

[1221] Step 5:

[1222] The server uploads the generated avatar video to the streaming server, encodes it, and prepares it for distribution. Once it is ready for distribution, a stream URL that users can access is generated. The input is the avatar video, and the output is the stream URL.

[1223] Step 6:

[1224] The user terminal provides a user interface that includes a video player, a comment input field, a question button, etc. The user uses the interface to watch videos and input comments and questions in real time. The input is the user's comments and questions, and the output is the corresponding interface operation.

[1225] Step 7:

[1226] The server receives comments and questions sent from the terminal in real time and adds the comments to a queue using WebSocket technology. The input is the comments and questions received from the user, and the output is the data added to the comment queue.

[1227] Step 8:

[1228] The server sequentially retrieves comments added to the queue and sends them to the avatar engine, where the avatar generates an appropriate response. The input is the comment retrieved from the comment queue, and the output is the generated avatar response.

[1229] Step 9:

[1230] The server analyzes the user's comments and questions using an emotion engine and adjusts the avatar's response based on the emotion. The input is the user's comments and questions, and the output is the adjusted avatar response. The emotion engine uses IBM Watson Natural Language Understanding and other technologies.

[1231] Step 10:

[1232] The user device visually displays the emotion recognition results, allowing the user to confirm the response that reflects their own emotional state. The input is the analysis result of the emotion engine, and the output is the visually displayed emotional information.

[1233] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1234] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1235] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1236] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1237] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1238] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1239] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1240] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1241] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1242] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1243] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1244] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1245] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1246] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1247] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1248] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1249] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1250] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1251] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1252] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1253] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1254] The following is further disclosed regarding the above embodiment.

[1255] (Claim 1)

[1256] a means for obtaining data from external sources;

[1257] a means for storing the acquired data;

[1258] means for filtering information from the stored data based on specific parameters;

[1259] means for generating the selected information as audio and video;

[1260] means for delivering the generated audio and video;

[1261] A way to process comments and questions received in real time during the broadcast and generate responses.

[1262] A system including:

[1263] (Claim 2)

[1264] 10. The system of claim 1, further comprising means for generating an avatar that visually represents the selected information.

[1265] (Claim 3)

[1266] 10. The system of claim 1, further comprising means for providing an interface for accepting input of comments and questions in real time.

[1267] "Example 1"

[1268] (Claim 1)

[1269] a means for obtaining data from external sources;

[1270] a means for storing the acquired data;

[1271] means for filtering information from the stored data based on specific parameters;

[1272] means for generating the selected information as audio and video;

[1273] means for delivering the generated audio and video;

[1274] A means of processing comments and questions received in real time during the broadcast and generating responses; and

[1275] A means for selecting news data from data stored in a database and generating a video in which an avatar reads the news aloud using speech synthesis and animation technologies;

[1276] A means to encode and distribute avatar videos uploaded to streaming services,

[1277] means for processing user comments and questions in response to the encoded video in real time and generating and redistributing avatar responses;

[1278] A system including:

[1279] (Claim 2)

[1280] 10. The system of claim 1, further comprising means for generating an avatar to visually represent and read aloud using speech synthesis and animation techniques.

[1281] (Claim 3)

[1282] 10. The system of claim 1, further comprising means for providing a user interface for accepting input of real-time comments and questions and processing the same to generate avatar responses.

[1283] "Application Example 1"

[1284] (Claim 1)

[1285] a means for obtaining data from external sources;

[1286] a means for storing the acquired data;

[1287] means for filtering information from the stored data based on specific parameters;

[1288] means for generating the selected information as audio and video;

[1289] means for delivering the generated audio and video;

[1290] A means of processing comments and questions received in real time during the broadcast and generating responses; and

[1291] A means for providing such data on a smartphone application;

[1292] A system including:

[1293] (Claim 2)

[1294] 10. The system of claim 1, further comprising means for generating an avatar that visually represents the selected information.

[1295] (Claim 3)

[1296] 10. The system of claim 1, further comprising means for providing an interface for accepting input of comments and questions in real time.

[1297] "Example 2: Combining Emotion Engines"

[1298] (Claim 1)

[1299] a means for obtaining data from external sources;

[1300] a means for storing the acquired data;

[1301] means for filtering information from the stored data based on specific parameters;

[1302] means for generating the selected information as audio and video;

[1303] means for delivering the generated audio and video;

[1304] A means of processing comments and questions received in real time during the broadcast and generating responses; and

[1305] a means for sentiment analysis of received comments and questions and tailoring responses based on the results of that analysis;

[1306] A system including:

[1307] (Claim 2)

[1308] 10. The system of claim 1, further comprising means for generating an avatar that visually represents the selected information.

[1309] (Claim 3)

[1310] 10. The system of claim 1, further comprising means for providing an interface for accepting input of comments and questions in real time.

[1311] "Application example 2 when combining emotion engines"

[1312] New Claims

[1313] (Claim 1)

[1314] a means for obtaining data from external sources;

[1315] a means for storing the acquired data;

[1316] means for filtering information from the stored data based on specific parameters;

[1317] means for generating the selected information as audio and video;

[1318] means for delivering the generated audio and video;

[1319] A means of processing comments and questions received in real time during the broadcast and generating responses; and

[1320] A means of recognizing user emotions from comments and questions and adjusting responses;

[1321] a means for visually displaying the emotion recognition results;

[1322] A system including:

[1323] (Claim 2)

[1324] 10. The system of claim 1, further comprising means for generating an avatar that visually represents the selected information.

[1325] (Claim 3)

[1326] 10. The system of claim 1, further comprising means for providing an interface for accepting input of comments and questions in real time. [Explanation of symbols]

[1327] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. a means for obtaining data from external sources; a means for storing the acquired data; means for filtering information from the stored data based on specific parameters; means for generating the selected information as audio and video; means for delivering the generated audio and video; A way to process comments and questions received in real time during the broadcast and generate responses. A system including:

2. 10. The system of claim 1, further comprising means for generating an avatar that visually represents the selected information.

3. 10. The system of claim 1, further comprising means for providing an interface for accepting input of comments and questions in real time.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A