System

The system transforms text news into real-time video and audio content using generative AI, allowing interactive engagement with VTuber characters and chat responses, addressing the need for visual and auditory news consumption.

JP2026017910APending Publication Date: 2026-02-05SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024118971
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-24
Publication Date
2026-02-05

AI Technical Summary

Technical Problem

Modern lifestyles lack effective methods for engaging with news visually and aurally, and there is a need for interactive two-way communication with news content.

Method used

A system that converts text news into real-time video and audio content using generative AI, allowing users to interact with VTuber characters and receive responses through a chat interface.

Benefits of technology

Enables users to enjoy news visually and aurally, interactively, and engage with others in real-time through chat comments, enhancing the news consumption experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026017910000001_ABST
    Figure 2026017910000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system comprising: means for obtaining text news in real-time; means for receiving a news article selected by a user and passing it to a generative artificial intelligence; means for the generative artificial intelligence generating video and audio content based on the news article; means for streaming the generated video and audio content to the user; and means for receiving chat comments from the user and generating an appropriate response.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] Today's busy lifestyles mean that many people don't have time to sit down and read the news. Furthermore, simply reading news articles as text lacks visual and auditory stimulation, making it difficult to engage. However, given the recent popularity of live streaming services such as VTubers, there is a demand for ways to enjoy the news visually and aurally. Furthermore, there is a lack of two-way communication methods for interacting with the news and exchanging opinions with other people. Therefore, a more convenient and engaging way to consume news is needed. [Means for solving the problem]

[0005] The present invention provides a means for acquiring text news in real time, receiving news articles selected by users, and passing them to a generation artificial intelligence. The system includes a system in which the generation artificial intelligence generates video and audio content based on the news articles and streams the generated content to users. It also provides a means for receiving chat comments from users and generating appropriate responses. This system allows the facial expressions and movements of a VTuber character to be changed in real time according to the news content, allowing users to enjoy the news articles both visually and aurally. It also includes an interface means for users to display and select from a list of news articles, allowing users to easily select and consume news of interest.

[0006] "Text news" refers to news information written in text format provided over the Internet or through other media.

[0007] "Real-time" refers to processing and transmission occurring immediately without delay.

[0008] "User" means any person or entity that utilizes the System to select news articles and receive video and audio content.

[0009] A "news article" is information written in text form about an event or topic.

[0010] "Generative AI" refers to an algorithm or system that automatically generates content in a specified format (e.g., video, audio) based on given data.

[0011] "Video and audio content" means digital media presented in a manner that appeals to both the visual and auditory senses.

[0012] "Streaming distribution" refers to a method of transmitting digital content in real time to a user's terminal via the Internet, making it possible to view the content instantly.

[0013] "Chat comments" means a means for users to send text messages in real time to share their opinions and thoughts.

[0014] An "appropriate response" is a reply or reaction that is generated in response to a chat comment from a user and that accurately corresponds to the meaning and context.

[0015] "Characters" are anthropomorphic digital avatars that visually read the news.

[0016] "Expressions and movements" refer to the emotions and actions that characters show depending on the news content.

[0017] "Interface means" refers to the operating screen or input device that allows the user to select news articles and interact with the system. [Brief explanation of the drawings]

[0018] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0019] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0020] First, the terms used in the following description will be explained.

[0021] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0022] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0023] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0024] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0025] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0026] [First embodiment]

[0027] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0028] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0029] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0030] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0031] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0032] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0033] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0034] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0035] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0036] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0037] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0038] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0039] The present invention is a system that acquires news articles in real time, converts them into video and audio content, and distributes them to users. This allows users to enjoy the news visually and aurally, and also allows them to interact with the news. The following describes an embodiment of the present invention.

[0040] System Configuration

[0041] 1. Server

[0042] The server has the ability to retrieve news articles from news sources in real time, obtain the latest article list using the news source's API, and store the data.

[0043] It also accepts news article selection requests from users and passes the selected article content to the generation AI.

[0044] The artificial intelligence has the ability to generate video and audio content for VTuber characters based on received news articles. The generated content is temporarily stored on the server.

[0045] The server has the function of distributing the generated video and audio content to users in streaming format.

[0046] In addition, it has the ability to receive chat comments sent by users, use generative AI to generate appropriate responses, and return them to the users.

[0047] 2. Terminal

[0048] The terminal provides a user interface and has the function of displaying a list of news articles. When the user selects an article of interest, the selection information is sent to the server.

[0049] It has the ability to play distributed video and audio content in real time.

[0050] It also provides an interface for users to send comments using a chat function.

[0051] 3. Users

[0052] The user selects news articles of interest through the terminal interface.

[0053] Selected news articles are viewed as real-time video and audio content.

[0054] You can send chat comments during the broadcast and receive responses from the server.

[0055] Program processing flow

[0056] Get news articles

[0057] The server calls the news provider's API to retrieve the latest news articles, parses them in JSON format, and saves the necessary information.

[0058] Accepting article conversion requests

[0059] The user selects news articles of interest from the terminal interface and the selection is transmitted to the server.

[0060] Content generation by generative AI

[0061] The server passes the received news article to the generation AI, which generates video and audio content for the VTuber character based on the article content. This includes creating a summary of the article, simulating the character's facial expressions and movements, and synthesizing voice.

[0062] Delivery of video and audio content

[0063] The server transmits the generated content to the user terminal via a streaming server, and the user watches the received content in real time.

[0064] Chat feature

[0065] Users can send chat comments in real time during the broadcast, and the server passes the received comments to a generation AI that generates an appropriate response, which is returned to the user via the chat interface.

[0066] Specific examples

[0067] 1. The user selects a news article.

[0068] The terminal displays a list of news articles and the user selects an article in the "Technology" category.

[0069] 2. The server passes the article to the generation AI.

[0070] The server passes the selected article to a generation AI to generate video and audio content.

[0071] 3. Distribution begins

[0072] The server distributes the generated content to the user through streaming, and the user views the content.

[0073] 4. Use of chat function

[0074] Users submit comments during a broadcast, and the server generates a response and sends it back to the user.

[0075] As described above, the present invention provides a new means for enjoying news visually and aurally, and allows users to react to the news interactively.

[0076] The processing flow will be explained below.

[0077] Step 1:

[0078] The server accesses the news provider's API to retrieve the latest news articles, sending an API request and receiving the response data in JSON format.

[0079] Step 2:

[0080] The server parses the JSON data of the retrieved news article, extracts the necessary information (title, content, URL, etc.), and stores it in an internal database or cache.

[0081] Step 3:

[0082] The device displays a list of news articles to the user on an interface, including the title of each article, a thumbnail image, and a brief summary.

[0083] Step 4:

[0084] The user selects an article of interest from the displayed news article list, and the ID (or URL) of the selected article is sent to the server via the terminal.

[0085] Step 5:

[0086] The server retrieves the corresponding article content from an internal database based on the article ID received from the user.

[0087] Step 6:

[0088] The server sends the retrieved article content to the generation AI module, which analyzes the article content, creates a summary, and then generates video and audio content for the VTuber character.

[0089] Step 7:

[0090] The generative AI simulates the facial expressions and movements of a VTuber character in real time based on a news article, and also generates audio reading the article content and integrates it into the video.

[0091] Step 8:

[0092] The server prepares the video and audio content received from the generation AI so that it can be provided in streaming format, and configures it for distribution through a streaming server.

[0093] Step 9:

[0094] The server then streams the prepared video and audio content in real time to the device, where the user can begin watching.

[0095] Step 10:

[0096] Users can submit comments through a chat interface while watching, and the comments are sent to the server in real time.

[0097] Step 11:

[0098] The server passes chat comments received from users to the generation AI and asks it to generate an appropriate response. The generation AI creates a response based on the content of the comment.

[0099] Step 12:

[0100] The response generated by the generation AI is sent to the server, which returns the response to the user through a chat interface.

[0101] Through these steps, news articles are transformed into video and audio content in real time and delivered interactively, allowing users to enjoy the news visually and audibly and react to it in real time.

[0102] Example 1

[0103] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0104] Users need to maximize their visual and auditory capabilities when accessing text data, but traditional methods do not adequately achieve this. Furthermore, users have limited options for interacting with the data in real time and engaging with it. Therefore, news and other text data must be delivered to users in a more interactive and engaging format.

[0105] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0106] In this invention, the server includes means for acquiring text data in real time, means for receiving data selected by the user and passing it to the generation AI, means for the generation AI to generate video and audio data based on the data, means for streaming the generated video and audio data to the user, and means for receiving comments from the user and generating appropriate responses, allowing the user to enjoy the text data of interest visually and audibly and to react interactively in real time.

[0107] "Text data" is digital data containing text information.

[0108] "Generative AI" is a system that uses artificial intelligence technology to generate moving images and audio data based on text data.

[0109] "Motion picture data" is data in a digital format that shows a sequence of images and provides visual information.

[0110] "Audio data" means data in a digital format that conveys information through hearing.

[0111] "Streaming distribution" is a technology that distributes digital media over the Internet in real time.

[0112] A "comment" is an opinion or question in text format that a user sends in real time in response to video or audio content.

[0113] MODE FOR CARRYING OUT THE INVENTION

[0114] This invention is a system that acquires text data in real time, generates video and audio data based on that data, and distributes them to users. This system allows users to enjoy the text data visually and audibly, and also allows them to respond interactively in real time.

[0115] Hardware and software used

[0116] 1. Server

[0117] It is used to call the news provider's API to obtain the latest news articles. This API can be from a general API service provider, for example.

[0118] Examples of generative AI include OpenAI's GPT-4 and similar AI models.

[0119] DeepFake technology is used to generate moving images, and Google Text-to-Speech and other voice synthesis engines are used.

[0120] Streaming services such as YouTube Live and Twitch are used for streaming.

[0121] 2. Terminal

[0122] It displays a list of news articles and provides a user interface for users to select articles they are interested in. This interface is implemented as a web application or smartphone app.

[0123] Processing flow

[0124] The server retrieves the latest news articles in real time using the news provider's API. The retrieved data is sent to the server in JSON format and analyzed. The analyzed article information is saved in the server's database. Specifically, for example, the server calls the endpoint of a service called "NewsAPI," sends a "GET" request, parses the retrieved JSON data, extracts the article title, text, date, etc., and saves them in the database.

[0125] The user browses through a list of news articles through the interface on the device and selects an article of interest. The ID information of the selected article is sent from the device to the server. For example, the user opens a news app on a smartphone or PC and selects an article in the "Technology" category. The device then sends the selection information to the server.

[0126] The server passes the selected article content to the generation AI. The generation AI generates video and audio content for the VTuber character based on the article content. This includes creating a summary of the article, simulating the character's facial expressions and movements, and processing voice synthesis. Specifically, it sends a prompt to the generation AI model, instructing it to summarize the article content. An example of a prompt is, "Please summarize the following news article: '(news article body)'."

[0127] The server sends the generated video and audio content to the user terminal via the streaming server, and the user watches the content in real time. Specifically, the server uploads the generated video file to the streaming server and sends a streaming URL to the user terminal. The user watches the content in real time via the URL.

[0128] Users can send chat comments in real time while watching. The server passes the received comments to the generation AI, which generates an appropriate response. The generated response is returned to the user via the chat interface. Specifically, the user sends a comment during the broadcast, and the server passes the comment to the generation AI, which generates a response. An example of a prompt is, "Please generate an appropriate response to the user's comment: 'Please tell us a specific application example of this technology.'"

[0129] In this way, the system provides a new means of visually and aurally enjoying text data, allowing users to respond interactively.

[0130] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0131] Step 1: Get news articles

[0132] The server calls the news provider's API to retrieve the latest news articles in real time. The retrieved data is sent to the server in JSON format and parsed. The retrieved JSON data includes information such as the article title, body text, and date. This data is then stored in a database.

[0133] Input: News source API endpoint

[0134] Output: Parsed news article data (title, body, date, etc.)

[0135] Specifically, for example, it uses the endpoint of a typical API service provider to send a "GET" request and parses the retrieved JSON data.

[0136] Step 2: Select an article

[0137] The user selects an article of interest from the list of news articles displayed through the interface on the terminal, and the ID information of the selected article is sent from the terminal to the server.

[0138] Input: The ID of the news article selected by the user

[0139] Output: ID information of selected articles sent to the server

[0140] Specifically, a user opens a news app on their smartphone or PC and selects an article in the "Technology" category, for example. At that time, the device sends a request including the ID of the selected article to the server.

[0141] Step 3: Submitting a content generation request

[0142] The server passes the selected article content to the generation AI, which then generates video and audio content for the VTuber character based on the article content, including creating a summary of the article, simulating the character's facial expressions and movements, and processing voice synthesis.

[0143] Input: Content of selected news article

[0144] Output: Prompt text passed to the generation AI

[0145] Specifically, the system sends a prompt to a generative AI model (e.g., GPT-4) to instruct it to summarize the article. An example prompt might be, "Please summarize the following news article: '(news article text)'."

[0146] Step 4: Generate video and audio content

[0147] The AI ​​generates video and audio content for the VTuber character based on the article content provided, including simulating the character's facial expressions and movements based on the summarized article content and generating audio using a speech synthesis engine.

[0148] Input: Summary of news article content

[0149] Output: The generated video and audio content

[0150] Specifically, it works by using a character simulation tool that uses DeepFake technology to generate voice using a voice synthesis engine (e.g., Google Text-to-Speech).

[0151] Step 5: Deliver your content

[0152] The server transmits the generated video and audio content to the user terminal via a streaming server, where the user can view the content in real time.

[0153] Input: Generated video and audio content

[0154] Output: Streaming URL sent in real time

[0155] Specifically, the generated video file is uploaded to a streaming server (such as YouTube Live or Twitch), and a streaming URL is sent to the user's device. The user can then watch the video in real time via that URL.

[0156] Step 6: Implementing the chat function

[0157] Users can send chat comments in real time while watching, and the server passes the received comments to a generation AI that generates an appropriate response, which is returned to the user via the chat interface.

[0158] Input: Chat comment sent by the user

[0159] Output: The generated response message

[0160] Specifically, while watching, the user sends a comment such as "Please tell me a specific application example of this technology," and the server passes the comment to the generation AI, which then generates an appropriate response. Example prompt: "Generate an appropriate response to the user's comment: 'Please tell me a specific application example of this technology.'"

[0161] (Application example 1)

[0162] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0163] Current news distribution methods are limited to providing text-based information, leaving many users lacking the means to enjoy news visually and aurally. Furthermore, they lack the ability to enjoy news articles interactively, making it difficult for users to efficiently understand news information. Furthermore, they lack the means to easily select specific news articles that interest them and share their reactions to them. These issues need to be addressed.

[0164] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0165] In this invention, the server includes means for acquiring text news in real time, means for receiving news articles selected by users and passing them to a generation AI, means for the generation AI to generate video and audio content based on the news articles, means for streaming the generated video and audio content to users, means for receiving chat comments from users and generating appropriate responses, and means for being installed in a smartphone application, which allows users to enjoy the news visually and audibly and to react to the news interactively.

[0166] "Means for obtaining text news in real time" refers to a system that obtains the latest news articles from news providers in real time via a dedicated API, and analyzes and stores them.

[0167] "Means for receiving news articles selected by the user and passing them on to the generating AI" refers to the communication and data processing functions for transmitting the news article information selected by the user to the generating AI.

[0168] "Generative AI" refers to AI technology for generating video and audio content based on received news articles. Specifically, it includes summarization, character facial expression and movement simulation, and voice synthesis.

[0169] "Means for generating video and audio content" refers to the mechanism by which the generative AI creates visual and audio multimedia content based on the content of news articles.

[0170] The "means for streaming the generated video and audio content to the user" refers to a streaming technology for delivering the generated multimedia content to the user terminal in real time.

[0171] "Means for receiving chat comments from users and generating appropriate responses" refers to a mechanism in which comments sent by users through the chat function are analyzed and the generation AI generates and replies to appropriate responses.

[0172] "Means installed in a smartphone application" refers to smartphone-specific software that incorporates the various functions mentioned above and provides users with an interactive news experience.

[0173] The "interface means for displaying a list of news articles and allowing the user to select an article of interest" is a user interface function that displays a list of news articles and allows the user to select a particular article.

[0174] This invention is a system that acquires news articles in real time, converts them into video and audio content, and distributes them to users. This system is composed of multiple components, such as a server, terminals, and generation AI, and its detailed configuration and processing are described below.

[0175] System Configuration

[0176] 1. Server

[0177] The server can obtain news articles from news providers in real time. This includes functions to obtain the latest article list using the news provider's API and store that data. It also accepts news article selection requests from users and passes the selected article content to the generation AI. The generation AI has the function to generate video and audio content based on the received news articles. The generated content is temporarily stored on the server and delivered to users in streaming format. It also receives chat comments sent by users, uses the generation AI to generate appropriate responses, and returns them to the user.

[0178] 2. Terminal

[0179] The device provides a user interface and can display a list of news articles. When a user selects an article of interest, the device transmits the selection information to the server. It can also play the distributed video and audio content in real time. It also provides an interface for users to send comments using a chat function. The device is installed as a smartphone application.

[0180] 3. Users

[0181] Users select news articles of interest through the device interface, and can view the selected news articles as video and audio content streamed in real time. Users can also send chat comments during the stream and receive responses from the server.

[0182] Program processing flow explanation

[0183] Get news articles

[0184] The server calls the news provider's API to retrieve the latest news articles. The retrieved articles are parsed in JSON format and the necessary information is saved. The server uses Node.js and Express, and an HTTP client such as Axios is used to communicate with the news provider.

[0185] Accepting article conversion requests

[0186] The device displays a list of news articles, and the user selects the news article they are interested in. The selection information is sent to the server. This process is implemented in a smartphone application using React Native.

[0187] Content generation by generative AI

[0188] The server passes the received news article to the generation AI, which generates video and audio content for the VTuber character based on the article content. This includes creating a summary of the article, simulating the character's facial expressions and movements, and synthesizing voice. The generation AI uses GPT-4 and deep learning frameworks such as TensorFlow and PyTorch.

[0189] Delivery of video and audio content

[0190] The server sends the generated content to the user's device via a streaming server (e.g., Wowza Streaming Engine), and the user watches the received content in real time.

[0191] Chat feature

[0192] Users can send chat comments in real time during the broadcast, and the server passes the received comments to the generation AI, which generates an appropriate response, which is returned to the user via the chat interface.

[0193] Examples of concrete examples and prompts

[0194] Specific examples

[0195] The user selects an article in the "Technology" category from a list of news articles. The server retrieves an article about the latest technology for self-driving cars from the news source and passes it to the generation AI. The generation AI summarizes the article and generates a video in which a VTuber character explains it. The user watches the generated video on a smartphone app and asks in chat, "Is the technology introduced here also used by other manufacturers?" The generation AI replies, "Yes, this technology is being adopted by many automakers."

[0196] Example prompts for generative AI models

[0197] Article content: Article about the latest technology in autonomous vehicles

[0198] Title: Evolution of in-vehicle AI technology in 2023

[0199] Article text: Advances in AI technology are enabling the latest self-driving cars to operate more safely and efficiently. In particular, real-time data processing combined with advanced sensor technology has improved the car's ability to accurately perceive its surroundings...

[0200] Prompt: Based on the article below, create a 5-minute news commentary video with a VTuber character.

[0201] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0202] Step 1:

[0203] The server retrieves the latest news articles in real time through the news provider's API.

[0204] How it works: The server periodically sends GET requests to the API endpoint to receive news article data in JSON format, which is then parsed to extract the necessary information (title, body, date, etc.) and store it in a database such as MongoDB.

[0205] Input: JSON data from the news source API.

[0206] Output: News article data stored in a database.

[0207] Step 2:

[0208] The terminal displays a list of news articles to the user.

[0209] Specific operation: The smartphone application on the device displays a list of news articles retrieved from the server on the user interface. The user selects the news article of interest. The selection information is sent from the device to the server.

[0210] Input: A list of news articles retrieved from the server.

[0211] Output: Information about the news article selected by the user.

[0212] Step 3:

[0213] The server passes the news articles selected by the user to the generating artificial intelligence.

[0214] Specific operation: The server calls an internal API to pass the selected news article to the generation AI module, which generates an appropriate prompt. The generation AI processes and analyzes the data based on the prompt and the news article content.

[0215] Input: Information about the news article selected by the user, and a prompt.

[0216] Output: The prompt and news article passed to the generation AI.

[0217] Step 4:

[0218] Generative AI generates video and audio content based on news articles.

[0219] How it works: The generative AI generates content based on prompts, summarizes articles, simulates the facial expressions and movements of VTuber characters, and performs voice synthesis. This results in the creation of visual and audio news commentary videos. The technologies used include GPT-4, TensorFlow, and PyTorch.

[0220] Input: Prompt text and news article content.

[0221] Output: The generated video and audio content.

[0222] Step 5:

[0223] The server streams the generated video and audio content to the user.

[0224] Specific operation: The generated video and audio content is temporarily stored on the server and then distributed in real time to the user's device via a streaming server (e.g., Wowza Streaming Engine).

[0225] Input: Generated video and audio content.

[0226] Output: The streaming content sent to the user device.

[0227] Step 6:

[0228] Users can view the distributed video and audio content and send chat comments.

[0229] Specific operation: A user watches a live video on a smartphone application and sends comments using the chat interface. The comments are sent to the server.

[0230] Input: The user's chat comment.

[0231] Output: Chat comments sent to the server.

[0232] Step 7:

[0233] The server passes the received chat comments to a generation artificial intelligence to generate an appropriate response.

[0234] Specific operation: The server analyzes the chat comments and passes them to the generation AI to generate an appropriate response, which is then sent back to the device and displayed to the user.

[0235] Input: The user's chat comment.

[0236] Output: The chat response generated by the generative AI.

[0237] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0238] The present invention is a system that acquires news articles in real time, converts them into video and audio content, and delivers them to users, and also combines it with an emotion engine that recognizes the user's emotions. This allows users to enjoy the news visually and aurally, and further enables them to receive responses that correspond to their emotions through an interactive experience. The following describes an embodiment of the present invention.

[0239] System Configuration

[0240] 1. Server

[0241] The server has the ability to retrieve news articles from news sources in real time, use the news source's API to obtain the latest article list, and analyze and store the data.

[0242] It also accepts news article selection requests from users and passes the selected article content to the generation artificial intelligence (generation AI).

[0243] The AI ​​has the ability to generate video and audio content for VTuber characters based on the received news articles. The generated content is temporarily stored on the server.

[0244] The server has the function of distributing the generated video and audio content to users in streaming format.

[0245] Furthermore, it has the ability to receive chat comments sent by users, recognize the user's emotions through an emotion engine, and then use the generative AI to generate an appropriate response based on that.

[0246] 2. Terminal

[0247] The terminal provides a user interface and has the function of displaying a list of news articles. When the user selects an article of interest, the selection information is sent to the server.

[0248] It has the ability to play distributed video and audio content in real time.

[0249] It also provides an interface for users to send comments using a chat function.

[0250] 3. Users

[0251] The user selects news articles of interest through the terminal interface.

[0252] Selected news articles are viewed as real-time video and audio content.

[0253] You can send chat comments during the broadcast and receive responses from the server.

[0254] 4. Emotion Engine

[0255] The emotion engine has the ability to recognize emotions from the user's facial expressions and voice, allowing it to grasp the user's emotions in real time.

[0256] The emotion engine has the ability to have the generative AI adjust the content of video and audio content based on the user's emotions.

[0257] Program processing flow

[0258] Get news articles

[0259] The server calls the news provider's API to retrieve the latest news articles. It sends an API request, receives the response data in JSON format, analyzes it, extracts the necessary information (title, content, URL, etc.), and stores it in an internal database.

[0260] Accepting article conversion requests

[0261] The user selects the news article of interest from the terminal interface, and the selection information is sent to the server via the terminal. The server receives the user's request and retrieves the corresponding article content from its internal database.

[0262] Content generation by generative AI

[0263] The server sends the retrieved article content to the generation AI, which analyzes the article content, creates a summary, and then generates video and audio content for the VTuber character. This process includes real-time simulation of the character's facial expressions and movements, as well as voice synthesis.

[0264] Emotion Engine Operation

[0265] The emotion engine analyzes the user's facial expressions and voice data in real time to recognize their emotions. The recognized emotion data is fed back to the generation AI, which then adjusts the content of the video and audio content based on this.

[0266] Delivery of video and audio content

[0267] The server prepares the video and audio content received from the generation AI so that it can be provided in streaming format. It configures the content for distribution through the streaming server and sends it to the device in real time. The user then begins watching the content on their device.

[0268] Chat feature

[0269] Users can send comments through a chat interface while watching. The comments are sent to the server in real time. The server then passes the chat comments received from the user to an emotion engine, which analyzes the user's emotions. Based on the results, the generative AI generates an appropriate response and returns it to the user via the server.

[0270] Specific examples

[0271] 1. The user selects a news article.

[0272] The terminal displays a list of news articles and the user selects an article in the "Technology" category.

[0273] 2. The server passes the article to the generation AI.

[0274] The server passes the selected article to a generation AI to generate video and audio content.

[0275] 3. Use of Emotion Engine

[0276] While a user is watching a video, the emotion engine recognizes emotions from the user's facial expressions and voice, and the generative AI adjusts the content in real time based on that.

[0277] 4. Distribution begins

[0278] The server distributes the generated content to the user through streaming, and the user views the content.

[0279] 5. Use of chat function

[0280] Users send comments during the broadcast, the server analyzes the user's emotions using an emotion engine, and the generative AI generates an appropriate response and returns it to the user.

[0281] The present invention provides a new means of enjoying the news visually and aurally, allowing users to not only react to the news interactively, but also providing appropriate responses and content according to the user's emotions.

[0282] The processing flow will be explained below.

[0283] The present invention is a system that combines a system that acquires news articles in real time and converts them into video and audio content that provides visual and auditory enjoyment to users with an emotion engine that recognizes user emotions. The processing flow of the present invention will be explained below by dividing it into specific steps.

[0284] Step 1:

[0285] The server accesses the news provider's API to retrieve the latest news articles, sends an API request, and receives the response data in JSON format.

[0286] Step 2:

[0287] The server analyzes the JSON data of the retrieved news article, extracts the necessary information (title, content, URL, etc.), and stores it in an internal database.

[0288] Step 3:

[0289] The device displays a list of news articles to the user on an interface, including the title of each article, a thumbnail image, and a brief summary.

[0290] Step 4:

[0291] The user selects an article of interest from the displayed news article list, and the ID (or URL) of the selected article is sent to the server via the terminal.

[0292] Step 5:

[0293] The server retrieves the corresponding article content from an internal database based on the article ID received from the user.

[0294] Step 6:

[0295] The server sends the retrieved article content to the generation AI module, which analyzes the article content, creates a summary, and generates video and audio content for the VTuber character.

[0296] Step 7:

[0297] The generative AI simulates the facial expressions and movements of a VTuber character in real time based on a news article, and also generates audio reading the article content and integrates it into the video.

[0298] Step 8:

[0299] The server prepares the video and audio content received from the generation AI so that it can be provided in streaming format. It sets up distribution through the streaming server.

[0300] Step 9:

[0301] The server streams the prepared video and audio content in real time to the terminal, and the user begins watching.

[0302] Step 10:

[0303] The emotion engine analyzes the user's facial expressions and voice in real time to recognize their emotions, and the recognized emotion data is fed back to the generation AI.

[0304] Step 11:

[0305] The generative AI adjusts the VTuber character's facial expressions and movements based on data from the emotion engine, dynamically changing video and audio content.

[0306] Step 12:

[0307] Users can submit comments through a chat interface while watching, and the comments are sent to the server in real time.

[0308] Step 13:

[0309] The server passes the chat comments received from the user to the emotion engine, which analyzes the user's emotions. Based on the results, the generative AI generates an appropriate response.

[0310] Step 14:

[0311] The responses generated by the AI ​​are returned to the user via the server, and are displayed in a chat interface, allowing the user to enjoy the news interactively.

[0312] Through these steps, news articles are converted into video and audio content in real time, and content tailored to the user's emotions is provided, allowing users to enjoy the news visually and audibly while also interacting with it.

[0313] Example 2

[0314] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0315] In recent years, there has been an increase in the number of ways to enjoy news visually and aurally, but users still rely on general text articles. This makes news viewing one-way and lacks an interactive experience. Furthermore, the inability to analyze emotions based on the news content or respond in real time means that a personalized viewing experience is not provided.

[0316] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0317] In this invention, the server includes means for acquiring text news in real time, means for receiving news articles selected by users and passing them to a generation artificial intelligence, means for the generation artificial intelligence to generate video and audio content based on the news articles, means for streaming the generated video and audio content to users, means for receiving chat comments from users and generating appropriate responses, and means for analyzing user emotions in real time using an emotion engine and adjusting content based on the analysis results. This allows users to enjoy the news visually and audibly and receive responses interactively. Furthermore, personalized content can be provided based on emotion analysis.

[0318] "Text news" refers to news articles that consist only of text information obtained from news sources.

[0319] A "server" is a computer system that provides data and services to multiple clients over a network.

[0320] "Users" are consumers who utilize the system to select news articles and view video and audio content.

[0321] "Generative artificial intelligence" refers to machine learning models and algorithms for generating video and audio content based on retrieved news articles.

[0322] "Video and audio content" means material for visual and auditory enjoyment generated by generative artificial intelligence.

[0323] "Streaming distribution" is a technology that transfers data sequentially and plays it back in real time.

[0324] "Chat comments" are text messages that users type and send in real time.

[0325] An "appropriate response" is an appropriate reply provided by a generative AI or system in response to a chat comment from a user.

[0326] An "emotion engine" is a system or algorithm that analyzes a user's facial expressions and voice data to recognize emotions.

[0327] An "interface" is a screen or operating means through which a user interacts with a system.

[0328] "Real-time" refers to the time nature of data and processing occurring immediately without delay.

[0329] The present invention is a system that acquires news articles in real time, converts them into video and audio content, and delivers them to users, and also combines it with an emotion engine that recognizes the user's emotions. This allows users to enjoy the news visually and aurally, and further enables them to receive responses that correspond to their emotions through an interactive experience. The following describes an embodiment of the present invention.

[0330] System configuration

[0331] 1. Server

[0332] The server has the ability to retrieve news articles from news sources in real time, use the news source's API to obtain the latest article list, and analyze and store the data.

[0333] It also accepts news article selection requests from users and passes the selected article content to the generation artificial intelligence (generation AI).

[0334] The AI ​​generator generates video and audio content for virtual characters based on received news articles. The generated content is temporarily stored on the server.

[0335] The server has the function of distributing the generated video and audio content to users in streaming format.

[0336] Furthermore, it has the ability to receive chat comments sent by users, recognize the user's emotions through an emotion engine, and then use the generative AI to generate an appropriate response based on that.

[0337] 2. Terminal

[0338] The terminal provides a user interface and has the function of displaying a list of news articles. When the user selects an article of interest, the selection information is sent to the server.

[0339] It has the ability to play distributed video and audio content in real time.

[0340] It also provides an interface for users to send comments using a chat function.

[0341] 3. Users

[0342] The user selects news articles of interest through the terminal interface.

[0343] Selected news articles are viewed as real-time video and audio content.

[0344] You can send chat comments during the broadcast and receive responses from the server.

[0345] 4. Emotion Engine

[0346] The emotion engine has the ability to recognize emotions from the user's facial expressions and voice, allowing it to grasp the user's emotions in real time.

[0347] The emotion engine has the ability to have the generative AI adjust the content of video and audio content based on the user's emotions.

[0348] Specific examples

[0349] For example, each component operates as follows:

[0350] 1. Select a news article

[0351] The user browses the list of news articles on the device interface and selects an article in the "Technology" category.

[0352] 2. Content generation using generative AI

[0353] The article selected by the user is sent to the server via the device, and the server passes the article to the AI ​​generator, which then generates video and audio content read by a virtual character.

[0354] 3. Emotion analysis using an emotion engine

[0355] While the user is watching the video, the emotion engine recognizes emotions from the user's facial expressions and voice, and this information is fed back to the generative AI, which then adjusts the content in real time.

[0356] 4. Streaming

[0357] The server distributes the generated content to the user in streaming format, and the user watches it in real time on the terminal.

[0358] 5. Chat Comments and Response Generation

[0359] Users can submit comments while watching, which are sent to the server, which uses an emotion engine to analyze the user's emotions, and the generative AI generates an appropriate response, which is then sent back to the user via the server.

[0360] Prompt Sentence Examples

[0361] "Create a video and audio piece that briefly summarizes the contents of a news article and has a virtual character read it aloud."

[0362] In this way, users can enjoy the news visually and audibly, while simultaneously receiving interactive responses. Personalized content based on sentiment analysis is also provided.

[0363] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0364] Step 1: Get news articles

[0365] The server calls the news provider's API to retrieve the latest news articles. As input, it sends a request with the API endpoint and required parameters. As output, it receives the news data in JSON format, parses it, and extracts the required information (title, content, URL, etc.). It then stores this data in an internal database.

[0366] Specific behavior:

[0367] The server sends a request to the News API using "https: / / newsapi.org / v2 / top-headlines?country=jp&apiKey=Your API Key".

[0368] The returned JSON data is analyzed, and the article title, content, URL, etc. are extracted and saved in the database.

[0369] Step 2: View and select a news article

[0370] The terminal displays a list of news articles. As input, it receives news article data retrieved from the server. As output, it displays the news article list in a user interface. The user selects news articles of interest and sends the selection information to the server.

[0371] Specific behavior:

[0372] The terminal receives the news article list sent from the server and displays it on the interface.

[0373] The user selects an article in the "technology" category, and the terminal transmits this selection information to the server.

[0374] Step 3: Send to article generation AI

[0375] The server receives the user's selection information. As input, it receives the ID and link of the news article selected by the user. As output, it retrieves the corresponding article content from the internal database and sends it to the generation AI.

[0376] Specific behavior:

[0377] The server retrieves relevant news articles from the database based on the selection information sent by the user.

[0378] The retrieved news articles are sent in text format to the generation AI.

[0379] Step 4: Content generation with generative AI

[0380] The generative AI generates video and audio content based on the submitted article. It receives the text data of the news article as input. It generates video and audio content of a virtual character as output. Specifically, it analyzes the article content, generates a summary, simulates the character's facial expressions and movements, and synthesizes voice.

[0381] Specific behavior:

[0382] The generative AI analyzes news articles, creates summaries, and generates videos with the content, "A new technology has been released. This technology is..."

[0383] The character's movements and facial expressions are specified, voice synthesis is performed, and the generated content is sent to the server.

[0384] Step 5: Prepare your content for distribution

[0385] The server receives the video and audio content generated by the generation AI. As input, it receives the video and audio data from the generation AI. As output, it converts these data into a streaming format and prepares it for distribution to the user's device through the streaming server.

[0386] Specific behavior:

[0387] The server receives the video and audio files from the generated AI and transfers them to the streaming server.

[0388] Configure streaming settings and generate a streaming link.

[0389] Step 6: Stream your content

[0390] The server streams the generated content to the user's device. As input, the server sends the generated content data to the streaming server and provides the streaming link to the user. As output, the server provides the content to be played in real time on the user's device.

[0391] Specific behavior:

[0392] The server provides a streaming link of the generated video content to the user's terminal.

[0393] Users can watch video and audio content in real time on their devices.

[0394] Step 7: Sending chat comments and generating replies

[0395] Users can send chat comments while watching a video. They input their comments using the device interface and send them to the server.

[0396] The server receives the sent chat comments and analyzes them using the emotion engine. As input, it receives chat comments from users. As output, it feeds back the analysis results to the generation AI, which generates an appropriate response.

[0397] Specific behavior:

[0398] The user sends a comment from the device saying, "This news is amazing!"

[0399] The server receives the comments and analyzes the emotion of "surprise" using an emotion engine.

[0400] The generative AI generates a response such as "Yes, modern technology is truly amazing," and the server sends that response to the user.

[0401] In this way, users can enjoy the news visually and audibly, while simultaneously receiving interactive responses. Personalized content based on sentiment analysis is also provided.

[0402] (Application example 2)

[0403] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0404] Conventional news delivery systems were capable of acquiring and providing news text information in real time, but were limited in the means to provide it as visually and aurally enjoyable video and audio content. Furthermore, they were not able to recognize users' emotions and generate interactive content or responses accordingly. Furthermore, while many users often feel the need for appropriate responses or feedback based on their emotions while watching the news, no system existed that could meet these needs.

[0405] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for acquiring text news in real time, means for receiving a news article selected by the user and passing it to the generation AI, means for the generation AI to generate video and audio content based on the news article, means for streaming the generated video and audio content to the user, means for receiving chat comments from the user and generating an appropriate response, and means for recognizing the user's emotions and adjusting the content based on the emotion. This enables highly real-time news delivery that can be enjoyed visually and audibly, and makes it possible to provide appropriate responses and interactive experiences according to the user's emotions.

[0406] "Text news" refers to news articles in written form distributed in real time over the Internet or through other communication means.

[0407] "User" refers to an individual or group of people who utilize the system of the present invention to select, view, and interact with news stories.

[0408] "Generative AI" refers to AI that has the ability to generate video and audio content based on input news articles.

[0409] "Video and audio content" means media that provides information through visual and audio means, including generated news feeds.

[0410] "Streaming distribution" is a method of providing video and audio content to users in real time over the Internet.

[0411] "Chat comments" refer to text messages that users enter while viewing content.

[0412] An "appropriate response" is an automatically generated reply or feedback that is generated in response to a chat comment or emotion from a user.

[0413] "Emotion recognition" refers to the process of analyzing and understanding a user's emotional state from data such as facial expressions and voice.

[0414] "Adjusting content" refers to changing the content or format of video and audio content in real time based on recognized user sentiment.

[0415] The present invention combines a system that acquires news articles in real time, converts them into video and audio content, and distributes them to users with an emotion engine that recognizes user emotions. Specific embodiments for carrying out the present invention will be described below.

[0416] System Configuration

[0417] The present invention is mainly composed of a server, a terminal, and a user. Each component will be explained below.

[0418] server

[0419] The server has the following functions:

[0420] 1. Real-time text news acquisition:

[0421] The server uses the news provider's API to retrieve the latest articles in real time. It sends an API request and receives the response data in JSON format. It extracts the necessary information (title, content, URL, etc.) and stores it in an internal database.

[0422] 2. Select news articles and pass them to the generation AI:

[0423] The system receives information about a news article selected by the user on their device and passes it to a generative artificial intelligence (generative AI). The generative AI generates video and audio content based on the article content. This generation process includes creating a summary of the news article, generating video of the VTuber character, and synthesizing voice.

[0424] 3. Emotion recognition and content adjustment:

[0425] The emotion engine recognizes emotions from the user's facial expressions and voice in real time and feeds the recognized emotion data back to the generation AI, which then adjusts the content of the video and audio content based on this feedback information.

[0426] 4. Streaming:

[0427] The generated video and audio content is distributed to the user's terminal in streaming format.

[0428] 5. Handling chat comments:

[0429] The system receives chat comments from users and analyzes their emotions with an emotion engine. Based on the results, the generative AI generates an appropriate response and returns it to the user via the server.

[0430] Terminal

[0431] The terminal has the following functions:

[0432] 1. Select a news article:

[0433] It provides a user interface to display a list of news articles, and when the user selects an article of interest, it sends the selection information to the server.

[0434] 2. Viewing content:

[0435] It has the ability to play distributed video and audio content in real time.

[0436] 3. Chat function:

[0437] It provides an interface that allows users to submit comments while watching, which are then sent to the server, which returns an appropriate response.

[0438] User

[0439] The user performs the following actions:

[0440] 1. Select a news article:

[0441] Select the news article of interest through the device interface.

[0442] 2. Content viewing and emotional feedback:

[0443] The selected news article is then viewed as real-time generated video and audio content, with facial expressions and voice acting analyzed by the emotion engine as the news is viewed.

[0444] 3. Sending chat comments:

[0445] Send comments while watching and receive responses from the generative AI.

[0446] Hardware and software used

[0447] On the server side, we use a server equipped with a high-performance GPU. The main software used is an API for retrieving news from news sources, generative AI (e.g., OpenAI's GPT-4), and an emotion recognition engine (e.g., DeepFace, OpenFace).

[0448] On the terminal side, mobile devices such as smartphones and tablets are used, which are equipped with cameras and microphones to collect data for emotion recognition.

[0449] Specific examples

[0450] 1. View news article list:

[0451] The device displays a list of the latest news, and the user selects "an article about the latest AI technology."

[0452] 2. Content Creation and Viewing:

[0453] The server uses AI to convert the selected articles into video and audio content, which are then streamed. Users can then watch the generated VTuber videos on their devices.

[0454] 3. Emotion Recognition and Feedback:

[0455] While watching, the user's facial expressions are analyzed by a camera, and a "happy" expression is recognized. This information is fed back to the generative AI, which then adjusts the video and audio content to better match the user's emotions.

[0456] 4. Chat comment response:

[0457] When a user types "This technology is amazing!" into the chat, the emotion engine analyzes the emotion and the generative AI responds appropriately with "Yes, the evolution of AI is truly amazing!"

[0458] Prompt Sentence Examples

[0459] "News article: 'The latest AI technology has been announced. This technology is...'"

[0460] "Summary: Please summarize this article."

[0461] "Video and Audio: Generate a script in which a VTuber character explains this news article in an easy-to-understand way."

[0462] The above is a specific embodiment for carrying out the present invention.

[0463] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0464] Step 1:

[0465] The server uses the news provider's API to obtain the latest news articles in real time. It sends an API request and receives the response data in JSON format. The server extracts necessary information such as the news article title, content, and URL, and stores it in an internal database. The input is the API request, and the output is the analyzed news article data.

[0466] Step 2:

[0467] The user selects an article of interest from the displayed news article list using the terminal interface. The selection information is sent from the terminal to the server. The input is the user's selection action, and the output is the selected news article information.

[0468] Step 3:

[0469] The server passes data based on the selected news article to the generation AI. The generation AI then analyzes the article content, creates a summary, and uses that to generate video and audio content for the VTuber character. The input is the selected news article data, and the output is the generated video and audio content.

[0470] Step 4:

[0471] The emotion engine starts up and analyzes the user's facial expressions and voice in real time. Data is acquired using the device's camera and microphone, and the emotion engine analyzes that data to recognize the user's emotions. The input is the user's facial and voice data, and the output is the recognized emotion data.

[0472] Step 5:

[0473] The recognized emotional data is then fed back to the generation AI, which then adjusts the content based on the fed-back emotional data. Specifically, it changes the character's facial expression and tone according to the user's emotions. The input is the recognized emotional data, and the output is the adjusted content.

[0474] Step 6:

[0475] The server distributes the generated video and audio content to the user's device in streaming format. This is done in real time through a streaming server. The input is the adjusted content, and the output is the video and audio to be distributed.

[0476] Step 7:

[0477] While watching a video, users can send comments using a chat interface. The comments are sent from the device to the server. The input is the user's chat comments, and the output is the comment data received by the server.

[0478] Step 8:

[0479] The server passes the received chat comments to the emotion engine for analysis. The emotion engine analyzes emotions based on the user's comments and feeds the results back to the generation AI. The input is the chat comments, and the output is the analyzed emotion data.

[0480] Step 9:

[0481] The generative AI generates an appropriate response based on the analyzed emotional data and returns it to the user through the server. The response is generated using a prompt sentence. The input is the analyzed emotional data, and the output is the response message generated by the AI.

[0482] Example prompt

[0483] "News article: 'The latest AI technology has been announced. This technology is...'"

[0484] "Summary: Please summarize this article."

[0485] "Video and Audio: Generate a script in which a VTuber character explains this news article in an easy-to-understand way."

[0486] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0487] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0488] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0489] [Second embodiment]

[0490] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0491] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0492] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0493] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0494] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0495] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0496] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0497] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0498] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0499] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0500] In the smart glasses 214, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0501] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0502] The present invention is a system that acquires news articles in real time, converts them into video and audio content, and distributes them to users. This allows users to enjoy the news visually and aurally, and also allows them to interact with the news. The following describes an embodiment of the present invention.

[0503] System Configuration

[0504] 1. Server

[0505] The server has the ability to retrieve news articles from news sources in real time, obtain the latest article list using the news source's API, and store the data.

[0506] It also accepts news article selection requests from users and passes the selected article content to the generation AI.

[0507] The artificial intelligence has the ability to generate video and audio content for VTuber characters based on received news articles. The generated content is temporarily stored on the server.

[0508] The server has the function of distributing the generated video and audio content to users in streaming format.

[0509] In addition, it has the ability to receive chat comments sent by users, use generative AI to generate appropriate responses, and return them to the users.

[0510] 2. Terminal

[0511] The terminal provides a user interface and has the function of displaying a list of news articles. When the user selects an article of interest, the selection information is sent to the server.

[0512] It has the ability to play distributed video and audio content in real time.

[0513] It also provides an interface for users to send comments using a chat function.

[0514] 3. Users

[0515] The user selects news articles of interest through the terminal interface.

[0516] Selected news articles are viewed as real-time video and audio content.

[0517] You can send chat comments during the broadcast and receive responses from the server.

[0518] Program processing flow

[0519] Get news articles

[0520] The server calls the news provider's API to retrieve the latest news articles, parses them in JSON format, and saves the necessary information.

[0521] Accepting article conversion requests

[0522] The user selects news articles of interest from the terminal interface and the selection is transmitted to the server.

[0523] Content generation by generative AI

[0524] The server passes the received news article to the generation AI, which generates video and audio content for the VTuber character based on the article content. This includes creating a summary of the article, simulating the character's facial expressions and movements, and synthesizing voice.

[0525] Delivery of video and audio content

[0526] The server transmits the generated content to the user terminal via a streaming server, and the user watches the received content in real time.

[0527] Chat feature

[0528] Users can send chat comments in real time during the broadcast, and the server passes the received comments to a generation AI that generates an appropriate response, which is returned to the user via the chat interface.

[0529] Specific examples

[0530] 1. The user selects a news article.

[0531] The terminal displays a list of news articles and the user selects an article in the "Technology" category.

[0532] 2. The server passes the article to the generation AI.

[0533] The server passes the selected article to a generation AI to generate video and audio content.

[0534] 3. Distribution begins

[0535] The server distributes the generated content to the user through streaming, and the user views the content.

[0536] 4. Use of chat function

[0537] Users submit comments during a broadcast, and the server generates a response and sends it back to the user.

[0538] As described above, the present invention provides a new means for enjoying news visually and aurally, and allows users to react to the news interactively.

[0539] The processing flow will be explained below.

[0540] Step 1:

[0541] The server accesses the news provider's API to retrieve the latest news articles, sending an API request and receiving the response data in JSON format.

[0542] Step 2:

[0543] The server parses the JSON data of the retrieved news article, extracts the necessary information (title, content, URL, etc.), and stores it in an internal database or cache.

[0544] Step 3:

[0545] The device displays a list of news articles to the user on an interface, including the title of each article, a thumbnail image, and a brief summary.

[0546] Step 4:

[0547] The user selects an article of interest from the displayed news article list, and the ID (or URL) of the selected article is sent to the server via the terminal.

[0548] Step 5:

[0549] The server retrieves the corresponding article content from an internal database based on the article ID received from the user.

[0550] Step 6:

[0551] The server sends the retrieved article content to the generation AI module, which analyzes the article content, creates a summary, and then generates video and audio content for the VTuber character.

[0552] Step 7:

[0553] The generative AI simulates the facial expressions and movements of a VTuber character in real time based on a news article, and also generates audio reading the article content and integrates it into the video.

[0554] Step 8:

[0555] The server prepares the video and audio content received from the generation AI so that it can be provided in streaming format, and configures it for distribution through a streaming server.

[0556] Step 9:

[0557] The server then streams the prepared video and audio content in real time to the device, where the user can begin watching.

[0558] Step 10:

[0559] Users can submit comments through a chat interface while watching, and the comments are sent to the server in real time.

[0560] Step 11:

[0561] The server passes chat comments received from users to the generation AI and asks it to generate an appropriate response. The generation AI creates a response based on the content of the comment.

[0562] Step 12:

[0563] The response generated by the generation AI is sent to the server, which returns the response to the user through a chat interface.

[0564] Through these steps, news articles are transformed into video and audio content in real time and delivered interactively, allowing users to enjoy the news visually and audibly and react to it in real time.

[0565] Example 1

[0566] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0567] Users need to maximize their visual and auditory capabilities when accessing text data, but traditional methods do not adequately achieve this. Furthermore, users have limited options for interacting with the data in real time and engaging with it. Therefore, news and other text data must be delivered to users in a more interactive and engaging format.

[0568] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0569] In this invention, the server includes means for acquiring text data in real time, means for receiving data selected by the user and passing it to the generation AI, means for the generation AI to generate video and audio data based on the data, means for streaming the generated video and audio data to the user, and means for receiving comments from the user and generating appropriate responses, allowing the user to enjoy the text data of interest visually and audibly and to react interactively in real time.

[0570] "Text data" is digital data containing text information.

[0571] "Generative AI" is a system that uses artificial intelligence technology to generate moving images and audio data based on text data.

[0572] "Motion picture data" is data in a digital format that shows a sequence of images and provides visual information.

[0573] "Audio data" means data in a digital format that conveys information through hearing.

[0574] "Streaming distribution" is a technology that distributes digital media over the Internet in real time.

[0575] A "comment" is an opinion or question in text format that a user sends in real time in response to video or audio content.

[0576] MODE FOR CARRYING OUT THE INVENTION

[0577] This invention is a system that acquires text data in real time, generates video and audio data based on that data, and distributes them to users. This system allows users to enjoy the text data visually and audibly, and also allows them to respond interactively in real time.

[0578] Hardware and software used

[0579] 1. Server

[0580] It is used to call the news provider's API to obtain the latest news articles. This API can be from a general API service provider, for example.

[0581] Examples of generative AI include OpenAI's GPT-4 and similar AI models.

[0582] DeepFake technology is used to generate moving images, and Google Text-to-Speech and other voice synthesis engines are used.

[0583] Streaming services such as YouTube Live and Twitch are used for streaming.

[0584] 2. Terminal

[0585] It displays a list of news articles and provides a user interface for users to select articles they are interested in. This interface is implemented as a web application or smartphone app.

[0586] Processing flow

[0587] The server retrieves the latest news articles in real time using the news provider's API. The retrieved data is sent to the server in JSON format and analyzed. The analyzed article information is saved in the server's database. Specifically, for example, the server calls the endpoint of a service called "NewsAPI," sends a "GET" request, parses the retrieved JSON data, extracts the article title, text, date, etc., and saves them in the database.

[0588] The user browses through a list of news articles through the interface on the device and selects an article of interest. The ID information of the selected article is sent from the device to the server. For example, the user opens a news app on a smartphone or PC and selects an article in the "Technology" category. The device then sends the selection information to the server.

[0589] The server passes the selected article content to the generation AI. The generation AI generates video and audio content for the VTuber character based on the article content. This includes creating a summary of the article, simulating the character's facial expressions and movements, and processing voice synthesis. Specifically, it sends a prompt to the generation AI model, instructing it to summarize the article content. An example of a prompt is, "Please summarize the following news article: '(news article body)'."

[0590] The server sends the generated video and audio content to the user terminal via the streaming server, and the user watches the content in real time. Specifically, the server uploads the generated video file to the streaming server and sends a streaming URL to the user terminal. The user watches the content in real time via the URL.

[0591] Users can send chat comments in real time while watching. The server passes the received comments to the generation AI, which generates an appropriate response. The generated response is returned to the user via the chat interface. Specifically, the user sends a comment during the broadcast, and the server passes the comment to the generation AI, which generates a response. An example of a prompt is, "Please generate an appropriate response to the user's comment: 'Please tell us a specific application example of this technology.'"

[0592] In this way, the system provides a new means of visually and aurally enjoying text data, allowing users to respond interactively.

[0593] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0594] Step 1: Get news articles

[0595] The server calls the news provider's API to retrieve the latest news articles in real time. The retrieved data is sent to the server in JSON format and parsed. The retrieved JSON data includes information such as the article title, body text, and date. This data is then stored in a database.

[0596] Input: News source API endpoint

[0597] Output: Parsed news article data (title, body, date, etc.)

[0598] Specifically, for example, it uses the endpoint of a typical API service provider to send a "GET" request and parses the retrieved JSON data.

[0599] Step 2: Select an article

[0600] The user selects an article of interest from the list of news articles displayed through the interface on the terminal, and the ID information of the selected article is sent from the terminal to the server.

[0601] Input: The ID of the news article selected by the user

[0602] Output: ID information of selected articles sent to the server

[0603] Specifically, a user opens a news app on their smartphone or PC and selects an article in the "Technology" category, for example. At that time, the device sends a request including the ID of the selected article to the server.

[0604] Step 3: Submitting a content generation request

[0605] The server passes the selected article content to the generation AI, which then generates video and audio content for the VTuber character based on the article content, including creating a summary of the article, simulating the character's facial expressions and movements, and processing voice synthesis.

[0606] Input: Content of selected news article

[0607] Output: Prompt text passed to the generation AI

[0608] Specifically, the system sends a prompt to a generative AI model (e.g., GPT-4) to instruct it to summarize the article. An example prompt might be, "Please summarize the following news article: '(news article text)'."

[0609] Step 4: Generate video and audio content

[0610] The AI ​​generates video and audio content for the VTuber character based on the article content provided, including simulating the character's facial expressions and movements based on the summarized article content and generating audio using a speech synthesis engine.

[0611] Input: Summary of news article content

[0612] Output: The generated video and audio content

[0613] Specifically, it works by using a character simulation tool that uses DeepFake technology to generate voice using a voice synthesis engine (e.g., Google Text-to-Speech).

[0614] Step 5: Deliver your content

[0615] The server transmits the generated video and audio content to the user terminal via a streaming server, where the user can view the content in real time.

[0616] Input: Generated video and audio content

[0617] Output: Streaming URL sent in real time

[0618] Specifically, the generated video file is uploaded to a streaming server (such as YouTube Live or Twitch), and a streaming URL is sent to the user's device. The user can then watch the video in real time via that URL.

[0619] Step 6: Implementing the chat function

[0620] Users can send chat comments in real time while watching, and the server passes the received comments to a generation AI that generates an appropriate response, which is returned to the user via the chat interface.

[0621] Input: Chat comment sent by the user

[0622] Output: The generated response message

[0623] Specifically, while watching, the user sends a comment such as "Please tell me a specific application example of this technology," and the server passes the comment to the generation AI, which then generates an appropriate response. Example prompt: "Generate an appropriate response to the user's comment: 'Please tell me a specific application example of this technology.'"

[0624] (Application example 1)

[0625] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0626] Current news distribution methods are limited to providing text-based information, leaving many users lacking the means to enjoy news visually and aurally. Furthermore, they lack the ability to enjoy news articles interactively, making it difficult for users to efficiently understand news information. Furthermore, they lack the means to easily select specific news articles that interest them and share their reactions to them. These issues need to be addressed.

[0627] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0628] In this invention, the server includes means for acquiring text news in real time, means for receiving news articles selected by users and passing them to a generation AI, means for the generation AI to generate video and audio content based on the news articles, means for streaming the generated video and audio content to users, means for receiving chat comments from users and generating appropriate responses, and means for being installed in a smartphone application, which allows users to enjoy the news visually and audibly and to react to the news interactively.

[0629] "Means for obtaining text news in real time" refers to a system that obtains the latest news articles from news providers in real time via a dedicated API, and analyzes and stores them.

[0630] "Means for receiving news articles selected by the user and passing them on to the generating AI" refers to the communication and data processing functions for transmitting the news article information selected by the user to the generating AI.

[0631] "Generative AI" refers to AI technology for generating video and audio content based on received news articles. Specifically, it includes summarization, character facial expression and movement simulation, and voice synthesis.

[0632] "Means for generating video and audio content" refers to the mechanism by which the generative AI creates visual and audio multimedia content based on the content of news articles.

[0633] The "means for streaming the generated video and audio content to the user" refers to a streaming technology for delivering the generated multimedia content to the user terminal in real time.

[0634] "Means for receiving chat comments from users and generating appropriate responses" refers to a mechanism in which comments sent by users through the chat function are analyzed and the generation AI generates and replies to appropriate responses.

[0635] "Means installed in a smartphone application" refers to smartphone-specific software that incorporates the various functions mentioned above and provides users with an interactive news experience.

[0636] The "interface means for displaying a list of news articles and allowing the user to select an article of interest" is a user interface function that displays a list of news articles and allows the user to select a particular article.

[0637] This invention is a system that acquires news articles in real time, converts them into video and audio content, and distributes them to users. This system is composed of multiple components, such as a server, terminals, and generation AI, and its detailed configuration and processing are described below.

[0638] System Configuration

[0639] 1. Server

[0640] The server can obtain news articles from news providers in real time. This includes functions to obtain the latest article list using the news provider's API and store that data. It also accepts news article selection requests from users and passes the selected article content to the generation AI. The generation AI has the function to generate video and audio content based on the received news articles. The generated content is temporarily stored on the server and delivered to users in streaming format. It also receives chat comments sent by users, uses the generation AI to generate appropriate responses, and returns them to the user.

[0641] 2. Terminal

[0642] The device provides a user interface and can display a list of news articles. When a user selects an article of interest, the device transmits the selection information to the server. It can also play the distributed video and audio content in real time. It also provides an interface for users to send comments using a chat function. The device is installed as a smartphone application.

[0643] 3. Users

[0644] Users select news articles of interest through the device interface, and can view the selected news articles as video and audio content streamed in real time. Users can also send chat comments during the stream and receive responses from the server.

[0645] Program processing flow explanation

[0646] Get news articles

[0647] The server calls the news provider's API to retrieve the latest news articles. The retrieved articles are parsed in JSON format and the necessary information is saved. The server uses Node.js and Express, and an HTTP client such as Axios is used to communicate with the news provider.

[0648] Accepting article conversion requests

[0649] The device displays a list of news articles, and the user selects the news article they are interested in. The selection information is sent to the server. This process is implemented in a smartphone application using React Native.

[0650] Content generation by generative AI

[0651] The server passes the received news article to the generation AI, which generates video and audio content for the VTuber character based on the article content. This includes creating a summary of the article, simulating the character's facial expressions and movements, and synthesizing voice. The generation AI uses GPT-4 and deep learning frameworks such as TensorFlow and PyTorch.

[0652] Delivery of video and audio content

[0653] The server sends the generated content to the user's device via a streaming server (e.g., Wowza Streaming Engine), and the user watches the received content in real time.

[0654] Chat feature

[0655] Users can send chat comments in real time during the broadcast, and the server passes the received comments to the generation AI, which generates an appropriate response, which is returned to the user via the chat interface.

[0656] Examples of concrete examples and prompts

[0657] Specific examples

[0658] The user selects an article in the "Technology" category from a list of news articles. The server retrieves an article about the latest technology for self-driving cars from the news source and passes it to the generation AI. The generation AI summarizes the article and generates a video in which a VTuber character explains it. The user watches the generated video on a smartphone app and asks in chat, "Is the technology introduced here also used by other manufacturers?" The generation AI replies, "Yes, this technology is being adopted by many automakers."

[0659] Example prompts for generative AI models

[0660] Article content: Article about the latest technology in autonomous vehicles

[0661] Title: Evolution of in-vehicle AI technology in 2023

[0662] Article text: Advances in AI technology are enabling the latest self-driving cars to operate more safely and efficiently. In particular, real-time data processing combined with advanced sensor technology has improved the car's ability to accurately perceive its surroundings...

[0663] Prompt: Based on the article below, create a 5-minute news commentary video with a VTuber character.

[0664] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0665] Step 1:

[0666] The server retrieves the latest news articles in real time through the news provider's API.

[0667] How it works: The server periodically sends GET requests to the API endpoint to receive news article data in JSON format, which is then parsed to extract the necessary information (title, body, date, etc.) and store it in a database such as MongoDB.

[0668] Input: JSON data from the news source API.

[0669] Output: News article data stored in a database.

[0670] Step 2:

[0671] The terminal displays a list of news articles to the user.

[0672] Specific operation: The smartphone application on the device displays a list of news articles retrieved from the server on the user interface. The user selects the news article of interest. The selection information is sent from the device to the server.

[0673] Input: A list of news articles retrieved from the server.

[0674] Output: Information about the news article selected by the user.

[0675] Step 3:

[0676] The server passes the news articles selected by the user to the generating artificial intelligence.

[0677] Specific operation: The server calls an internal API to pass the selected news article to the generation AI module, which generates an appropriate prompt. The generation AI processes and analyzes the data based on the prompt and the news article content.

[0678] Input: Information about the news article selected by the user, and a prompt.

[0679] Output: The prompt and news article passed to the generation AI.

[0680] Step 4:

[0681] Generative AI generates video and audio content based on news articles.

[0682] How it works: The generative AI generates content based on prompts, summarizes articles, simulates the facial expressions and movements of VTuber characters, and performs voice synthesis. This results in the creation of visual and audio news commentary videos. The technologies used include GPT-4, TensorFlow, and PyTorch.

[0683] Input: Prompt text and news article content.

[0684] Output: The generated video and audio content.

[0685] Step 5:

[0686] The server streams the generated video and audio content to the user.

[0687] Specific operation: The generated video and audio content is temporarily stored on the server and then distributed in real time to the user's device via a streaming server (e.g., Wowza Streaming Engine).

[0688] Input: Generated video and audio content.

[0689] Output: The streaming content sent to the user device.

[0690] Step 6:

[0691] Users can view the distributed video and audio content and send chat comments.

[0692] Specific operation: A user watches a live video on a smartphone application and sends comments using the chat interface. The comments are sent to the server.

[0693] Input: The user's chat comment.

[0694] Output: Chat comments sent to the server.

[0695] Step 7:

[0696] The server passes the received chat comments to a generation artificial intelligence to generate an appropriate response.

[0697] Specific operation: The server analyzes the chat comments and passes them to the generation AI to generate an appropriate response, which is then sent back to the device and displayed to the user.

[0698] Input: The user's chat comment.

[0699] Output: The chat response generated by the generative AI.

[0700] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0701] The present invention is a system that acquires news articles in real time, converts them into video and audio content, and delivers them to users, and also combines it with an emotion engine that recognizes the user's emotions. This allows users to enjoy the news visually and aurally, and further enables them to receive responses that correspond to their emotions through an interactive experience. The following describes an embodiment of the present invention.

[0702] System Configuration

[0703] 1. Server

[0704] The server has the ability to retrieve news articles from news sources in real time, use the news source's API to obtain the latest article list, and analyze and store the data.

[0705] It also accepts news article selection requests from users and passes the selected article content to the generation artificial intelligence (generation AI).

[0706] The AI ​​has the ability to generate video and audio content for VTuber characters based on the received news articles. The generated content is temporarily stored on the server.

[0707] The server has the function of distributing the generated video and audio content to users in streaming format.

[0708] Furthermore, it has the ability to receive chat comments sent by users, recognize the user's emotions through an emotion engine, and then use the generative AI to generate an appropriate response based on that.

[0709] 2. Terminal

[0710] The terminal provides a user interface and has the function of displaying a list of news articles. When the user selects an article of interest, the selection information is sent to the server.

[0711] It has the ability to play distributed video and audio content in real time.

[0712] It also provides an interface for users to send comments using a chat function.

[0713] 3. Users

[0714] The user selects news articles of interest through the terminal interface.

[0715] Selected news articles are viewed as real-time video and audio content.

[0716] You can send chat comments during the broadcast and receive responses from the server.

[0717] 4. Emotion Engine

[0718] The emotion engine has the ability to recognize emotions from the user's facial expressions and voice, allowing it to grasp the user's emotions in real time.

[0719] The emotion engine has the ability to have the generative AI adjust the content of video and audio content based on the user's emotions.

[0720] Program processing flow

[0721] Get news articles

[0722] The server calls the news provider's API to retrieve the latest news articles. It sends an API request, receives the response data in JSON format, analyzes it, extracts the necessary information (title, content, URL, etc.), and stores it in an internal database.

[0723] Accepting article conversion requests

[0724] The user selects the news article of interest from the terminal interface, and the selection information is sent to the server via the terminal. The server receives the user's request and retrieves the corresponding article content from its internal database.

[0725] Content generation by generative AI

[0726] The server sends the retrieved article content to the generation AI, which analyzes the article content, creates a summary, and then generates video and audio content for the VTuber character. This process includes real-time simulation of the character's facial expressions and movements, as well as voice synthesis.

[0727] Emotion Engine Operation

[0728] The emotion engine analyzes the user's facial expressions and voice data in real time to recognize their emotions. The recognized emotion data is fed back to the generation AI, which then adjusts the content of the video and audio content based on this.

[0729] Delivery of video and audio content

[0730] The server prepares the video and audio content received from the generation AI so that it can be provided in streaming format. It configures the content for distribution through the streaming server and sends it to the device in real time. The user then begins watching the content on their device.

[0731] Chat feature

[0732] Users can send comments through a chat interface while watching. The comments are sent to the server in real time. The server then passes the chat comments received from the user to an emotion engine, which analyzes the user's emotions. Based on the results, the generative AI generates an appropriate response and returns it to the user via the server.

[0733] Specific examples

[0734] 1. The user selects a news article.

[0735] The terminal displays a list of news articles and the user selects an article in the "Technology" category.

[0736] 2. The server passes the article to the generation AI.

[0737] The server passes the selected article to a generation AI to generate video and audio content.

[0738] 3. Use of Emotion Engine

[0739] While a user is watching a video, the emotion engine recognizes emotions from the user's facial expressions and voice, and the generative AI adjusts the content in real time based on that.

[0740] 4. Distribution begins

[0741] The server distributes the generated content to the user through streaming, and the user views the content.

[0742] 5. Use of chat function

[0743] Users send comments during the broadcast, the server analyzes the user's emotions using an emotion engine, and the generative AI generates an appropriate response and returns it to the user.

[0744] The present invention provides a new means of enjoying the news visually and aurally, allowing users to not only react to the news interactively, but also providing appropriate responses and content according to the user's emotions.

[0745] The processing flow will be explained below.

[0746] The present invention is a system that combines a system that acquires news articles in real time and converts them into video and audio content that provides visual and auditory enjoyment to users with an emotion engine that recognizes user emotions. The processing flow of the present invention will be explained below by dividing it into specific steps.

[0747] Step 1:

[0748] The server accesses the news provider's API to retrieve the latest news articles, sends an API request, and receives the response data in JSON format.

[0749] Step 2:

[0750] The server analyzes the JSON data of the retrieved news article, extracts the necessary information (title, content, URL, etc.), and stores it in an internal database.

[0751] Step 3:

[0752] The device displays a list of news articles to the user on an interface, including the title of each article, a thumbnail image, and a brief summary.

[0753] Step 4:

[0754] The user selects an article of interest from the displayed news article list, and the ID (or URL) of the selected article is sent to the server via the terminal.

[0755] Step 5:

[0756] The server retrieves the corresponding article content from an internal database based on the article ID received from the user.

[0757] Step 6:

[0758] The server sends the retrieved article content to the generation AI module, which analyzes the article content, creates a summary, and generates video and audio content for the VTuber character.

[0759] Step 7:

[0760] The generative AI simulates the facial expressions and movements of a VTuber character in real time based on a news article, and also generates audio reading the article content and integrates it into the video.

[0761] Step 8:

[0762] The server prepares the video and audio content received from the generation AI so that it can be provided in streaming format. It sets up distribution through the streaming server.

[0763] Step 9:

[0764] The server streams the prepared video and audio content in real time to the terminal, and the user begins watching.

[0765] Step 10:

[0766] The emotion engine analyzes the user's facial expressions and voice in real time to recognize their emotions, and the recognized emotion data is fed back to the generation AI.

[0767] Step 11:

[0768] The generative AI adjusts the VTuber character's facial expressions and movements based on data from the emotion engine, dynamically changing video and audio content.

[0769] Step 12:

[0770] Users can submit comments through a chat interface while watching, and the comments are sent to the server in real time.

[0771] Step 13:

[0772] The server passes the chat comments received from the user to the emotion engine, which analyzes the user's emotions. Based on the results, the generative AI generates an appropriate response.

[0773] Step 14:

[0774] The responses generated by the AI ​​are returned to the user via the server, and are displayed in a chat interface, allowing the user to enjoy the news interactively.

[0775] Through these steps, news articles are converted into video and audio content in real time, and content tailored to the user's emotions is provided, allowing users to enjoy the news visually and audibly while also interacting with it.

[0776] Example 2

[0777] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0778] In recent years, there has been an increase in the number of ways to enjoy news visually and aurally, but users still rely on general text articles. This makes news viewing one-way and lacks an interactive experience. Furthermore, the inability to analyze emotions based on the news content or respond in real time means that a personalized viewing experience is not provided.

[0779] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0780] In this invention, the server includes means for acquiring text news in real time, means for receiving news articles selected by users and passing them to a generation artificial intelligence, means for the generation artificial intelligence to generate video and audio content based on the news articles, means for streaming the generated video and audio content to users, means for receiving chat comments from users and generating appropriate responses, and means for analyzing user emotions in real time using an emotion engine and adjusting content based on the analysis results. This allows users to enjoy the news visually and audibly and receive responses interactively. Furthermore, personalized content can be provided based on emotion analysis.

[0781] "Text news" refers to news articles that consist only of text information obtained from news sources.

[0782] A "server" is a computer system that provides data and services to multiple clients over a network.

[0783] "Users" are consumers who utilize the system to select news articles and view video and audio content.

[0784] "Generative artificial intelligence" refers to machine learning models and algorithms for generating video and audio content based on retrieved news articles.

[0785] "Video and audio content" means material for visual and auditory enjoyment generated by generative artificial intelligence.

[0786] "Streaming distribution" is a technology that transfers data sequentially and plays it back in real time.

[0787] "Chat comments" are text messages that users type and send in real time.

[0788] An "appropriate response" is an appropriate reply provided by a generative AI or system in response to a chat comment from a user.

[0789] An "emotion engine" is a system or algorithm that analyzes a user's facial expressions and voice data to recognize emotions.

[0790] An "interface" is a screen or operating means through which a user interacts with a system.

[0791] "Real-time" refers to the time nature of data and processing occurring immediately without delay.

[0792] The present invention is a system that acquires news articles in real time, converts them into video and audio content, and delivers them to users, and also combines it with an emotion engine that recognizes the user's emotions. This allows users to enjoy the news visually and aurally, and further enables them to receive responses that correspond to their emotions through an interactive experience. The following describes an embodiment of the present invention.

[0793] System configuration

[0794] 1. Server

[0795] The server has the ability to retrieve news articles from news sources in real time, use the news source's API to obtain the latest article list, and analyze and store the data.

[0796] It also accepts news article selection requests from users and passes the selected article content to the generation artificial intelligence (generation AI).

[0797] The AI ​​generator generates video and audio content for virtual characters based on received news articles. The generated content is temporarily stored on the server.

[0798] The server has the function of distributing the generated video and audio content to users in streaming format.

[0799] Furthermore, it has the ability to receive chat comments sent by users, recognize the user's emotions through an emotion engine, and then use the generative AI to generate an appropriate response based on that.

[0800] 2. Terminal

[0801] The terminal provides a user interface and has the function of displaying a list of news articles. When the user selects an article of interest, the selection information is sent to the server.

[0802] It has the ability to play distributed video and audio content in real time.

[0803] It also provides an interface for users to send comments using a chat function.

[0804] 3. Users

[0805] The user selects news articles of interest through the terminal interface.

[0806] Selected news articles are viewed as real-time video and audio content.

[0807] You can send chat comments during the broadcast and receive responses from the server.

[0808] 4. Emotion Engine

[0809] The emotion engine has the ability to recognize emotions from the user's facial expressions and voice, allowing it to grasp the user's emotions in real time.

[0810] The emotion engine has the ability to have the generative AI adjust the content of video and audio content based on the user's emotions.

[0811] Specific examples

[0812] For example, each component operates as follows:

[0813] 1. Select a news article

[0814] The user browses the list of news articles on the device interface and selects an article in the "Technology" category.

[0815] 2. Content generation using generative AI

[0816] The article selected by the user is sent to the server via the device, and the server passes the article to the AI ​​generator, which then generates video and audio content read by a virtual character.

[0817] 3. Emotion analysis using an emotion engine

[0818] While the user is watching the video, the emotion engine recognizes emotions from the user's facial expressions and voice, and this information is fed back to the generative AI, which then adjusts the content in real time.

[0819] 4. Streaming

[0820] The server distributes the generated content to the user in streaming format, and the user watches it in real time on the terminal.

[0821] 5. Chat Comments and Response Generation

[0822] Users can submit comments while watching, which are sent to the server, which uses an emotion engine to analyze the user's emotions, and the generative AI generates an appropriate response, which is then sent back to the user via the server.

[0823] Prompt Sentence Examples

[0824] "Create a video and audio piece that briefly summarizes the contents of a news article and has a virtual character read it aloud."

[0825] In this way, users can enjoy the news visually and audibly, while simultaneously receiving interactive responses. Personalized content based on sentiment analysis is also provided.

[0826] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0827] Step 1: Get news articles

[0828] The server calls the news provider's API to retrieve the latest news articles. As input, it sends a request with the API endpoint and required parameters. As output, it receives the news data in JSON format, parses it, and extracts the required information (title, content, URL, etc.). It then stores this data in an internal database.

[0829] Specific behavior:

[0830] The server sends a request to the News API using "https: / / newsapi.org / v2 / top-headlines?country=jp&apiKey=Your API Key".

[0831] The returned JSON data is analyzed, and the article title, content, URL, etc. are extracted and saved in the database.

[0832] Step 2: View and select a news article

[0833] The terminal displays a list of news articles. As input, it receives news article data retrieved from the server. As output, it displays the news article list in a user interface. The user selects news articles of interest and sends the selection information to the server.

[0834] Specific behavior:

[0835] The terminal receives the news article list sent from the server and displays it on the interface.

[0836] The user selects an article in the "technology" category, and the terminal transmits this selection information to the server.

[0837] Step 3: Send to article generation AI

[0838] The server receives the user's selection information. As input, it receives the ID and link of the news article selected by the user. As output, it retrieves the corresponding article content from the internal database and sends it to the generation AI.

[0839] Specific behavior:

[0840] The server retrieves relevant news articles from the database based on the selection information sent by the user.

[0841] The retrieved news articles are sent in text format to the generation AI.

[0842] Step 4: Content generation with generative AI

[0843] The generative AI generates video and audio content based on the submitted article. It receives the text data of the news article as input. It generates video and audio content of a virtual character as output. Specifically, it analyzes the article content, generates a summary, simulates the character's facial expressions and movements, and synthesizes voice.

[0844] Specific behavior:

[0845] The generative AI analyzes news articles, creates summaries, and generates videos with the content, "A new technology has been released. This technology is..."

[0846] The character's movements and facial expressions are specified, voice synthesis is performed, and the generated content is sent to the server.

[0847] Step 5: Prepare your content for distribution

[0848] The server receives the video and audio content generated by the generation AI. As input, it receives the video and audio data from the generation AI. As output, it converts these data into a streaming format and prepares it for distribution to the user's device through the streaming server.

[0849] Specific behavior:

[0850] The server receives the video and audio files from the generated AI and transfers them to the streaming server.

[0851] Configure streaming settings and generate a streaming link.

[0852] Step 6: Stream your content

[0853] The server streams the generated content to the user's device. As input, the server sends the generated content data to the streaming server and provides the streaming link to the user. As output, the server provides the content to be played in real time on the user's device.

[0854] Specific behavior:

[0855] The server provides a streaming link of the generated video content to the user's terminal.

[0856] Users can watch video and audio content in real time on their devices.

[0857] Step 7: Sending chat comments and generating replies

[0858] Users can send chat comments while watching a video. They input their comments using the device interface and send them to the server.

[0859] The server receives the sent chat comments and analyzes them using the emotion engine. As input, it receives chat comments from users. As output, it feeds back the analysis results to the generation AI, which generates an appropriate response.

[0860] Specific behavior:

[0861] The user sends a comment from the device saying, "This news is amazing!"

[0862] The server receives the comments and analyzes the emotion of "surprise" using an emotion engine.

[0863] The generative AI generates a response such as "Yes, modern technology is truly amazing," and the server sends that response to the user.

[0864] In this way, users can enjoy the news visually and audibly, while simultaneously receiving interactive responses. Personalized content based on sentiment analysis is also provided.

[0865] (Application example 2)

[0866] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0867] Conventional news delivery systems were capable of acquiring and providing news text information in real time, but were limited in the means to provide it as visually and aurally enjoyable video and audio content. Furthermore, they were not able to recognize users' emotions and generate interactive content or responses accordingly. Furthermore, while many users often feel the need for appropriate responses or feedback based on their emotions while watching the news, no system existed that could meet these needs.

[0868] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for acquiring text news in real time, means for receiving a news article selected by the user and passing it to the generation AI, means for the generation AI to generate video and audio content based on the news article, means for streaming the generated video and audio content to the user, means for receiving chat comments from the user and generating an appropriate response, and means for recognizing the user's emotions and adjusting the content based on the emotion. This enables highly real-time news delivery that can be enjoyed visually and audibly, and makes it possible to provide appropriate responses and interactive experiences according to the user's emotions.

[0869] "Text news" refers to news articles in written form distributed in real time over the Internet or through other communication means.

[0870] "User" refers to an individual or group of people who utilize the system of the present invention to select, view, and interact with news stories.

[0871] "Generative AI" refers to AI that has the ability to generate video and audio content based on input news articles.

[0872] "Video and audio content" means media that provides information through visual and audio means, including generated news feeds.

[0873] "Streaming distribution" is a method of providing video and audio content to users in real time over the Internet.

[0874] "Chat comments" refer to text messages that users enter while viewing content.

[0875] An "appropriate response" is an automatically generated reply or feedback that is generated in response to a chat comment or emotion from a user.

[0876] "Emotion recognition" refers to the process of analyzing and understanding a user's emotional state from data such as facial expressions and voice.

[0877] "Adjusting content" refers to changing the content or format of video and audio content in real time based on recognized user sentiment.

[0878] The present invention combines a system that acquires news articles in real time, converts them into video and audio content, and distributes them to users with an emotion engine that recognizes user emotions. Specific embodiments for carrying out the present invention will be described below.

[0879] System Configuration

[0880] The present invention is mainly composed of a server, a terminal, and a user. Each component will be explained below.

[0881] server

[0882] The server has the following functions:

[0883] 1. Real-time text news acquisition:

[0884] The server uses the news provider's API to retrieve the latest articles in real time. It sends an API request and receives the response data in JSON format. It extracts the necessary information (title, content, URL, etc.) and stores it in an internal database.

[0885] 2. Select news articles and pass them to the generation AI:

[0886] The system receives information about a news article selected by the user on their device and passes it to a generative artificial intelligence (generative AI). The generative AI generates video and audio content based on the article content. This generation process includes creating a summary of the news article, generating video of the VTuber character, and synthesizing voice.

[0887] 3. Emotion recognition and content adjustment:

[0888] The emotion engine recognizes emotions from the user's facial expressions and voice in real time and feeds the recognized emotion data back to the generation AI, which then adjusts the content of the video and audio content based on this feedback information.

[0889] 4. Streaming:

[0890] The generated video and audio content is distributed to the user's terminal in streaming format.

[0891] 5. Handling chat comments:

[0892] The system receives chat comments from users and analyzes their emotions with an emotion engine. Based on the results, the generative AI generates an appropriate response and returns it to the user via the server.

[0893] Terminal

[0894] The terminal has the following functions:

[0895] 1. Select a news article:

[0896] It provides a user interface to display a list of news articles, and when the user selects an article of interest, it sends the selection information to the server.

[0897] 2. Viewing content:

[0898] It has the ability to play distributed video and audio content in real time.

[0899] 3. Chat function:

[0900] It provides an interface that allows users to submit comments while watching, which are then sent to the server, which returns an appropriate response.

[0901] User

[0902] The user performs the following actions:

[0903] 1. Select a news article:

[0904] Select the news article of interest through the device interface.

[0905] 2. Content viewing and emotional feedback:

[0906] The selected news article is then viewed as real-time generated video and audio content, with facial expressions and voice acting analyzed by the emotion engine as the news is viewed.

[0907] 3. Sending chat comments:

[0908] Send comments while watching and receive responses from the generative AI.

[0909] Hardware and software used

[0910] On the server side, we use a server equipped with a high-performance GPU. The main software used is an API for retrieving news from news sources, generative AI (e.g., OpenAI's GPT-4), and an emotion recognition engine (e.g., DeepFace, OpenFace).

[0911] On the terminal side, mobile devices such as smartphones and tablets are used, which are equipped with cameras and microphones to collect data for emotion recognition.

[0912] Specific examples

[0913] 1. View news article list:

[0914] The device displays a list of the latest news, and the user selects "an article about the latest AI technology."

[0915] 2. Content Creation and Viewing:

[0916] The server uses AI to convert the selected articles into video and audio content, which are then streamed. Users can then watch the generated VTuber videos on their devices.

[0917] 3. Emotion Recognition and Feedback:

[0918] While watching, the user's facial expressions are analyzed by a camera, and a "happy" expression is recognized. This information is fed back to the generative AI, which then adjusts the video and audio content to better match the user's emotions.

[0919] 4. Chat comment response:

[0920] When a user types "This technology is amazing!" into the chat, the emotion engine analyzes the emotion and the generative AI responds appropriately with "Yes, the evolution of AI is truly amazing!"

[0921] Prompt Sentence Examples

[0922] "News article: 'The latest AI technology has been announced. This technology is...'"

[0923] "Summary: Please summarize this article."

[0924] "Video and Audio: Generate a script in which a VTuber character explains this news article in an easy-to-understand way."

[0925] The above is a specific embodiment for carrying out the present invention.

[0926] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0927] Step 1:

[0928] The server uses the news provider's API to obtain the latest news articles in real time. It sends an API request and receives the response data in JSON format. The server extracts necessary information such as the news article title, content, and URL, and stores it in an internal database. The input is the API request, and the output is the analyzed news article data.

[0929] Step 2:

[0930] The user selects an article of interest from the displayed news article list using the terminal interface. The selection information is sent from the terminal to the server. The input is the user's selection action, and the output is the selected news article information.

[0931] Step 3:

[0932] The server passes data based on the selected news article to the generation AI. The generation AI then analyzes the article content, creates a summary, and uses that to generate video and audio content for the VTuber character. The input is the selected news article data, and the output is the generated video and audio content.

[0933] Step 4:

[0934] The emotion engine starts up and analyzes the user's facial expressions and voice in real time. Data is acquired using the device's camera and microphone, and the emotion engine analyzes that data to recognize the user's emotions. The input is the user's facial and voice data, and the output is the recognized emotion data.

[0935] Step 5:

[0936] The recognized emotional data is then fed back to the generation AI, which then adjusts the content based on the fed-back emotional data. Specifically, it changes the character's facial expression and tone according to the user's emotions. The input is the recognized emotional data, and the output is the adjusted content.

[0937] Step 6:

[0938] The server distributes the generated video and audio content to the user's device in streaming format. This is done in real time through a streaming server. The input is the adjusted content, and the output is the video and audio to be distributed.

[0939] Step 7:

[0940] While watching a video, users can send comments using a chat interface. The comments are sent from the device to the server. The input is the user's chat comments, and the output is the comment data received by the server.

[0941] Step 8:

[0942] The server passes the received chat comments to the emotion engine for analysis. The emotion engine analyzes emotions based on the user's comments and feeds the results back to the generation AI. The input is the chat comments, and the output is the analyzed emotion data.

[0943] Step 9:

[0944] The generative AI generates an appropriate response based on the analyzed emotional data and returns it to the user through the server. The response is generated using a prompt sentence. The input is the analyzed emotional data, and the output is the response message generated by the AI.

[0945] Example prompt

[0946] "News article: 'The latest AI technology has been announced. This technology is...'"

[0947] "Summary: Please summarize this article."

[0948] "Video and Audio: Generate a script in which a VTuber character explains this news article in an easy-to-understand way."

[0949] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0950] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0951] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0952] [Third embodiment]

[0953] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0954] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0955] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0956] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0957] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0958] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0959] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0960] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0961] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0962] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0963] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0964] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0965] The present invention is a system that acquires news articles in real time, converts them into video and audio content, and distributes them to users. This allows users to enjoy the news visually and aurally, and also allows them to interact with the news. The following describes an embodiment of the present invention.

[0966] System Configuration

[0967] 1. Server

[0968] The server has the ability to retrieve news articles from news sources in real time, obtain the latest article list using the news source's API, and store the data.

[0969] It also accepts news article selection requests from users and passes the selected article content to the generation AI.

[0970] The artificial intelligence has the ability to generate video and audio content for VTuber characters based on received news articles. The generated content is temporarily stored on the server.

[0971] The server has the function of distributing the generated video and audio content to users in streaming format.

[0972] In addition, it has the ability to receive chat comments sent by users, use generative AI to generate appropriate responses, and return them to the users.

[0973] 2. Terminal

[0974] The terminal provides a user interface and has the function of displaying a list of news articles. When the user selects an article of interest, the selection information is sent to the server.

[0975] It has the ability to play distributed video and audio content in real time.

[0976] It also provides an interface for users to send comments using a chat function.

[0977] 3. Users

[0978] The user selects news articles of interest through the terminal interface.

[0979] Selected news articles are viewed as real-time video and audio content.

[0980] You can send chat comments during the broadcast and receive responses from the server.

[0981] Program processing flow

[0982] Get news articles

[0983] The server calls the news provider's API to retrieve the latest news articles, parses them in JSON format, and saves the necessary information.

[0984] Accepting article conversion requests

[0985] The user selects news articles of interest from the terminal interface and the selection is transmitted to the server.

[0986] Content generation by generative AI

[0987] The server passes the received news article to the generation AI, which generates video and audio content for the VTuber character based on the article content. This includes creating a summary of the article, simulating the character's facial expressions and movements, and synthesizing voice.

[0988] Delivery of video and audio content

[0989] The server transmits the generated content to the user terminal via a streaming server, and the user watches the received content in real time.

[0990] Chat feature

[0991] Users can send chat comments in real time during the broadcast, and the server passes the received comments to a generation AI that generates an appropriate response, which is returned to the user via the chat interface.

[0992] Specific examples

[0993] 1. The user selects a news article.

[0994] The terminal displays a list of news articles and the user selects an article in the "Technology" category.

[0995] 2. The server passes the article to the generation AI.

[0996] The server passes the selected article to a generation AI to generate video and audio content.

[0997] 3. Distribution begins

[0998] The server distributes the generated content to the user through streaming, and the user views the content.

[0999] 4. Use of chat function

[1000] Users submit comments during a broadcast, and the server generates a response and sends it back to the user.

[1001] As described above, the present invention provides a new means for enjoying news visually and aurally, and allows users to react to the news interactively.

[1002] The processing flow will be explained below.

[1003] Step 1:

[1004] The server accesses the news provider's API to retrieve the latest news articles, sending an API request and receiving the response data in JSON format.

[1005] Step 2:

[1006] The server parses the JSON data of the retrieved news article, extracts the necessary information (title, content, URL, etc.), and stores it in an internal database or cache.

[1007] Step 3:

[1008] The device displays a list of news articles to the user on an interface, including the title of each article, a thumbnail image, and a brief summary.

[1009] Step 4:

[1010] The user selects an article of interest from the displayed news article list, and the ID (or URL) of the selected article is sent to the server via the terminal.

[1011] Step 5:

[1012] The server retrieves the corresponding article content from an internal database based on the article ID received from the user.

[1013] Step 6:

[1014] The server sends the retrieved article content to the generation AI module, which analyzes the article content, creates a summary, and then generates video and audio content for the VTuber character.

[1015] Step 7:

[1016] The generative AI simulates the facial expressions and movements of a VTuber character in real time based on a news article, and also generates audio reading the article content and integrates it into the video.

[1017] Step 8:

[1018] The server prepares the video and audio content received from the generation AI so that it can be provided in streaming format, and configures it for distribution through a streaming server.

[1019] Step 9:

[1020] The server then streams the prepared video and audio content in real time to the device, where the user can begin watching.

[1021] Step 10:

[1022] Users can submit comments through a chat interface while watching, and the comments are sent to the server in real time.

[1023] Step 11:

[1024] The server passes chat comments received from users to the generation AI and asks it to generate an appropriate response. The generation AI creates a response based on the content of the comment.

[1025] Step 12:

[1026] The response generated by the generation AI is sent to the server, which returns the response to the user through a chat interface.

[1027] Through these steps, news articles are transformed into video and audio content in real time and delivered interactively, allowing users to enjoy the news visually and audibly and react to it in real time.

[1028] Example 1

[1029] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1030] Users need to maximize their visual and auditory capabilities when accessing text data, but traditional methods do not adequately achieve this. Furthermore, users have limited options for interacting with the data in real time and engaging with it. Therefore, news and other text data must be delivered to users in a more interactive and engaging format.

[1031] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1032] In this invention, the server includes means for acquiring text data in real time, means for receiving data selected by the user and passing it to the generation AI, means for the generation AI to generate video and audio data based on the data, means for streaming the generated video and audio data to the user, and means for receiving comments from the user and generating appropriate responses, allowing the user to enjoy the text data of interest visually and audibly and to react interactively in real time.

[1033] "Text data" is digital data containing text information.

[1034] "Generative AI" is a system that uses artificial intelligence technology to generate moving images and audio data based on text data.

[1035] "Motion picture data" is data in a digital format that shows a sequence of images and provides visual information.

[1036] "Audio data" means data in a digital format that conveys information through hearing.

[1037] "Streaming distribution" is a technology that distributes digital media over the Internet in real time.

[1038] A "comment" is an opinion or question in text format that a user sends in real time in response to video or audio content.

[1039] MODE FOR CARRYING OUT THE INVENTION

[1040] This invention is a system that acquires text data in real time, generates video and audio data based on that data, and distributes them to users. This system allows users to enjoy the text data visually and audibly, and also allows them to respond interactively in real time.

[1041] Hardware and software used

[1042] 1. Server

[1043] It is used to call the news provider's API to obtain the latest news articles. This API can be from a general API service provider, for example.

[1044] Examples of generative AI include OpenAI's GPT-4 and similar AI models.

[1045] DeepFake technology is used to generate moving images, and Google Text-to-Speech and other voice synthesis engines are used.

[1046] Streaming services such as YouTube Live and Twitch are used for streaming.

[1047] 2. Terminal

[1048] It displays a list of news articles and provides a user interface for users to select articles they are interested in. This interface is implemented as a web application or smartphone app.

[1049] Processing flow

[1050] The server retrieves the latest news articles in real time using the news provider's API. The retrieved data is sent to the server in JSON format and analyzed. The analyzed article information is saved in the server's database. Specifically, for example, the server calls the endpoint of a service called "NewsAPI," sends a "GET" request, parses the retrieved JSON data, extracts the article title, text, date, etc., and saves them in the database.

[1051] The user browses through a list of news articles through the interface on the device and selects an article of interest. The ID information of the selected article is sent from the device to the server. For example, the user opens a news app on a smartphone or PC and selects an article in the "Technology" category. The device then sends the selection information to the server.

[1052] The server passes the selected article content to the generation AI. The generation AI generates video and audio content for the VTuber character based on the article content. This includes creating a summary of the article, simulating the character's facial expressions and movements, and processing voice synthesis. Specifically, it sends a prompt to the generation AI model, instructing it to summarize the article content. An example of a prompt is, "Please summarize the following news article: '(news article body)'."

[1053] The server sends the generated video and audio content to the user terminal via the streaming server, and the user watches the content in real time. Specifically, the server uploads the generated video file to the streaming server and sends a streaming URL to the user terminal. The user watches the content in real time via the URL.

[1054] Users can send chat comments in real time while watching. The server passes the received comments to the generation AI, which generates an appropriate response. The generated response is returned to the user via the chat interface. Specifically, the user sends a comment during the broadcast, and the server passes the comment to the generation AI, which generates a response. An example of a prompt is, "Please generate an appropriate response to the user's comment: 'Please tell us a specific application example of this technology.'"

[1055] In this way, the system provides a new means of visually and aurally enjoying text data, allowing users to respond interactively.

[1056] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1057] Step 1: Get news articles

[1058] The server calls the news provider's API to retrieve the latest news articles in real time. The retrieved data is sent to the server in JSON format and parsed. The retrieved JSON data includes information such as the article title, body text, and date. This data is then stored in a database.

[1059] Input: News source API endpoint

[1060] Output: Parsed news article data (title, body, date, etc.)

[1061] Specifically, for example, it uses the endpoint of a typical API service provider to send a "GET" request and parses the retrieved JSON data.

[1062] Step 2: Select an article

[1063] The user selects an article of interest from the list of news articles displayed through the interface on the terminal, and the ID information of the selected article is sent from the terminal to the server.

[1064] Input: The ID of the news article selected by the user

[1065] Output: ID information of selected articles sent to the server

[1066] Specifically, a user opens a news app on their smartphone or PC and selects an article in the "Technology" category, for example. At that time, the device sends a request including the ID of the selected article to the server.

[1067] Step 3: Submitting a content generation request

[1068] The server passes the selected article content to the generation AI, which then generates video and audio content for the VTuber character based on the article content, including creating a summary of the article, simulating the character's facial expressions and movements, and processing voice synthesis.

[1069] Input: Content of selected news article

[1070] Output: Prompt text passed to the generation AI

[1071] Specifically, the system sends a prompt to a generative AI model (e.g., GPT-4) to instruct it to summarize the article. An example prompt might be, "Please summarize the following news article: '(news article text)'."

[1072] Step 4: Generate video and audio content

[1073] The AI ​​generates video and audio content for the VTuber character based on the article content provided, including simulating the character's facial expressions and movements based on the summarized article content and generating audio using a speech synthesis engine.

[1074] Input: Summary of news article content

[1075] Output: The generated video and audio content

[1076] Specifically, it works by using a character simulation tool that uses DeepFake technology to generate voice using a voice synthesis engine (e.g., Google Text-to-Speech).

[1077] Step 5: Deliver your content

[1078] The server transmits the generated video and audio content to the user terminal via a streaming server, where the user can view the content in real time.

[1079] Input: Generated video and audio content

[1080] Output: Streaming URL sent in real time

[1081] Specifically, the generated video file is uploaded to a streaming server (such as YouTube Live or Twitch), and a streaming URL is sent to the user's device. The user can then watch the video in real time via that URL.

[1082] Step 6: Implementing the chat function

[1083] Users can send chat comments in real time while watching, and the server passes the received comments to a generation AI that generates an appropriate response, which is returned to the user via the chat interface.

[1084] Input: Chat comment sent by the user

[1085] Output: The generated response message

[1086] Specifically, while watching, the user sends a comment such as "Please tell me a specific application example of this technology," and the server passes the comment to the generation AI, which then generates an appropriate response. Example prompt: "Generate an appropriate response to the user's comment: 'Please tell me a specific application example of this technology.'"

[1087] (Application example 1)

[1088] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1089] Current news distribution methods are limited to providing text-based information, leaving many users lacking the means to enjoy news visually and aurally. Furthermore, they lack the ability to enjoy news articles interactively, making it difficult for users to efficiently understand news information. Furthermore, they lack the means to easily select specific news articles that interest them and share their reactions to them. These issues need to be addressed.

[1090] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1091] In this invention, the server includes means for acquiring text news in real time, means for receiving news articles selected by users and passing them to a generation AI, means for the generation AI to generate video and audio content based on the news articles, means for streaming the generated video and audio content to users, means for receiving chat comments from users and generating appropriate responses, and means for being installed in a smartphone application, which allows users to enjoy the news visually and audibly and to react to the news interactively.

[1092] "Means for obtaining text news in real time" refers to a system that obtains the latest news articles from news providers in real time via a dedicated API, and analyzes and stores them.

[1093] "Means for receiving news articles selected by the user and passing them on to the generating AI" refers to the communication and data processing functions for transmitting the news article information selected by the user to the generating AI.

[1094] "Generative AI" refers to AI technology for generating video and audio content based on received news articles. Specifically, it includes summarization, character facial expression and movement simulation, and voice synthesis.

[1095] "Means for generating video and audio content" refers to the mechanism by which the generative AI creates visual and audio multimedia content based on the content of news articles.

[1096] The "means for streaming the generated video and audio content to the user" refers to a streaming technology for delivering the generated multimedia content to the user terminal in real time.

[1097] "Means for receiving chat comments from users and generating appropriate responses" refers to a mechanism in which comments sent by users through the chat function are analyzed and the generation AI generates and replies to appropriate responses.

[1098] "Means installed in a smartphone application" refers to smartphone-specific software that incorporates the various functions mentioned above and provides users with an interactive news experience.

[1099] The "interface means for displaying a list of news articles and allowing the user to select an article of interest" is a user interface function that displays a list of news articles and allows the user to select a particular article.

[1100] This invention is a system that acquires news articles in real time, converts them into video and audio content, and distributes them to users. This system is composed of multiple components, such as a server, terminals, and generation AI, and its detailed configuration and processing are described below.

[1101] System Configuration

[1102] 1. Server

[1103] The server can obtain news articles from news providers in real time. This includes functions to obtain the latest article list using the news provider's API and store that data. It also accepts news article selection requests from users and passes the selected article content to the generation AI. The generation AI has the function to generate video and audio content based on the received news articles. The generated content is temporarily stored on the server and delivered to users in streaming format. It also receives chat comments sent by users, uses the generation AI to generate appropriate responses, and returns them to the user.

[1104] 2. Terminal

[1105] The device provides a user interface and can display a list of news articles. When a user selects an article of interest, the device transmits the selection information to the server. It can also play the distributed video and audio content in real time. It also provides an interface for users to send comments using a chat function. The device is installed as a smartphone application.

[1106] 3. Users

[1107] Users select news articles of interest through the device interface, and can view the selected news articles as video and audio content streamed in real time. Users can also send chat comments during the stream and receive responses from the server.

[1108] Program processing flow explanation

[1109] Get news articles

[1110] The server calls the news provider's API to retrieve the latest news articles. The retrieved articles are parsed in JSON format and the necessary information is saved. The server uses Node.js and Express, and an HTTP client such as Axios is used to communicate with the news provider.

[1111] Accepting article conversion requests

[1112] The device displays a list of news articles, and the user selects the news article they are interested in. The selection information is sent to the server. This process is implemented in a smartphone application using React Native.

[1113] Content generation by generative AI

[1114] The server passes the received news article to the generation AI, which generates video and audio content for the VTuber character based on the article content. This includes creating a summary of the article, simulating the character's facial expressions and movements, and synthesizing voice. The generation AI uses GPT-4 and deep learning frameworks such as TensorFlow and PyTorch.

[1115] Delivery of video and audio content

[1116] The server sends the generated content to the user's device via a streaming server (e.g., Wowza Streaming Engine), and the user watches the received content in real time.

[1117] Chat feature

[1118] Users can send chat comments in real time during the broadcast, and the server passes the received comments to the generation AI, which generates an appropriate response, which is returned to the user via the chat interface.

[1119] Examples of concrete examples and prompts

[1120] Specific examples

[1121] The user selects an article in the "Technology" category from a list of news articles. The server retrieves an article about the latest technology for self-driving cars from the news source and passes it to the generation AI. The generation AI summarizes the article and generates a video in which a VTuber character explains it. The user watches the generated video on a smartphone app and asks in chat, "Is the technology introduced here also used by other manufacturers?" The generation AI replies, "Yes, this technology is being adopted by many automakers."

[1122] Example prompts for generative AI models

[1123] Article content: Article about the latest technology in autonomous vehicles

[1124] Title: Evolution of in-vehicle AI technology in 2023

[1125] Article text: Advances in AI technology are enabling the latest self-driving cars to operate more safely and efficiently. In particular, real-time data processing combined with advanced sensor technology has improved the car's ability to accurately perceive its surroundings...

[1126] Prompt: Based on the article below, create a 5-minute news commentary video with a VTuber character.

[1127] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1128] Step 1:

[1129] The server retrieves the latest news articles in real time through the news provider's API.

[1130] How it works: The server periodically sends GET requests to the API endpoint to receive news article data in JSON format, which is then parsed to extract the necessary information (title, body, date, etc.) and store it in a database such as MongoDB.

[1131] Input: JSON data from the news source API.

[1132] Output: News article data stored in a database.

[1133] Step 2:

[1134] The terminal displays a list of news articles to the user.

[1135] Specific operation: The smartphone application on the device displays a list of news articles retrieved from the server on the user interface. The user selects the news article of interest. The selection information is sent from the device to the server.

[1136] Input: A list of news articles retrieved from the server.

[1137] Output: Information about the news article selected by the user.

[1138] Step 3:

[1139] The server passes the news articles selected by the user to the generating artificial intelligence.

[1140] Specific operation: The server calls an internal API to pass the selected news article to the generation AI module, which generates an appropriate prompt. The generation AI processes and analyzes the data based on the prompt and the news article content.

[1141] Input: Information about the news article selected by the user, and a prompt.

[1142] Output: The prompt and news article passed to the generation AI.

[1143] Step 4:

[1144] Generative AI generates video and audio content based on news articles.

[1145] How it works: The generative AI generates content based on prompts, summarizes articles, simulates the facial expressions and movements of VTuber characters, and performs voice synthesis. This results in the creation of visual and audio news commentary videos. The technologies used include GPT-4, TensorFlow, and PyTorch.

[1146] Input: Prompt text and news article content.

[1147] Output: The generated video and audio content.

[1148] Step 5:

[1149] The server streams the generated video and audio content to the user.

[1150] Specific operation: The generated video and audio content is temporarily stored on the server and then distributed in real time to the user's device via a streaming server (e.g., Wowza Streaming Engine).

[1151] Input: Generated video and audio content.

[1152] Output: The streaming content sent to the user device.

[1153] Step 6:

[1154] Users can view the distributed video and audio content and send chat comments.

[1155] Specific operation: A user watches a live video on a smartphone application and sends comments using the chat interface. The comments are sent to the server.

[1156] Input: The user's chat comment.

[1157] Output: Chat comments sent to the server.

[1158] Step 7:

[1159] The server passes the received chat comments to a generation artificial intelligence to generate an appropriate response.

[1160] Specific operation: The server analyzes the chat comments and passes them to the generation AI to generate an appropriate response, which is then sent back to the device and displayed to the user.

[1161] Input: The user's chat comment.

[1162] Output: The chat response generated by the generative AI.

[1163] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1164] The present invention is a system that acquires news articles in real time, converts them into video and audio content, and delivers them to users, and also combines it with an emotion engine that recognizes the user's emotions. This allows users to enjoy the news visually and aurally, and further enables them to receive responses that correspond to their emotions through an interactive experience. The following describes an embodiment of the present invention.

[1165] System Configuration

[1166] 1. Server

[1167] The server has the ability to retrieve news articles from news sources in real time, use the news source's API to obtain the latest article list, and analyze and store the data.

[1168] It also accepts news article selection requests from users and passes the selected article content to the generation artificial intelligence (generation AI).

[1169] The AI ​​has the ability to generate video and audio content for VTuber characters based on the received news articles. The generated content is temporarily stored on the server.

[1170] The server has the function of distributing the generated video and audio content to users in streaming format.

[1171] Furthermore, it has the ability to receive chat comments sent by users, recognize the user's emotions through an emotion engine, and then use the generative AI to generate an appropriate response based on that.

[1172] 2. Terminal

[1173] The terminal provides a user interface and has the function of displaying a list of news articles. When the user selects an article of interest, the selection information is sent to the server.

[1174] It has the ability to play distributed video and audio content in real time.

[1175] It also provides an interface for users to send comments using a chat function.

[1176] 3. Users

[1177] The user selects news articles of interest through the terminal interface.

[1178] Selected news articles are viewed as real-time video and audio content.

[1179] You can send chat comments during the broadcast and receive responses from the server.

[1180] 4. Emotion Engine

[1181] The emotion engine has the ability to recognize emotions from the user's facial expressions and voice, allowing it to grasp the user's emotions in real time.

[1182] The emotion engine has the ability to have the generative AI adjust the content of video and audio content based on the user's emotions.

[1183] Program processing flow

[1184] Get news articles

[1185] The server calls the news provider's API to retrieve the latest news articles. It sends an API request, receives the response data in JSON format, analyzes it, extracts the necessary information (title, content, URL, etc.), and stores it in an internal database.

[1186] Accepting article conversion requests

[1187] The user selects the news article of interest from the terminal interface, and the selection information is sent to the server via the terminal. The server receives the user's request and retrieves the corresponding article content from its internal database.

[1188] Content generation by generative AI

[1189] The server sends the retrieved article content to the generation AI, which analyzes the article content, creates a summary, and then generates video and audio content for the VTuber character. This process includes real-time simulation of the character's facial expressions and movements, as well as voice synthesis.

[1190] Emotion Engine Operation

[1191] The emotion engine analyzes the user's facial expressions and voice data in real time to recognize their emotions. The recognized emotion data is fed back to the generation AI, which then adjusts the content of the video and audio content based on this.

[1192] Delivery of video and audio content

[1193] The server prepares the video and audio content received from the generation AI so that it can be provided in streaming format. It configures the content for distribution through the streaming server and sends it to the device in real time. The user then begins watching the content on their device.

[1194] Chat feature

[1195] Users can send comments through a chat interface while watching. The comments are sent to the server in real time. The server then passes the chat comments received from the user to an emotion engine, which analyzes the user's emotions. Based on the results, the generative AI generates an appropriate response and returns it to the user via the server.

[1196] Specific examples

[1197] 1. The user selects a news article.

[1198] The terminal displays a list of news articles and the user selects an article in the "Technology" category.

[1199] 2. The server passes the article to the generation AI.

[1200] The server passes the selected article to a generation AI to generate video and audio content.

[1201] 3. Use of Emotion Engine

[1202] While a user is watching a video, the emotion engine recognizes emotions from the user's facial expressions and voice, and the generative AI adjusts the content in real time based on that.

[1203] 4. Distribution begins

[1204] The server distributes the generated content to the user through streaming, and the user views the content.

[1205] 5. Use of chat function

[1206] Users send comments during the broadcast, the server analyzes the user's emotions using an emotion engine, and the generative AI generates an appropriate response and returns it to the user.

[1207] The present invention provides a new means of enjoying the news visually and aurally, allowing users to not only react to the news interactively, but also providing appropriate responses and content according to the user's emotions.

[1208] The processing flow will be explained below.

[1209] The present invention is a system that combines a system that acquires news articles in real time and converts them into video and audio content that provides visual and auditory enjoyment to users with an emotion engine that recognizes user emotions. The processing flow of the present invention will be explained below by dividing it into specific steps.

[1210] Step 1:

[1211] The server accesses the news provider's API to retrieve the latest news articles, sends an API request, and receives the response data in JSON format.

[1212] Step 2:

[1213] The server analyzes the JSON data of the retrieved news article, extracts the necessary information (title, content, URL, etc.), and stores it in an internal database.

[1214] Step 3:

[1215] The device displays a list of news articles to the user on an interface, including the title of each article, a thumbnail image, and a brief summary.

[1216] Step 4:

[1217] The user selects an article of interest from the displayed news article list, and the ID (or URL) of the selected article is sent to the server via the terminal.

[1218] Step 5:

[1219] The server retrieves the corresponding article content from an internal database based on the article ID received from the user.

[1220] Step 6:

[1221] The server sends the retrieved article content to the generation AI module, which analyzes the article content, creates a summary, and generates video and audio content for the VTuber character.

[1222] Step 7:

[1223] The generative AI simulates the facial expressions and movements of a VTuber character in real time based on a news article, and also generates audio reading the article content and integrates it into the video.

[1224] Step 8:

[1225] The server prepares the video and audio content received from the generation AI so that it can be provided in streaming format. It sets up distribution through the streaming server.

[1226] Step 9:

[1227] The server streams the prepared video and audio content in real time to the terminal, and the user begins watching.

[1228] Step 10:

[1229] The emotion engine analyzes the user's facial expressions and voice in real time to recognize their emotions, and the recognized emotion data is fed back to the generation AI.

[1230] Step 11:

[1231] The generative AI adjusts the VTuber character's facial expressions and movements based on data from the emotion engine, dynamically changing video and audio content.

[1232] Step 12:

[1233] Users can submit comments through a chat interface while watching, and the comments are sent to the server in real time.

[1234] Step 13:

[1235] The server passes the chat comments received from the user to the emotion engine, which analyzes the user's emotions. Based on the results, the generative AI generates an appropriate response.

[1236] Step 14:

[1237] The responses generated by the AI ​​are returned to the user via the server, and are displayed in a chat interface, allowing the user to enjoy the news interactively.

[1238] Through these steps, news articles are converted into video and audio content in real time, and content tailored to the user's emotions is provided, allowing users to enjoy the news visually and audibly while also interacting with it.

[1239] Example 2

[1240] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1241] In recent years, there has been an increase in the number of ways to enjoy news visually and aurally, but users still rely on general text articles. This makes news viewing one-way and lacks an interactive experience. Furthermore, the inability to analyze emotions based on the news content or respond in real time means that a personalized viewing experience is not provided.

[1242] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1243] In this invention, the server includes means for acquiring text news in real time, means for receiving news articles selected by users and passing them to a generation artificial intelligence, means for the generation artificial intelligence to generate video and audio content based on the news articles, means for streaming the generated video and audio content to users, means for receiving chat comments from users and generating appropriate responses, and means for analyzing user emotions in real time using an emotion engine and adjusting content based on the analysis results. This allows users to enjoy the news visually and audibly and receive responses interactively. Furthermore, personalized content can be provided based on emotion analysis.

[1244] "Text news" refers to news articles that consist only of text information obtained from news sources.

[1245] A "server" is a computer system that provides data and services to multiple clients over a network.

[1246] "Users" are consumers who utilize the system to select news articles and view video and audio content.

[1247] "Generative artificial intelligence" refers to machine learning models and algorithms for generating video and audio content based on retrieved news articles.

[1248] "Video and audio content" means material for visual and auditory enjoyment generated by generative artificial intelligence.

[1249] "Streaming distribution" is a technology that transfers data sequentially and plays it back in real time.

[1250] "Chat comments" are text messages that users type and send in real time.

[1251] An "appropriate response" is an appropriate reply provided by a generative AI or system in response to a chat comment from a user.

[1252] An "emotion engine" is a system or algorithm that analyzes a user's facial expressions and voice data to recognize emotions.

[1253] An "interface" is a screen or operating means through which a user interacts with a system.

[1254] "Real-time" refers to the time nature of data and processing occurring immediately without delay.

[1255] The present invention is a system that acquires news articles in real time, converts them into video and audio content, and delivers them to users, and also combines it with an emotion engine that recognizes the user's emotions. This allows users to enjoy the news visually and aurally, and further enables them to receive responses that correspond to their emotions through an interactive experience. The following describes an embodiment of the present invention.

[1256] System configuration

[1257] 1. Server

[1258] The server has the ability to retrieve news articles from news sources in real time, use the news source's API to obtain the latest article list, and analyze and store the data.

[1259] It also accepts news article selection requests from users and passes the selected article content to the generation artificial intelligence (generation AI).

[1260] The AI ​​generator generates video and audio content for virtual characters based on received news articles. The generated content is temporarily stored on the server.

[1261] The server has the function of distributing the generated video and audio content to users in streaming format.

[1262] Furthermore, it has the ability to receive chat comments sent by users, recognize the user's emotions through an emotion engine, and then use the generative AI to generate an appropriate response based on that.

[1263] 2. Terminal

[1264] The terminal provides a user interface and has the function of displaying a list of news articles. When the user selects an article of interest, the selection information is sent to the server.

[1265] It has the ability to play distributed video and audio content in real time.

[1266] It also provides an interface for users to send comments using a chat function.

[1267] 3. Users

[1268] The user selects news articles of interest through the terminal interface.

[1269] Selected news articles are viewed as real-time video and audio content.

[1270] You can send chat comments during the broadcast and receive responses from the server.

[1271] 4. Emotion Engine

[1272] The emotion engine has the ability to recognize emotions from the user's facial expressions and voice, allowing it to grasp the user's emotions in real time.

[1273] The emotion engine has the ability to have the generative AI adjust the content of video and audio content based on the user's emotions.

[1274] Specific examples

[1275] For example, each component operates as follows:

[1276] 1. Select a news article

[1277] The user browses the list of news articles on the device interface and selects an article in the "Technology" category.

[1278] 2. Content generation using generative AI

[1279] The article selected by the user is sent to the server via the device, and the server passes the article to the AI ​​generator, which then generates video and audio content read by a virtual character.

[1280] 3. Emotion analysis using an emotion engine

[1281] While the user is watching the video, the emotion engine recognizes emotions from the user's facial expressions and voice, and this information is fed back to the generative AI, which then adjusts the content in real time.

[1282] 4. Streaming

[1283] The server distributes the generated content to the user in streaming format, and the user watches it in real time on the terminal.

[1284] 5. Chat Comments and Response Generation

[1285] Users can submit comments while watching, which are sent to the server, which uses an emotion engine to analyze the user's emotions, and the generative AI generates an appropriate response, which is then sent back to the user via the server.

[1286] Prompt Sentence Examples

[1287] "Create a video and audio piece that briefly summarizes the contents of a news article and has a virtual character read it aloud."

[1288] In this way, users can enjoy the news visually and audibly, while simultaneously receiving interactive responses. Personalized content based on sentiment analysis is also provided.

[1289] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1290] Step 1: Get news articles

[1291] The server calls the news provider's API to retrieve the latest news articles. As input, it sends a request with the API endpoint and required parameters. As output, it receives the news data in JSON format, parses it, and extracts the required information (title, content, URL, etc.). It then stores this data in an internal database.

[1292] Specific behavior:

[1293] The server sends a request to the News API using "https: / / newsapi.org / v2 / top-headlines?country=jp&apiKey=Your API Key".

[1294] The returned JSON data is analyzed, and the article title, content, URL, etc. are extracted and saved in the database.

[1295] Step 2: View and select a news article

[1296] The terminal displays a list of news articles. As input, it receives news article data retrieved from the server. As output, it displays the news article list in a user interface. The user selects news articles of interest and sends the selection information to the server.

[1297] Specific behavior:

[1298] The terminal receives the news article list sent from the server and displays it on the interface.

[1299] The user selects an article in the "technology" category, and the terminal transmits this selection information to the server.

[1300] Step 3: Send to article generation AI

[1301] The server receives the user's selection information. As input, it receives the ID and link of the news article selected by the user. As output, it retrieves the corresponding article content from the internal database and sends it to the generation AI.

[1302] Specific behavior:

[1303] The server retrieves relevant news articles from the database based on the selection information sent by the user.

[1304] The retrieved news articles are sent in text format to the generation AI.

[1305] Step 4: Content generation with generative AI

[1306] The generative AI generates video and audio content based on the submitted article. It receives the text data of the news article as input. It generates video and audio content of a virtual character as output. Specifically, it analyzes the article content, generates a summary, simulates the character's facial expressions and movements, and synthesizes voice.

[1307] Specific behavior:

[1308] The generative AI analyzes news articles, creates summaries, and generates videos with the content, "A new technology has been released. This technology is..."

[1309] The character's movements and facial expressions are specified, voice synthesis is performed, and the generated content is sent to the server.

[1310] Step 5: Prepare your content for distribution

[1311] The server receives the video and audio content generated by the generation AI. As input, it receives the video and audio data from the generation AI. As output, it converts these data into a streaming format and prepares it for distribution to the user's device through the streaming server.

[1312] Specific behavior:

[1313] The server receives the video and audio files from the generated AI and transfers them to the streaming server.

[1314] Configure streaming settings and generate a streaming link.

[1315] Step 6: Stream your content

[1316] The server streams the generated content to the user's device. As input, the server sends the generated content data to the streaming server and provides the streaming link to the user. As output, the server provides the content to be played in real time on the user's device.

[1317] Specific behavior:

[1318] The server provides a streaming link of the generated video content to the user's terminal.

[1319] Users can watch video and audio content in real time on their devices.

[1320] Step 7: Sending chat comments and generating replies

[1321] Users can send chat comments while watching a video. They input their comments using the device interface and send them to the server.

[1322] The server receives the sent chat comments and analyzes them using the emotion engine. As input, it receives chat comments from users. As output, it feeds back the analysis results to the generation AI, which generates an appropriate response.

[1323] Specific behavior:

[1324] The user sends a comment from the device saying, "This news is amazing!"

[1325] The server receives the comments and analyzes the emotion of "surprise" using an emotion engine.

[1326] The generative AI generates a response such as "Yes, modern technology is truly amazing," and the server sends that response to the user.

[1327] In this way, users can enjoy the news visually and audibly, while simultaneously receiving interactive responses. Personalized content based on sentiment analysis is also provided.

[1328] (Application example 2)

[1329] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1330] Conventional news delivery systems were capable of acquiring and providing news text information in real time, but were limited in the means to provide it as visually and aurally enjoyable video and audio content. Furthermore, they were not able to recognize users' emotions and generate interactive content or responses accordingly. Furthermore, while many users often feel the need for appropriate responses or feedback based on their emotions while watching the news, no system existed that could meet these needs.

[1331] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for acquiring text news in real time, means for receiving a news article selected by the user and passing it to the generation AI, means for the generation AI to generate video and audio content based on the news article, means for streaming the generated video and audio content to the user, means for receiving chat comments from the user and generating an appropriate response, and means for recognizing the user's emotions and adjusting the content based on the emotion. This enables highly real-time news delivery that can be enjoyed visually and audibly, and makes it possible to provide appropriate responses and interactive experiences according to the user's emotions.

[1332] "Text news" refers to news articles in written form distributed in real time over the Internet or through other communication means.

[1333] "User" refers to an individual or group of people who utilize the system of the present invention to select, view, and interact with news stories.

[1334] "Generative AI" refers to AI that has the ability to generate video and audio content based on input news articles.

[1335] "Video and audio content" means media that provides information through visual and audio means, including generated news feeds.

[1336] "Streaming distribution" is a method of providing video and audio content to users in real time over the Internet.

[1337] "Chat comments" refer to text messages that users enter while viewing content.

[1338] An "appropriate response" is an automatically generated reply or feedback that is generated in response to a chat comment or emotion from a user.

[1339] "Emotion recognition" refers to the process of analyzing and understanding a user's emotional state from data such as facial expressions and voice.

[1340] "Adjusting content" refers to changing the content or format of video and audio content in real time based on recognized user sentiment.

[1341] The present invention combines a system that acquires news articles in real time, converts them into video and audio content, and distributes them to users with an emotion engine that recognizes user emotions. Specific embodiments for carrying out the present invention will be described below.

[1342] System Configuration

[1343] The present invention is mainly composed of a server, a terminal, and a user. Each component will be explained below.

[1344] server

[1345] The server has the following functions:

[1346] 1. Real-time text news acquisition:

[1347] The server uses the news provider's API to retrieve the latest articles in real time. It sends an API request and receives the response data in JSON format. It extracts the necessary information (title, content, URL, etc.) and stores it in an internal database.

[1348] 2. Select news articles and pass them to the generation AI:

[1349] The system receives information about a news article selected by the user on their device and passes it to a generative artificial intelligence (generative AI). The generative AI generates video and audio content based on the article content. This generation process includes creating a summary of the news article, generating video of the VTuber character, and synthesizing voice.

[1350] 3. Emotion recognition and content adjustment:

[1351] The emotion engine recognizes emotions from the user's facial expressions and voice in real time and feeds the recognized emotion data back to the generation AI, which then adjusts the content of the video and audio content based on this feedback information.

[1352] 4. Streaming:

[1353] The generated video and audio content is distributed to the user's terminal in streaming format.

[1354] 5. Handling chat comments:

[1355] The system receives chat comments from users and analyzes their emotions with an emotion engine. Based on the results, the generative AI generates an appropriate response and returns it to the user via the server.

[1356] Terminal

[1357] The terminal has the following functions:

[1358] 1. Select a news article:

[1359] It provides a user interface to display a list of news articles, and when the user selects an article of interest, it sends the selection information to the server.

[1360] 2. Viewing content:

[1361] It has the ability to play distributed video and audio content in real time.

[1362] 3. Chat function:

[1363] It provides an interface that allows users to submit comments while watching, which are then sent to the server, which returns an appropriate response.

[1364] User

[1365] The user performs the following actions:

[1366] 1. Select a news article:

[1367] Select the news article of interest through the device interface.

[1368] 2. Content viewing and emotional feedback:

[1369] The selected news article is then viewed as real-time generated video and audio content, with facial expressions and voice acting analyzed by the emotion engine as the news is viewed.

[1370] 3. Sending chat comments:

[1371] Send comments while watching and receive responses from the generative AI.

[1372] Hardware and software used

[1373] On the server side, we use a server equipped with a high-performance GPU. The main software used is an API for retrieving news from news sources, generative AI (e.g., OpenAI's GPT-4), and an emotion recognition engine (e.g., DeepFace, OpenFace).

[1374] On the terminal side, mobile devices such as smartphones and tablets are used, which are equipped with cameras and microphones to collect data for emotion recognition.

[1375] Specific examples

[1376] 1. View news article list:

[1377] The device displays a list of the latest news, and the user selects "an article about the latest AI technology."

[1378] 2. Content Creation and Viewing:

[1379] The server uses AI to convert the selected articles into video and audio content, which are then streamed. Users can then watch the generated VTuber videos on their devices.

[1380] 3. Emotion Recognition and Feedback:

[1381] While watching, the user's facial expressions are analyzed by a camera, and a "happy" expression is recognized. This information is fed back to the generative AI, which then adjusts the video and audio content to better match the user's emotions.

[1382] 4. Chat comment response:

[1383] When a user types "This technology is amazing!" into the chat, the emotion engine analyzes the emotion and the generative AI responds appropriately with "Yes, the evolution of AI is truly amazing!"

[1384] Prompt Sentence Examples

[1385] "News article: 'The latest AI technology has been announced. This technology is...'"

[1386] "Summary: Please summarize this article."

[1387] "Video and Audio: Generate a script in which a VTuber character explains this news article in an easy-to-understand way."

[1388] The above is a specific embodiment for carrying out the present invention.

[1389] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1390] Step 1:

[1391] The server uses the news provider's API to obtain the latest news articles in real time. It sends an API request and receives the response data in JSON format. The server extracts necessary information such as the news article title, content, and URL, and stores it in an internal database. The input is the API request, and the output is the analyzed news article data.

[1392] Step 2:

[1393] The user selects an article of interest from the displayed news article list using the terminal interface. The selection information is sent from the terminal to the server. The input is the user's selection action, and the output is the selected news article information.

[1394] Step 3:

[1395] The server passes data based on the selected news article to the generation AI. The generation AI then analyzes the article content, creates a summary, and uses that to generate video and audio content for the VTuber character. The input is the selected news article data, and the output is the generated video and audio content.

[1396] Step 4:

[1397] The emotion engine starts up and analyzes the user's facial expressions and voice in real time. Data is acquired using the device's camera and microphone, and the emotion engine analyzes that data to recognize the user's emotions. The input is the user's facial and voice data, and the output is the recognized emotion data.

[1398] Step 5:

[1399] The recognized emotional data is then fed back to the generation AI, which then adjusts the content based on the fed-back emotional data. Specifically, it changes the character's facial expression and tone according to the user's emotions. The input is the recognized emotional data, and the output is the adjusted content.

[1400] Step 6:

[1401] The server distributes the generated video and audio content to the user's device in streaming format. This is done in real time through a streaming server. The input is the adjusted content, and the output is the video and audio to be distributed.

[1402] Step 7:

[1403] While watching a video, users can send comments using a chat interface. The comments are sent from the device to the server. The input is the user's chat comments, and the output is the comment data received by the server.

[1404] Step 8:

[1405] The server passes the received chat comments to the emotion engine for analysis. The emotion engine analyzes emotions based on the user's comments and feeds the results back to the generation AI. The input is the chat comments, and the output is the analyzed emotion data.

[1406] Step 9:

[1407] The generative AI generates an appropriate response based on the analyzed emotional data and returns it to the user through the server. The response is generated using a prompt sentence. The input is the analyzed emotional data, and the output is the response message generated by the AI.

[1408] Example prompt

[1409] "News article: 'The latest AI technology has been announced. This technology is...'"

[1410] "Summary: Please summarize this article."

[1411] "Video and Audio: Generate a script in which a VTuber character explains this news article in an easy-to-understand way."

[1412] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1413] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1414] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1415] [Fourth embodiment]

[1416] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1417] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1418] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1419] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1420] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1421] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1422] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1423] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1424] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1425] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1426] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1427] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1428] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1429] The present invention is a system that acquires news articles in real time, converts them into video and audio content, and distributes them to users. This allows users to enjoy the news visually and aurally, and also allows them to interact with the news. The following describes an embodiment of the present invention.

[1430] System Configuration

[1431] 1. Server

[1432] The server has the ability to retrieve news articles from news sources in real time, obtain the latest article list using the news source's API, and store the data.

[1433] It also accepts news article selection requests from users and passes the selected article content to the generation AI.

[1434] The artificial intelligence has the ability to generate video and audio content for VTuber characters based on received news articles. The generated content is temporarily stored on the server.

[1435] The server has the function of distributing the generated video and audio content to users in streaming format.

[1436] In addition, it has the ability to receive chat comments sent by users, use generative AI to generate appropriate responses, and return them to the users.

[1437] 2. Terminal

[1438] The terminal provides a user interface and has the function of displaying a list of news articles. When the user selects an article of interest, the selection information is sent to the server.

[1439] It has the ability to play distributed video and audio content in real time.

[1440] It also provides an interface for users to send comments using a chat function.

[1441] 3. Users

[1442] The user selects news articles of interest through the terminal interface.

[1443] Selected news articles are viewed as real-time video and audio content.

[1444] You can send chat comments during the broadcast and receive responses from the server.

[1445] Program processing flow

[1446] Get news articles

[1447] The server calls the news provider's API to retrieve the latest news articles, parses them in JSON format, and saves the necessary information.

[1448] Accepting article conversion requests

[1449] The user selects news articles of interest from the terminal interface and the selection is transmitted to the server.

[1450] Content generation by generative AI

[1451] The server passes the received news article to the generation AI, which generates video and audio content for the VTuber character based on the article content. This includes creating a summary of the article, simulating the character's facial expressions and movements, and synthesizing voice.

[1452] Delivery of video and audio content

[1453] The server transmits the generated content to the user terminal via a streaming server, and the user watches the received content in real time.

[1454] Chat feature

[1455] Users can send chat comments in real time during the broadcast, and the server passes the received comments to a generation AI that generates an appropriate response, which is returned to the user via the chat interface.

[1456] Specific examples

[1457] 1. The user selects a news article.

[1458] The terminal displays a list of news articles and the user selects an article in the "Technology" category.

[1459] 2. The server passes the article to the generation AI.

[1460] The server passes the selected article to a generation AI to generate video and audio content.

[1461] 3. Distribution begins

[1462] The server distributes the generated content to the user through streaming, and the user views the content.

[1463] 4. Use of chat function

[1464] Users submit comments during a broadcast, and the server generates a response and sends it back to the user.

[1465] As described above, the present invention provides a new means for enjoying news visually and aurally, and allows users to react to the news interactively.

[1466] The processing flow will be explained below.

[1467] Step 1:

[1468] The server accesses the news provider's API to retrieve the latest news articles, sending an API request and receiving the response data in JSON format.

[1469] Step 2:

[1470] The server parses the JSON data of the retrieved news article, extracts the necessary information (title, content, URL, etc.), and stores it in an internal database or cache.

[1471] Step 3:

[1472] The device displays a list of news articles to the user on an interface, including the title of each article, a thumbnail image, and a brief summary.

[1473] Step 4:

[1474] The user selects an article of interest from the displayed news article list, and the ID (or URL) of the selected article is sent to the server via the terminal.

[1475] Step 5:

[1476] The server retrieves the corresponding article content from an internal database based on the article ID received from the user.

[1477] Step 6:

[1478] The server sends the retrieved article content to the generation AI module, which analyzes the article content, creates a summary, and then generates video and audio content for the VTuber character.

[1479] Step 7:

[1480] The generative AI simulates the facial expressions and movements of a VTuber character in real time based on a news article, and also generates audio reading the article content and integrates it into the video.

[1481] Step 8:

[1482] The server prepares the video and audio content received from the generation AI so that it can be provided in streaming format, and configures it for distribution through a streaming server.

[1483] Step 9:

[1484] The server then streams the prepared video and audio content in real time to the device, where the user can begin watching.

[1485] Step 10:

[1486] Users can submit comments through a chat interface while watching, and the comments are sent to the server in real time.

[1487] Step 11:

[1488] The server passes chat comments received from users to the generation AI and asks it to generate an appropriate response. The generation AI creates a response based on the content of the comment.

[1489] Step 12:

[1490] The response generated by the generation AI is sent to the server, which returns the response to the user through a chat interface.

[1491] Through these steps, news articles are transformed into video and audio content in real time and delivered interactively, allowing users to enjoy the news visually and audibly and react to it in real time.

[1492] Example 1

[1493] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1494] Users need to maximize their visual and auditory capabilities when accessing text data, but traditional methods do not adequately achieve this. Furthermore, users have limited options for interacting with the data in real time and engaging with it. Therefore, news and other text data must be delivered to users in a more interactive and engaging format.

[1495] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1496] In this invention, the server includes means for acquiring text data in real time, means for receiving data selected by the user and passing it to the generation AI, means for the generation AI to generate video and audio data based on the data, means for streaming the generated video and audio data to the user, and means for receiving comments from the user and generating appropriate responses, allowing the user to enjoy the text data of interest visually and audibly and to react interactively in real time.

[1497] "Text data" is digital data containing text information.

[1498] "Generative AI" is a system that uses artificial intelligence technology to generate moving images and audio data based on text data.

[1499] "Motion picture data" is data in a digital format that shows a sequence of images and provides visual information.

[1500] "Audio data" means data in a digital format that conveys information through hearing.

[1501] "Streaming distribution" is a technology that distributes digital media over the Internet in real time.

[1502] A "comment" is an opinion or question in text format that a user sends in real time in response to video or audio content.

[1503] MODE FOR CARRYING OUT THE INVENTION

[1504] This invention is a system that acquires text data in real time, generates video and audio data based on that data, and distributes them to users. This system allows users to enjoy the text data visually and audibly, and also allows them to respond interactively in real time.

[1505] Hardware and software used

[1506] 1. Server

[1507] It is used to call the news provider's API to obtain the latest news articles. This API can be from a general API service provider, for example.

[1508] Examples of generative AI include OpenAI's GPT-4 and similar AI models.

[1509] DeepFake technology is used to generate moving images, and Google Text-to-Speech and other voice synthesis engines are used.

[1510] Streaming services such as YouTube Live and Twitch are used for streaming.

[1511] 2. Terminal

[1512] It displays a list of news articles and provides a user interface for users to select articles they are interested in. This interface is implemented as a web application or smartphone app.

[1513] Processing flow

[1514] The server retrieves the latest news articles in real time using the news provider's API. The retrieved data is sent to the server in JSON format and analyzed. The analyzed article information is saved in the server's database. Specifically, for example, the server calls the endpoint of a service called "NewsAPI," sends a "GET" request, parses the retrieved JSON data, extracts the article title, text, date, etc., and saves them in the database.

[1515] The user browses through a list of news articles through the interface on the device and selects an article of interest. The ID information of the selected article is sent from the device to the server. For example, the user opens a news app on a smartphone or PC and selects an article in the "Technology" category. The device then sends the selection information to the server.

[1516] The server passes the selected article content to the generation AI. The generation AI generates video and audio content for the VTuber character based on the article content. This includes creating a summary of the article, simulating the character's facial expressions and movements, and processing voice synthesis. Specifically, it sends a prompt to the generation AI model, instructing it to summarize the article content. An example of a prompt is, "Please summarize the following news article: '(news article body)'."

[1517] The server sends the generated video and audio content to the user terminal via the streaming server, and the user watches the content in real time. Specifically, the server uploads the generated video file to the streaming server and sends a streaming URL to the user terminal. The user watches the content in real time via the URL.

[1518] Users can send chat comments in real time while watching. The server passes the received comments to the generation AI, which generates an appropriate response. The generated response is returned to the user via the chat interface. Specifically, the user sends a comment during the broadcast, and the server passes the comment to the generation AI, which generates a response. An example of a prompt is, "Please generate an appropriate response to the user's comment: 'Please tell us a specific application example of this technology.'"

[1519] In this way, the system provides a new means of visually and aurally enjoying text data, allowing users to respond interactively.

[1520] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1521] Step 1: Get news articles

[1522] The server calls the news provider's API to retrieve the latest news articles in real time. The retrieved data is sent to the server in JSON format and parsed. The retrieved JSON data includes information such as the article title, body text, and date. This data is then stored in a database.

[1523] Input: News source API endpoint

[1524] Output: Parsed news article data (title, body, date, etc.)

[1525] Specifically, for example, it uses the endpoint of a typical API service provider to send a "GET" request and parses the retrieved JSON data.

[1526] Step 2: Select an article

[1527] The user selects an article of interest from the list of news articles displayed through the interface on the terminal, and the ID information of the selected article is sent from the terminal to the server.

[1528] Input: The ID of the news article selected by the user

[1529] Output: ID information of selected articles sent to the server

[1530] Specifically, a user opens a news app on their smartphone or PC and selects an article in the "Technology" category, for example. At that time, the device sends a request including the ID of the selected article to the server.

[1531] Step 3: Submitting a content generation request

[1532] The server passes the selected article content to the generation AI, which then generates video and audio content for the VTuber character based on the article content, including creating a summary of the article, simulating the character's facial expressions and movements, and processing voice synthesis.

[1533] Input: Content of selected news article

[1534] Output: Prompt text passed to the generation AI

[1535] Specifically, the system sends a prompt to a generative AI model (e.g., GPT-4) to instruct it to summarize the article. An example prompt might be, "Please summarize the following news article: '(news article text)'."

[1536] Step 4: Generate video and audio content

[1537] The AI ​​generates video and audio content for the VTuber character based on the article content provided, including simulating the character's facial expressions and movements based on the summarized article content and generating audio using a speech synthesis engine.

[1538] Input: Summary of news article content

[1539] Output: The generated video and audio content

[1540] Specifically, it works by using a character simulation tool that uses DeepFake technology to generate voice using a voice synthesis engine (e.g., Google Text-to-Speech).

[1541] Step 5: Deliver your content

[1542] The server transmits the generated video and audio content to the user terminal via a streaming server, where the user can view the content in real time.

[1543] Input: Generated video and audio content

[1544] Output: Streaming URL sent in real time

[1545] Specifically, the generated video file is uploaded to a streaming server (such as YouTube Live or Twitch), and a streaming URL is sent to the user's device. The user can then watch the video in real time via that URL.

[1546] Step 6: Implementing the chat function

[1547] Users can send chat comments in real time while watching, and the server passes the received comments to a generation AI that generates an appropriate response, which is returned to the user via the chat interface.

[1548] Input: Chat comment sent by the user

[1549] Output: The generated response message

[1550] Specifically, while watching, the user sends a comment such as "Please tell me a specific application example of this technology," and the server passes the comment to the generation AI, which then generates an appropriate response. Example prompt: "Generate an appropriate response to the user's comment: 'Please tell me a specific application example of this technology.'"

[1551] (Application example 1)

[1552] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1553] Current news distribution methods are limited to providing text-based information, leaving many users lacking the means to enjoy news visually and aurally. Furthermore, they lack the ability to enjoy news articles interactively, making it difficult for users to efficiently understand news information. Furthermore, they lack the means to easily select specific news articles that interest them and share their reactions to them. These issues need to be addressed.

[1554] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1555] In this invention, the server includes means for acquiring text news in real time, means for receiving news articles selected by users and passing them to a generation AI, means for the generation AI to generate video and audio content based on the news articles, means for streaming the generated video and audio content to users, means for receiving chat comments from users and generating appropriate responses, and means for being installed in a smartphone application, which allows users to enjoy the news visually and audibly and to react to the news interactively.

[1556] "Means for obtaining text news in real time" refers to a system that obtains the latest news articles from news providers in real time via a dedicated API, and analyzes and stores them.

[1557] "Means for receiving news articles selected by the user and passing them on to the generating AI" refers to the communication and data processing functions for transmitting the news article information selected by the user to the generating AI.

[1558] "Generative AI" refers to AI technology for generating video and audio content based on received news articles. Specifically, it includes summarization, character facial expression and movement simulation, and voice synthesis.

[1559] "Means for generating video and audio content" refers to the mechanism by which the generative AI creates visual and audio multimedia content based on the content of news articles.

[1560] The "means for streaming the generated video and audio content to the user" refers to a streaming technology for delivering the generated multimedia content to the user terminal in real time.

[1561] "Means for receiving chat comments from users and generating appropriate responses" refers to a mechanism in which comments sent by users through the chat function are analyzed and the generation AI generates and replies to appropriate responses.

[1562] "Means installed in a smartphone application" refers to smartphone-specific software that incorporates the various functions mentioned above and provides users with an interactive news experience.

[1563] The "interface means for displaying a list of news articles and allowing the user to select an article of interest" is a user interface function that displays a list of news articles and allows the user to select a particular article.

[1564] This invention is a system that acquires news articles in real time, converts them into video and audio content, and distributes them to users. This system is composed of multiple components, such as a server, terminals, and generation AI, and its detailed configuration and processing are described below.

[1565] System Configuration

[1566] 1. Server

[1567] The server can obtain news articles from news providers in real time. This includes functions to obtain the latest article list using the news provider's API and store that data. It also accepts news article selection requests from users and passes the selected article content to the generation AI. The generation AI has the function to generate video and audio content based on the received news articles. The generated content is temporarily stored on the server and delivered to users in streaming format. It also receives chat comments sent by users, uses the generation AI to generate appropriate responses, and returns them to the user.

[1568] 2. Terminal

[1569] The device provides a user interface and can display a list of news articles. When a user selects an article of interest, the device transmits the selection information to the server. It can also play the distributed video and audio content in real time. It also provides an interface for users to send comments using a chat function. The device is installed as a smartphone application.

[1570] 3. Users

[1571] Users select news articles of interest through the device interface, and can view the selected news articles as video and audio content streamed in real time. Users can also send chat comments during the stream and receive responses from the server.

[1572] Program processing flow explanation

[1573] Get news articles

[1574] The server calls the news provider's API to retrieve the latest news articles. The retrieved articles are parsed in JSON format and the necessary information is saved. The server uses Node.js and Express, and an HTTP client such as Axios is used to communicate with the news provider.

[1575] Accepting article conversion requests

[1576] The device displays a list of news articles, and the user selects the news article they are interested in. The selection information is sent to the server. This process is implemented in a smartphone application using React Native.

[1577] Content generation by generative AI

[1578] The server passes the received news article to the generation AI, which generates video and audio content for the VTuber character based on the article content. This includes creating a summary of the article, simulating the character's facial expressions and movements, and synthesizing voice. The generation AI uses GPT-4 and deep learning frameworks such as TensorFlow and PyTorch.

[1579] Delivery of video and audio content

[1580] The server sends the generated content to the user's device via a streaming server (e.g., Wowza Streaming Engine), and the user watches the received content in real time.

[1581] Chat feature

[1582] Users can send chat comments in real time during the broadcast, and the server passes the received comments to the generation AI, which generates an appropriate response, which is returned to the user via the chat interface.

[1583] Examples of concrete examples and prompts

[1584] Specific examples

[1585] The user selects an article in the "Technology" category from a list of news articles. The server retrieves an article about the latest technology for self-driving cars from the news source and passes it to the generation AI. The generation AI summarizes the article and generates a video in which a VTuber character explains it. The user watches the generated video on a smartphone app and asks in chat, "Is the technology introduced here also used by other manufacturers?" The generation AI replies, "Yes, this technology is being adopted by many automakers."

[1586] Example prompts for generative AI models

[1587] Article content: Article about the latest technology in autonomous vehicles

[1588] Title: Evolution of in-vehicle AI technology in 2023

[1589] Article text: Advances in AI technology are enabling the latest self-driving cars to operate more safely and efficiently. In particular, real-time data processing combined with advanced sensor technology has improved the car's ability to accurately perceive its surroundings...

[1590] Prompt: Based on the article below, create a 5-minute news commentary video with a VTuber character.

[1591] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1592] Step 1:

[1593] The server retrieves the latest news articles in real time through the news provider's API.

[1594] How it works: The server periodically sends GET requests to the API endpoint to receive news article data in JSON format, which is then parsed to extract the necessary information (title, body, date, etc.) and store it in a database such as MongoDB.

[1595] Input: JSON data from the news source API.

[1596] Output: News article data stored in a database.

[1597] Step 2:

[1598] The terminal displays a list of news articles to the user.

[1599] Specific operation: The smartphone application on the device displays a list of news articles retrieved from the server on the user interface. The user selects the news article of interest. The selection information is sent from the device to the server.

[1600] Input: A list of news articles retrieved from the server.

[1601] Output: Information about the news article selected by the user.

[1602] Step 3:

[1603] The server passes the news articles selected by the user to the generating artificial intelligence.

[1604] Specific operation: The server calls an internal API to pass the selected news article to the generation AI module, which generates an appropriate prompt. The generation AI processes and analyzes the data based on the prompt and the news article content.

[1605] Input: Information about the news article selected by the user, and a prompt.

[1606] Output: The prompt and news article passed to the generation AI.

[1607] Step 4:

[1608] Generative AI generates video and audio content based on news articles.

[1609] How it works: The generative AI generates content based on prompts, summarizes articles, simulates the facial expressions and movements of VTuber characters, and performs voice synthesis. This results in the creation of visual and audio news commentary videos. The technologies used include GPT-4, TensorFlow, and PyTorch.

[1610] Input: Prompt text and news article content.

[1611] Output: The generated video and audio content.

[1612] Step 5:

[1613] The server streams the generated video and audio content to the user.

[1614] Specific operation: The generated video and audio content is temporarily stored on the server and then distributed in real time to the user's device via a streaming server (e.g., Wowza Streaming Engine).

[1615] Input: Generated video and audio content.

[1616] Output: The streaming content sent to the user device.

[1617] Step 6:

[1618] Users can view the distributed video and audio content and send chat comments.

[1619] Specific operation: A user watches a live video on a smartphone application and sends comments using the chat interface. The comments are sent to the server.

[1620] Input: The user's chat comment.

[1621] Output: Chat comments sent to the server.

[1622] Step 7:

[1623] The server passes the received chat comments to a generation artificial intelligence to generate an appropriate response.

[1624] Specific operation: The server analyzes the chat comments and passes them to the generation AI to generate an appropriate response, which is then sent back to the device and displayed to the user.

[1625] Input: The user's chat comment.

[1626] Output: The chat response generated by the generative AI.

[1627] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1628] The present invention is a system that acquires news articles in real time, converts them into video and audio content, and delivers them to users, and also combines it with an emotion engine that recognizes the user's emotions. This allows users to enjoy the news visually and aurally, and further enables them to receive responses that correspond to their emotions through an interactive experience. The following describes an embodiment of the present invention.

[1629] System Configuration

[1630] 1. Server

[1631] The server has the ability to retrieve news articles from news sources in real time, use the news source's API to obtain the latest article list, and analyze and store the data.

[1632] It also accepts news article selection requests from users and passes the selected article content to the generation artificial intelligence (generation AI).

[1633] The AI ​​has the ability to generate video and audio content for VTuber characters based on the received news articles. The generated content is temporarily stored on the server.

[1634] The server has the function of distributing the generated video and audio content to users in streaming format.

[1635] Furthermore, it has the ability to receive chat comments sent by users, recognize the user's emotions through an emotion engine, and then use the generative AI to generate an appropriate response based on that.

[1636] 2. Terminal

[1637] The terminal provides a user interface and has the function of displaying a list of news articles. When the user selects an article of interest, the selection information is sent to the server.

[1638] It has the ability to play distributed video and audio content in real time.

[1639] It also provides an interface for users to send comments using a chat function.

[1640] 3. Users

[1641] The user selects news articles of interest through the terminal interface.

[1642] Selected news articles are viewed as real-time video and audio content.

[1643] You can send chat comments during the broadcast and receive responses from the server.

[1644] 4. Emotion Engine

[1645] The emotion engine has the ability to recognize emotions from the user's facial expressions and voice, allowing it to grasp the user's emotions in real time.

[1646] The emotion engine has the ability to have the generative AI adjust the content of video and audio content based on the user's emotions.

[1647] Program processing flow

[1648] Get news articles

[1649] The server calls the news provider's API to retrieve the latest news articles. It sends an API request, receives the response data in JSON format, analyzes it, extracts the necessary information (title, content, URL, etc.), and stores it in an internal database.

[1650] Accepting article conversion requests

[1651] The user selects the news article of interest from the terminal interface, and the selection information is sent to the server via the terminal. The server receives the user's request and retrieves the corresponding article content from its internal database.

[1652] Content generation by generative AI

[1653] The server sends the retrieved article content to the generation AI, which analyzes the article content, creates a summary, and then generates video and audio content for the VTuber character. This process includes real-time simulation of the character's facial expressions and movements, as well as voice synthesis.

[1654] Emotion Engine Operation

[1655] The emotion engine analyzes the user's facial expressions and voice data in real time to recognize their emotions. The recognized emotion data is fed back to the generation AI, which then adjusts the content of the video and audio content based on this.

[1656] Delivery of video and audio content

[1657] The server prepares the video and audio content received from the generation AI so that it can be provided in streaming format. It configures the content for distribution through the streaming server and sends it to the device in real time. The user then begins watching the content on their device.

[1658] Chat feature

[1659] Users can send comments through a chat interface while watching. The comments are sent to the server in real time. The server then passes the chat comments received from the user to an emotion engine, which analyzes the user's emotions. Based on the results, the generative AI generates an appropriate response and returns it to the user via the server.

[1660] Specific examples

[1661] 1. The user selects a news article.

[1662] The terminal displays a list of news articles and the user selects an article in the "Technology" category.

[1663] 2. The server passes the article to the generation AI.

[1664] The server passes the selected article to a generation AI to generate video and audio content.

[1665] 3. Use of Emotion Engine

[1666] While a user is watching a video, the emotion engine recognizes emotions from the user's facial expressions and voice, and the generative AI adjusts the content in real time based on that.

[1667] 4. Distribution begins

[1668] The server distributes the generated content to the user through streaming, and the user views the content.

[1669] 5. Use of chat function

[1670] Users send comments during the broadcast, the server analyzes the user's emotions using an emotion engine, and the generative AI generates an appropriate response and returns it to the user.

[1671] The present invention provides a new means of enjoying the news visually and aurally, allowing users to not only react to the news interactively, but also providing appropriate responses and content according to the user's emotions.

[1672] The processing flow will be explained below.

[1673] The present invention is a system that combines a system that acquires news articles in real time and converts them into video and audio content that provides visual and auditory enjoyment to users with an emotion engine that recognizes user emotions. The processing flow of the present invention will be explained below by dividing it into specific steps.

[1674] Step 1:

[1675] The server accesses the news provider's API to retrieve the latest news articles, sends an API request, and receives the response data in JSON format.

[1676] Step 2:

[1677] The server analyzes the JSON data of the retrieved news article, extracts the necessary information (title, content, URL, etc.), and stores it in an internal database.

[1678] Step 3:

[1679] The device displays a list of news articles to the user on an interface, including the title of each article, a thumbnail image, and a brief summary.

[1680] Step 4:

[1681] The user selects an article of interest from the displayed news article list, and the ID (or URL) of the selected article is sent to the server via the terminal.

[1682] Step 5:

[1683] The server retrieves the corresponding article content from an internal database based on the article ID received from the user.

[1684] Step 6:

[1685] The server sends the retrieved article content to the generation AI module, which analyzes the article content, creates a summary, and generates video and audio content for the VTuber character.

[1686] Step 7:

[1687] The generative AI simulates the facial expressions and movements of a VTuber character in real time based on a news article, and also generates audio reading the article content and integrates it into the video.

[1688] Step 8:

[1689] The server prepares the video and audio content received from the generation AI so that it can be provided in streaming format. It sets up distribution through the streaming server.

[1690] Step 9:

[1691] The server streams the prepared video and audio content in real time to the terminal, and the user begins watching.

[1692] Step 10:

[1693] The emotion engine analyzes the user's facial expressions and voice in real time to recognize their emotions, and the recognized emotion data is fed back to the generation AI.

[1694] Step 11:

[1695] The generative AI adjusts the VTuber character's facial expressions and movements based on data from the emotion engine, dynamically changing video and audio content.

[1696] Step 12:

[1697] Users can submit comments through a chat interface while watching, and the comments are sent to the server in real time.

[1698] Step 13:

[1699] The server passes the chat comments received from the user to the emotion engine, which analyzes the user's emotions. Based on the results, the generative AI generates an appropriate response.

[1700] Step 14:

[1701] The responses generated by the AI ​​are returned to the user via the server, and are displayed in a chat interface, allowing the user to enjoy the news interactively.

[1702] Through these steps, news articles are converted into video and audio content in real time, and content tailored to the user's emotions is provided, allowing users to enjoy the news visually and audibly while also interacting with it.

[1703] Example 2

[1704] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1705] In recent years, there has been an increase in the number of ways to enjoy news visually and aurally, but users still rely on general text articles. This makes news viewing one-way and lacks an interactive experience. Furthermore, the inability to analyze emotions based on the news content or respond in real time means that a personalized viewing experience is not provided.

[1706] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1707] In this invention, the server includes means for acquiring text news in real time, means for receiving news articles selected by users and passing them to a generation artificial intelligence, means for the generation artificial intelligence to generate video and audio content based on the news articles, means for streaming the generated video and audio content to users, means for receiving chat comments from users and generating appropriate responses, and means for analyzing user emotions in real time using an emotion engine and adjusting content based on the analysis results. This allows users to enjoy the news visually and audibly and receive responses interactively. Furthermore, personalized content can be provided based on emotion analysis.

[1708] "Text news" refers to news articles that consist only of text information obtained from news sources.

[1709] A "server" is a computer system that provides data and services to multiple clients over a network.

[1710] "Users" are consumers who utilize the system to select news articles and view video and audio content.

[1711] "Generative artificial intelligence" refers to machine learning models and algorithms for generating video and audio content based on retrieved news articles.

[1712] "Video and audio content" means material for visual and auditory enjoyment generated by generative artificial intelligence.

[1713] "Streaming distribution" is a technology that transfers data sequentially and plays it back in real time.

[1714] "Chat comments" are text messages that users type and send in real time.

[1715] An "appropriate response" is an appropriate reply provided by a generative AI or system in response to a chat comment from a user.

[1716] An "emotion engine" is a system or algorithm that analyzes a user's facial expressions and voice data to recognize emotions.

[1717] An "interface" is a screen or operating means through which a user interacts with a system.

[1718] "Real-time" refers to the time nature of data and processing occurring immediately without delay.

[1719] The present invention is a system that acquires news articles in real time, converts them into video and audio content, and delivers them to users, and also combines it with an emotion engine that recognizes the user's emotions. This allows users to enjoy the news visually and aurally, and further enables them to receive responses that correspond to their emotions through an interactive experience. The following describes an embodiment of the present invention.

[1720] System configuration

[1721] 1. Server

[1722] The server has the ability to retrieve news articles from news sources in real time, use the news source's API to obtain the latest article list, and analyze and store the data.

[1723] It also accepts news article selection requests from users and passes the selected article content to the generation artificial intelligence (generation AI).

[1724] The AI ​​generator generates video and audio content for virtual characters based on received news articles. The generated content is temporarily stored on the server.

[1725] The server has the function of distributing the generated video and audio content to users in streaming format.

[1726] Furthermore, it has the ability to receive chat comments sent by users, recognize the user's emotions through an emotion engine, and then use the generative AI to generate an appropriate response based on that.

[1727] 2. Terminal

[1728] The terminal provides a user interface and has the function of displaying a list of news articles. When the user selects an article of interest, the selection information is sent to the server.

[1729] It has the ability to play distributed video and audio content in real time.

[1730] It also provides an interface for users to send comments using a chat function.

[1731] 3. Users

[1732] The user selects news articles of interest through the terminal interface.

[1733] Selected news articles are viewed as real-time video and audio content.

[1734] You can send chat comments during the broadcast and receive responses from the server.

[1735] 4. Emotion Engine

[1736] The emotion engine has the ability to recognize emotions from the user's facial expressions and voice, allowing it to grasp the user's emotions in real time.

[1737] The emotion engine has the ability to have the generative AI adjust the content of video and audio content based on the user's emotions.

[1738] Specific examples

[1739] For example, each component operates as follows:

[1740] 1. Select a news article

[1741] The user browses the list of news articles on the device interface and selects an article in the "Technology" category.

[1742] 2. Content generation using generative AI

[1743] The article selected by the user is sent to the server via the device, and the server passes the article to the AI ​​generator, which then generates video and audio content read by a virtual character.

[1744] 3. Emotion analysis using an emotion engine

[1745] While the user is watching the video, the emotion engine recognizes emotions from the user's facial expressions and voice, and this information is fed back to the generative AI, which then adjusts the content in real time.

[1746] 4. Streaming

[1747] The server distributes the generated content to the user in streaming format, and the user watches it in real time on the terminal.

[1748] 5. Chat Comments and Response Generation

[1749] Users can submit comments while watching, which are sent to the server, which uses an emotion engine to analyze the user's emotions, and the generative AI generates an appropriate response, which is then sent back to the user via the server.

[1750] Prompt Sentence Examples

[1751] "Create a video and audio piece that briefly summarizes the contents of a news article and has a virtual character read it aloud."

[1752] In this way, users can enjoy the news visually and audibly, while simultaneously receiving interactive responses. Personalized content based on sentiment analysis is also provided.

[1753] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1754] Step 1: Get news articles

[1755] The server calls the news provider's API to retrieve the latest news articles. As input, it sends a request with the API endpoint and required parameters. As output, it receives the news data in JSON format, parses it, and extracts the required information (title, content, URL, etc.). It then stores this data in an internal database.

[1756] Specific behavior:

[1757] The server sends a request to the News API using "https: / / newsapi.org / v2 / top-headlines?country=jp&apiKey=Your API Key".

[1758] The returned JSON data is analyzed, and the article title, content, URL, etc. are extracted and saved in the database.

[1759] Step 2: View and select a news article

[1760] The terminal displays a list of news articles. As input, it receives news article data retrieved from the server. As output, it displays the news article list in a user interface. The user selects news articles of interest and sends the selection information to the server.

[1761] Specific behavior:

[1762] The terminal receives the news article list sent from the server and displays it on the interface.

[1763] The user selects an article in the "technology" category, and the terminal transmits this selection information to the server.

[1764] Step 3: Send to article generation AI

[1765] The server receives the user's selection information. As input, it receives the ID and link of the news article selected by the user. As output, it retrieves the corresponding article content from the internal database and sends it to the generation AI.

[1766] Specific behavior:

[1767] The server retrieves relevant news articles from the database based on the selection information sent by the user.

[1768] The retrieved news articles are sent in text format to the generation AI.

[1769] Step 4: Content generation with generative AI

[1770] The generative AI generates video and audio content based on the submitted article. It receives the text data of the news article as input. It generates video and audio content of a virtual character as output. Specifically, it analyzes the article content, generates a summary, simulates the character's facial expressions and movements, and synthesizes voice.

[1771] Specific behavior:

[1772] The generative AI analyzes news articles, creates summaries, and generates videos with the content, "A new technology has been released. This technology is..."

[1773] The character's movements and facial expressions are specified, voice synthesis is performed, and the generated content is sent to the server.

[1774] Step 5: Prepare your content for distribution

[1775] The server receives the video and audio content generated by the generation AI. As input, it receives the video and audio data from the generation AI. As output, it converts these data into a streaming format and prepares it for distribution to the user's device through the streaming server.

[1776] Specific behavior:

[1777] The server receives the video and audio files from the generated AI and transfers them to the streaming server.

[1778] Configure streaming settings and generate a streaming link.

[1779] Step 6: Stream your content

[1780] The server streams the generated content to the user's device. As input, the server sends the generated content data to the streaming server and provides the streaming link to the user. As output, the server provides the content to be played in real time on the user's device.

[1781] Specific behavior:

[1782] The server provides a streaming link of the generated video content to the user's terminal.

[1783] Users can watch video and audio content in real time on their devices.

[1784] Step 7: Sending chat comments and generating replies

[1785] Users can send chat comments while watching a video. They input their comments using the device interface and send them to the server.

[1786] The server receives the sent chat comments and analyzes them using the emotion engine. As input, it receives chat comments from users. As output, it feeds back the analysis results to the generation AI, which generates an appropriate response.

[1787] Specific behavior:

[1788] The user sends a comment from the device saying, "This news is amazing!"

[1789] The server receives the comments and analyzes the emotion of "surprise" using an emotion engine.

[1790] The generative AI generates a response such as "Yes, modern technology is truly amazing," and the server sends that response to the user.

[1791] In this way, users can enjoy the news visually and audibly, while simultaneously receiving interactive responses. Personalized content based on sentiment analysis is also provided.

[1792] (Application example 2)

[1793] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1794] Conventional news delivery systems were capable of acquiring and providing news text information in real time, but were limited in the means to provide it as visually and aurally enjoyable video and audio content. Furthermore, they were not able to recognize users' emotions and generate interactive content or responses accordingly. Furthermore, while many users often feel the need for appropriate responses or feedback based on their emotions while watching the news, no system existed that could meet these needs.

[1795] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for acquiring text news in real time, means for receiving a news article selected by the user and passing it to the generation AI, means for the generation AI to generate video and audio content based on the news article, means for streaming the generated video and audio content to the user, means for receiving chat comments from the user and generating an appropriate response, and means for recognizing the user's emotions and adjusting the content based on the emotion. This enables highly real-time news delivery that can be enjoyed visually and audibly, and makes it possible to provide appropriate responses and interactive experiences according to the user's emotions.

[1796] "Text news" refers to news articles in written form distributed in real time over the Internet or through other communication means.

[1797] "User" refers to an individual or group of people who utilize the system of the present invention to select, view, and interact with news stories.

[1798] "Generative AI" refers to AI that has the ability to generate video and audio content based on input news articles.

[1799] "Video and audio content" means media that provides information through visual and audio means, including generated news feeds.

[1800] "Streaming distribution" is a method of providing video and audio content to users in real time over the Internet.

[1801] "Chat comments" refer to text messages that users enter while viewing content.

[1802] An "appropriate response" is an automatically generated reply or feedback that is generated in response to a chat comment or emotion from a user.

[1803] "Emotion recognition" refers to the process of analyzing and understanding a user's emotional state from data such as facial expressions and voice.

[1804] "Adjusting content" refers to changing the content or format of video and audio content in real time based on recognized user sentiment.

[1805] The present invention combines a system that acquires news articles in real time, converts them into video and audio content, and distributes them to users with an emotion engine that recognizes user emotions. Specific embodiments for carrying out the present invention will be described below.

[1806] System Configuration

[1807] The present invention is mainly composed of a server, a terminal, and a user. Each component will be explained below.

[1808] server

[1809] The server has the following functions:

[1810] 1. Real-time text news acquisition:

[1811] The server uses the news provider's API to retrieve the latest articles in real time. It sends an API request and receives the response data in JSON format. It extracts the necessary information (title, content, URL, etc.) and stores it in an internal database.

[1812] 2. Select news articles and pass them to the generation AI:

[1813] The system receives information about a news article selected by the user on their device and passes it to a generative artificial intelligence (generative AI). The generative AI generates video and audio content based on the article content. This generation process includes creating a summary of the news article, generating video of the VTuber character, and synthesizing voice.

[1814] 3. Emotion recognition and content adjustment:

[1815] The emotion engine recognizes emotions from the user's facial expressions and voice in real time and feeds the recognized emotion data back to the generation AI, which then adjusts the content of the video and audio content based on this feedback information.

[1816] 4. Streaming:

[1817] The generated video and audio content is distributed to the user's terminal in streaming format.

[1818] 5. Handling chat comments:

[1819] The system receives chat comments from users and analyzes their emotions with an emotion engine. Based on the results, the generative AI generates an appropriate response and returns it to the user via the server.

[1820] Terminal

[1821] The terminal has the following functions:

[1822] 1. Select a news article:

[1823] It provides a user interface to display a list of news articles, and when the user selects an article of interest, it sends the selection information to the server.

[1824] 2. Viewing content:

[1825] It has the ability to play distributed video and audio content in real time.

[1826] 3. Chat function:

[1827] It provides an interface that allows users to submit comments while watching, which are then sent to the server, which returns an appropriate response.

[1828] User

[1829] The user performs the following actions:

[1830] 1. Select a news article:

[1831] Select the news article of interest through the device interface.

[1832] 2. Content viewing and emotional feedback:

[1833] The selected news article is then viewed as real-time generated video and audio content, with facial expressions and voice acting analyzed by the emotion engine as the news is viewed.

[1834] 3. Sending chat comments:

[1835] Send comments while watching and receive responses from the generative AI.

[1836] Hardware and software used

[1837] On the server side, we use a server equipped with a high-performance GPU. The main software used is an API for retrieving news from news sources, generative AI (e.g., OpenAI's GPT-4), and an emotion recognition engine (e.g., DeepFace, OpenFace).

[1838] On the terminal side, mobile devices such as smartphones and tablets are used, which are equipped with cameras and microphones to collect data for emotion recognition.

[1839] Specific examples

[1840] 1. View news article list:

[1841] The device displays a list of the latest news, and the user selects "an article about the latest AI technology."

[1842] 2. Content Creation and Viewing:

[1843] The server uses AI to convert the selected articles into video and audio content, which are then streamed. Users can then watch the generated VTuber videos on their devices.

[1844] 3. Emotion Recognition and Feedback:

[1845] While watching, the user's facial expressions are analyzed by a camera, and a "happy" expression is recognized. This information is fed back to the generative AI, which then adjusts the video and audio content to better match the user's emotions.

[1846] 4. Chat comment response:

[1847] When a user types "This technology is amazing!" into the chat, the emotion engine analyzes the emotion and the generative AI responds appropriately with "Yes, the evolution of AI is truly amazing!"

[1848] Prompt Sentence Examples

[1849] "News article: 'The latest AI technology has been announced. This technology is...'"

[1850] "Summary: Please summarize this article."

[1851] "Video and Audio: Generate a script in which a VTuber character explains this news article in an easy-to-understand way."

[1852] The above is a specific embodiment for carrying out the present invention.

[1853] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1854] Step 1:

[1855] The server uses the news provider's API to obtain the latest news articles in real time. It sends an API request and receives the response data in JSON format. The server extracts necessary information such as the news article title, content, and URL, and stores it in an internal database. The input is the API request, and the output is the analyzed news article data.

[1856] Step 2:

[1857] The user selects an article of interest from the displayed news article list using the terminal interface. The selection information is sent from the terminal to the server. The input is the user's selection action, and the output is the selected news article information.

[1858] Step 3:

[1859] The server passes data based on the selected news article to the generation AI. The generation AI then analyzes the article content, creates a summary, and uses that to generate video and audio content for the VTuber character. The input is the selected news article data, and the output is the generated video and audio content.

[1860] Step 4:

[1861] The emotion engine starts up and analyzes the user's facial expressions and voice in real time. Data is acquired using the device's camera and microphone, and the emotion engine analyzes that data to recognize the user's emotions. The input is the user's facial and voice data, and the output is the recognized emotion data.

[1862] Step 5:

[1863] The recognized emotional data is then fed back to the generation AI, which then adjusts the content based on the fed-back emotional data. Specifically, it changes the character's facial expression and tone according to the user's emotions. The input is the recognized emotional data, and the output is the adjusted content.

[1864] Step 6:

[1865] The server distributes the generated video and audio content to the user's device in streaming format. This is done in real time through a streaming server. The input is the adjusted content, and the output is the video and audio to be distributed.

[1866] Step 7:

[1867] While watching a video, users can send comments using a chat interface. The comments are sent from the device to the server. The input is the user's chat comments, and the output is the comment data received by the server.

[1868] Step 8:

[1869] The server passes the received chat comments to the emotion engine for analysis. The emotion engine analyzes emotions based on the user's comments and feeds the results back to the generation AI. The input is the chat comments, and the output is the analyzed emotion data.

[1870] Step 9:

[1871] The generative AI generates an appropriate response based on the analyzed emotional data and returns it to the user through the server. The response is generated using a prompt sentence. The input is the analyzed emotional data, and the output is the response message generated by the AI.

[1872] Example prompt

[1873] "News article: 'The latest AI technology has been announced. This technology is...'"

[1874] "Summary: Please summarize this article."

[1875] "Video and Audio: Generate a script in which a VTuber character explains this news article in an easy-to-understand way."

[1876] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1877] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1878] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1879] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1880] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1881] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1882] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1883] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1884] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1885] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1886] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1887] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1888] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1889] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1890] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1891] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1892] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1893] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1894] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1895] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1896] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1897] The following is further disclosed regarding the above embodiment.

[1898] (Claim 1)

[1899] A means of obtaining text news in real time;

[1900] a means for receiving a news article selected by the user and passing it to the generating artificial intelligence;

[1901] A means for generative artificial intelligence to generate video and audio content based on news articles;

[1902] means for streaming the generated video and audio content to a user;

[1903] means for receiving chat comments from users and generating appropriate responses;

[1904] A system including:

[1905] (Claim 2)

[1906] 2. The system of claim 1, wherein the generating artificial intelligence further comprises means for changing the character's facial expressions and movements in real time according to the content of the news article.

[1907] (Claim 3)

[1908] 10. The system of claim 1, wherein the user terminal further comprises interface means for displaying a list of news articles and for allowing the user to select an article of interest.

[1909] "Example 1"

[1910] (Claim 1)

[1911] a means for acquiring text data in real time;

[1912] A means to receive the data selected by the user and pass it to the generating AI;

[1913] A means for the generation AI to generate video and audio data based on the data;

[1914] means for streaming the generated video and audio data to a user;

[1915] means for receiving comments from users and generating appropriate responses;

[1916] A system including:

[1917] (Claim 2)

[1918] The system of claim 1, wherein the generating AI further includes means for changing the character's expressions and actions in real time according to the content of the data.

[1919] (Claim 3)

[1920] 10. The system of claim 1, wherein the user terminal further comprises interface means for displaying a list of data and for allowing the user to select items of interest.

[1921] "Application Example 1"

[1922] (Claim 1)

[1923] A means of obtaining text news in real time;

[1924] a means for receiving a news article selected by the user and passing it to the generating artificial intelligence;

[1925] A means for generative artificial intelligence to generate video and audio content based on news articles;

[1926] means for streaming the generated video and audio content to a user;

[1927] means for receiving chat comments from users and generating appropriate responses;

[1928] a means for being installed in a smartphone application;

[1929] A system including:

[1930] (Claim 2)

[1931] 2. The system of claim 1, wherein the generating artificial intelligence further comprises means for changing the character's facial expressions and movements in real time according to the content of the news article.

[1932] (Claim 3)

[1933] 10. The system of claim 1, wherein the user terminal further comprises interface means for displaying a list of news articles and for allowing the user to select an article of interest.

[1934] "Example 2: Combining Emotion Engines"

[1935] (Claim 1)

[1936] A means of obtaining text news in real time;

[1937] a means for receiving a news article selected by the user and passing it to the generating artificial intelligence;

[1938] A means for generative artificial intelligence to generate video and audio content based on news articles;

[1939] means for streaming the generated video and audio content to a user;

[1940] means for receiving chat comments from users and generating appropriate responses;

[1941] A means for analyzing user emotions in real time using an emotion engine and adjusting content based on the analysis results;

[1942] A system including:

[1943] (Claim 2)

[1944] 2. The system of claim 1, wherein the generating artificial intelligence further comprises means for changing the character's facial expressions and movements in real time according to the content of the news article.

[1945] (Claim 3)

[1946] 10. The system of claim 1, wherein the user terminal further comprises interface means for displaying a list of news articles and for allowing the user to select an article of interest.

[1947] "Application example 2 when combining emotion engines"

[1948] New Claims

[1949] (Claim 1)

[1950] A means of obtaining text news in real time;

[1951] a means for receiving a news article selected by the user and passing it to the generating artificial intelligence;

[1952] A means for generative artificial intelligence to generate video and audio content based on news articles;

[1953] means for streaming the generated video and audio content to a user;

[1954] means for receiving chat comments from users and generating appropriate responses;

[1955] means for recognizing a user's emotion and adjusting content based on the emotion;

[1956] A system including:

[1957] (Claim 2)

[1958] 2. The system of claim 1, wherein the generating artificial intelligence further comprises means for changing the character's facial expressions and movements in real time according to the content of the news article.

[1959] (Claim 3)

[1960] 10. The system of claim 1, wherein the user terminal further comprises interface means for displaying a list of news articles and for allowing the user to select an article of interest. [Explanation of symbols]

[1961] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. A means of obtaining text news in real time; a means for receiving a news article selected by the user and passing it to the generating artificial intelligence; A means for generative artificial intelligence to generate video and audio content based on news articles; means for streaming the generated video and audio content to a user; means for receiving chat comments from users and generating appropriate responses; A system including:

2. 2. The system of claim 1, wherein the generating artificial intelligence further comprises means for changing the character's facial expressions and movements in real time according to the content of the news article.

3. 10. The system of claim 1, wherein the user terminal further comprises interface means for displaying a list of news articles and for allowing the user to select an article of interest.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A