System

The system addresses bias and comprehension issues in news commentary by generating videos with multiple perspectives, providing comprehensive and unbiased news understanding.

JP2026023958APending Publication Date: 2026-02-13SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024126279
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-08-01
Publication Date
2026-02-13

AI Technical Summary

Technical Problem

Current news aggregation and commentary services struggle with bias, difficulty in understanding text-based commentary, and lack of multifaceted perspectives, making it hard for users to grasp news comprehensively.

Method used

A system that acquires news information, summarizes it, generates explanatory text from multiple commentary characters' perspectives, and creates videos to present news commentary from various viewpoints, using natural language processing and generative AI models.

Benefits of technology

Enables users to understand news from multiple perspectives easily, reducing bias and ensuring timely delivery of unbiased, multifaceted commentary.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026023958000001_ABST
    Figure 2026023958000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: This system includes a means for acquiring news information, a means for summarizing the acquired news information, a means for generating a text for explaining the news information summarized by using a plurality of explanatory characters from various viewpoints, a means for generating a moving image by using the explanatory text of the character, and a means for distributing the generated moving image.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] Current news aggregation and commentary services have the problem of making it difficult to judge the bias and importance of news. Another issue is that users are easily influenced by commentary that is biased toward a particular opinion. Furthermore, text-based commentary can be difficult to understand, and many users are looking for unbiased, multifaceted news commentary. The purpose of this invention is to solve these problems. [Means for solving the problem]

[0005] The present invention provides a system including means for acquiring news information, means for summarizing the acquired news information, means for generating text that uses multiple commentary characters to explain the summarized news information from various perspectives, means for generating video using the character commentary text, and means for distributing the generated video. This allows users to view news commentary from various perspectives in video format, thereby deepening their understanding of the news and enabling them to obtain information without being biased toward any particular opinion.

[0006] "Means for acquiring news information" means a program or system for automatically collecting news articles from news sources on the Internet.

[0007] "Means for summarizing acquired news information" refers to natural language processing techniques and algorithms that extract important parts from acquired news information and summarize them in a concise format.

[0008] "A means for generating text that explains summarized news information from various perspectives using multiple commentary characters" refers to a program or system that generates news information as text from the perspectives of multiple characters with different backgrounds and positions.

[0009] "Means for generating video using character explanatory text" refers to a program or system for creating video content by combining character voice, animation, and graphics based on the generated explanatory text.

[0010] "Means for distributing the generated video" refers to a server or interface for uploading the created explanatory video to a platform accessible by users and distributing it. [Brief explanation of the drawings]

[0011] [Figure 1]1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0012] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0013] First, the terms used in the following description will be explained.

[0014] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0015] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0016] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0017] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0018] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0019] [First embodiment]

[0020] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0021] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0022] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0023] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0024] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0025] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0026] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0027] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0028] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0029] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0030] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0031] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0032] MODE FOR CARRYING OUT THE INVENTION

[0033] The system for implementing the present invention includes a series of processing means for acquiring news information, summarizing it, generating videos in which multiple commentary characters provide commentary from various perspectives, and finally distributing the videos. Below, the processing steps of the server, terminal, and user are explained with concrete examples.

[0034] 1. Acquiring news information

[0035] The server retrieves news information from a news source. This is done, for example, by using the news site's API. The server sends a request to the news source at a specific time (for example, every morning at 8:00) to retrieve the latest news information. The retrieved news is then stored in a database.

[0036] Examples:

[0037] The server accesses a news API on the Internet and receives the latest news headlines and article content.

[0038] 2. News Summary

[0039] The server summarizes the news information it has acquired. It uses natural language processing technology to extract key points from news articles and summarize them in a concise format. The results of this summary are also stored in the database.

[0040] Examples:

[0041] The server analyzes the title, date, time, and text of the news article, extracts key keywords and key points, and generates a summary of about 100 characters.

[0042] 3. Generating explanatory text

[0043] The server loads multiple commentary characters and generates commentary text for each commentary character, allowing for commentary from a variety of perspectives, including government, consumer, and expert perspectives.

[0044] Examples:

[0045] The server generates text explaining the news from the perspective of character A, and then generates explanatory text from the perspective of character B.

[0046] 4. Generating explanatory videos

[0047] The server generates a video based on the generated explanatory text, combining voice-overs and animations of each character. Related graphics and images are also added to the video, making it visually easy to understand.

[0048] Examples:

[0049] The server synthesizes the text of commentary character A into voice and combines the voice with the character's animation to generate a video file.

[0050] 5. Preparing and Streaming Videos

[0051] The server uploads the generated video to a dedicated video site, where advertisements are inserted and the video is ready to be distributed for free under an advertising model. Users can access the video site and watch the video.

[0052] Examples:

[0053] The server uploads the explanatory video to the video site and configures it to insert an advertisement in the first five seconds. Users access the site and select and watch the news explanatory video that interests them.

[0054] This system allows users to receive news from multiple perspectives in an easy-to-understand format. The server periodically retrieves news and quickly provides the latest information, ensuring that fresh news commentary videos are always delivered.

[0055] The processing flow will be explained below.

[0056] Specific processing steps of the program

[0057] Step 1: Get news information

[0058] 1.1. The server checks the news source's API endpoint.

[0059] 1.2. The server sends an API request on a specified schedule (e.g., every morning at 8:00).

[0060] 1.3. The server receives the news data and extracts information such as the title, text, date and time, and URL.

[0061] 1.4. The server stores the extracted news data in a database.

[0062] Specific behavior:

[0063] The server accesses the "News API" and requests the latest news. It analyzes the received news data, extracts the necessary information (title, text, date and time, URL), and records it in the database.

[0064] Step 2: Summarize the news

[0065] 2.1. The server loads the stored news article.

[0066] 2.2. The server parses the article using a natural language processing library.

[0067] 2.3. The server extracts key points and keywords and generates a news summary.

[0068] 2.4. The server associates the generated summaries with the original news data and stores them in a database.

[0069] Specific behavior:

[0070] The server applies natural language processing (NLP) techniques to summarize news articles, extracting key points from the articles, and then creates a concise summary based on the extracted information and stores it in a database.

[0071] Step 3: Generate explanatory text

[0072] 3.1. The server loads data for multiple commentary characters (e.g., character profiles and attributes).

[0073] 3.2. The server analyzes the news from each character's perspective.

[0074] 3.3. The server generates explanatory text from each character's perspective using a natural language generation model.

[0075] 3.4. The server stores the generated explanatory text in a database.

[0076] Specific behavior:

[0077] The server reads the attributes of the commentary characters (e.g., government officials, consumers, experts, etc.) and analyzes the news article from each character's perspective. Based on the analysis results, it uses natural language generation technology to create commentary text and records it in a database.

[0078] Step 4: Generate an explainer video

[0079] 4.1. The server retrieves the generated description text and character information.

[0080] 4.2. The server launches a video generation engine (e.g., OpenCV, FFmpeg).

[0081] 4.3. The server generates the character's voice data and combines it with the animation.

[0082] 4.4. The server builds the explainer video and adds any necessary graphics and effects.

[0083] 4.5. The server saves the completed video file in storage.

[0084] Specific behavior:

[0085] The server converts the explanatory text into audio data using speech synthesis technology and synchronizes it with the corresponding character animation. It also incorporates images and graphs related to the explanatory content into the video, generating the final explanatory video and saving it in storage.

[0086] Step 5: Prepare and stream your video

[0087] 5.1. The server accesses the management interface of the dedicated video site.

[0088] 5.2. The server uploads the generated video to the site.

[0089] 5.3. The server configures the insertion of ads into the ad slots.

[0090] 5.4. The server sets the video to public and makes it available for viewing.

[0091] Specific behavior:

[0092] The server logs in to the video site and uploads the saved instructional video. It then sets up the video so that ads are displayed when the video is played, allowing for free distribution under an advertising model. It also sets up the public settings so that new videos can be viewed when users access the site.

[0093] Step 6: User Watches Video

[0094] 6.1. The user accesses a dedicated video site.

[0095] 6.2. The user selects the news commentary video that interests them.

[0096] 6.3. The user plays the video and watches the news commentary.

[0097] Specific behavior:

[0098] A user visits a video site and finds a video they are interested in by category or from the latest news list. When they select a video and start playing it, an advertisement is displayed first, and then they can watch a news commentary video.

[0099] Example 1

[0100] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0101] Conventional news distribution systems are limited to providing information from a single perspective, making it difficult for users to understand the news from a variety of perspectives. Furthermore, there is a lack of automated means for quickly and effectively delivering the latest news information, and periodic updates are often performed manually. This can result in delayed or biased information being provided to users.

[0102] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0103] In this invention, the server includes means for acquiring news information, means for summarizing the news information, means for generating text that uses multiple commentary characters to explain the summarized news information from various perspectives, means for generating videos using the character commentary texts, and means for distributing the generated videos, thereby enabling users to understand the news from multiple perspectives and always receive the latest news commentary videos quickly.

[0104] "Means for obtaining news information" refers to a function for automatically obtaining the latest news information from news sources or APIs on the Internet.

[0105] The "means for summarizing news information" is a function that uses natural language processing technology to extract important points from retrieved news articles and summarize them in a concise form.

[0106] "A means for generating text that explains summarized news information from various perspectives using multiple commentary characters" is a function that uses a generative AI model to generate explanatory text to explain news information from various perspectives, such as the government perspective, consumer perspective, and expert perspective.

[0107] "Means for generating videos using character explanatory text" is a function that combines voice synthesis technology and animation based on the generated explanatory text to create videos that are easy to understand visually and aurally.

[0108] The "means for distributing the generated video" is a function for uploading the generated video to a dedicated video site and distributing the video so that users can view it.

[0109] MODE FOR CARRYING OUT THE INVENTION

[0110] The system for implementing this invention includes a series of processing means for acquiring news information, summarizing it, generating videos in which multiple commentary characters provide commentary from various perspectives, and finally distributing the videos. This system functions through the following processing steps by the server, terminal, and user.

[0111] The server sends a request to the API of a news site on the Internet and automatically obtains the latest news information. For example, an HTTP GET request is made to the news site's API endpoint, and the obtained news data is stored in a database. This ensures that the latest news information is always recorded.

[0112] The server then uses natural language processing (NLP) techniques to summarize the retrieved news information. For example, NLP models such as BERT are used to extract key points from news articles and summarize them in a concise format. The summary results are also stored in a database for further processing.

[0113] Next, the server loads multiple commentary characters and generates news commentary text from each character's perspective. Generative AI models used here include (for example) GPT-2. Users can receive news commentary from different commentary characters' perspectives, such as the government perspective, consumer perspective, and expert perspective.

[0114] The server generates videos based on the generated explanatory text, combining each character's voice reading with animation. The Google Text-to-Speech API is used for voice synthesis, and FFmpeg is used as video editing software to generate the video. Related graphics and images are also added to the generated video, making it visually easy to understand.

[0115] As a final step, the server uploads the generated video to a dedicated video site. For example, the video can be uploaded using the YouTube API and configured to insert advertisements. Users can then access this video site and select and watch news commentary videos of their interest.

[0116] Example prompt sentence:

[0117] Below is a major news article. Use this article as a starting point to generate commentary from the government's perspective, the consumer's perspective, and the expert's perspective.

[0118] News article title: {News title}

[0119] News article content: {News content}

[0120] This system allows users to receive news from multiple perspectives in an easy-to-understand format. The server periodically retrieves news and quickly provides the latest information, ensuring that fresh news commentary videos are always delivered.

[0121] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0122] Step 1:

[0123] The server sends a request to the API of a news site on the Internet to obtain the latest news information. The input is the API endpoint URL, and the output is news data in JSON format. Specifically, the server sends an HTTP GET request according to a regular schedule (for example, every morning at 8:00), and the obtained news data is stored in a database.

[0124] Input: News API endpoint URL

[0125] Output: News data (JSON format)

[0126] Specific operation: The server sends an HTTP GET request and receives news data.

[0127] Step 2:

[0128] The server summarizes the news information it obtains. The input is news data, and the output is summarized news text. Natural language processing (NLP) technology is used to extract key points from news articles and summarize them in a concise form. The generated summary is also stored in a database.

[0129] Input: News data

[0130] Output: Summary text

[0131] What it does: The server uses an NLP model to analyze a news article and generate a summary.

[0132] Step 3:

[0133] The server loads multiple commentary characters and generates commentary text from each character's perspective based on the summarized news information. The input is the summary text, and the output is the commentary text. A generative AI model is used to create commentary from various perspectives.

[0134] Input: Summary text

[0135] Output:Descriptive text

[0136] Specific operation: The server inputs a prompt sentence into the generative AI model and generates explanatory text from various perspectives.

[0137] Example prompt sentence:

[0138] Below is a major news article. Use this article as a starting point to generate commentary from the government's perspective, the consumer's perspective, and the expert's perspective.

[0139] News article title: {News title}

[0140] News article content: {News content}

[0141] Step 4:

[0142] The server generates a video that combines voice reading and animation based on the generated explanatory text. The input is the explanatory text, and the output is a completed video file. It integrates speech synthesis technology (e.g., Google Text-to-Speech API) and animation to create a visual video.

[0143] Input:Descriptive text

[0144] Output: Video file

[0145] Specific operation: The server synthesizes explanatory text into speech and combines it with character animation to generate a video.

[0146] Step 5:

[0147] The server uploads the generated video to a dedicated video site and prepares it for distribution. The input is a video file, and the output is the published video content. The video is uploaded using the API of the video distribution platform (for example, YouTube API).

[0148] Input: Video file

[0149] Output: Published video content

[0150] Specific operation: The server uploads the video file to the video site and performs settings such as inserting advertisements.

[0151] (Application example 1)

[0152] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0153] In today's world, news information is extremely diverse, making it difficult for users to efficiently understand it. Furthermore, while commentary on news topics requires commentary from multiple perspectives, there is a lack of easy ways to view such information. Furthermore, systems that provide commentary on news topics from multiple perspectives, rather than a single one, would be effective in helping users gain a deeper understanding of specific news topics, but such systems are currently limited.

[0154] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0155] In this invention, the server includes means for acquiring news information, means for summarizing the acquired news information, means for generating text that uses multiple commentary characters to explain the summarized news information from various perspectives, means for generating a video using the character's commentary text, means for distributing the generated video, means for playing the news commentary video on a smart device, means for generating commentary text using a generative AI model, and means for using prompt sentences as input to the generative AI model, thereby enabling users to easily watch news commentary from multiple perspectives in video format.

[0156] "News information" is information about current events and happenings obtained from media such as newspapers, television, and websites.

[0157] A "summary" is a short sentence that succinctly summarizes the content of the original news article.

[0158] A "commentary character" is a fictional character with a particular perspective or expertise who provides multifaceted commentary on news information.

[0159] "Explanatory text" is a sentence used by the commentary character to explain the news information.

[0160] "Video" refers to a video file created based on the explanatory text of the commentary character.

[0161] A "smart device" is an electronic device with internet connectivity, such as a smartphone, tablet, or smart glasses.

[0162] A "generative AI model" is an algorithm that uses artificial intelligence techniques to generate output data of a specific format from input data.

[0163] A "prompt sentence" is initial input data that is input to a generative AI model to elicit a specific output.

[0164] A system embodying the present invention combines various types of hardware and software to provide users with video news commentary from multiple viewpoints.

[0165] 1. System Programming

[0166] We use a program that acquires news information, summarizes it, generates explanatory text by multiple commentators, creates a video based on that text, and performs a series of processes to distribute that video. This program has the following components:

[0167] 2. Server processing

[0168] The server retrieves the latest news information from a news API via the Internet and stores it in a database. Next, it uses natural language processing technology to summarize the news article, and generates commentary text from a different perspective by a commentary character based on the summary. Based on this commentary text, a video is created using speech synthesis technology (e.g., gTTS) and video generation technology (e.g., moviepy). The generated video is then uploaded to a video distribution platform.

[0169] Specific hardware and software usage examples:

[0170] Hardware: Servers, smart devices (smartphones, tablets, etc.)

[0171] Software: News APIs (e.g., NewsAPI), natural language processing libraries (e.g., spaCy), speech synthesis tools (e.g., gTTS), video editing libraries (e.g., moviepy)

[0172] 3. User terminal processing

[0173] Users access the video distribution platform through a dedicated application on their smart devices to watch news commentary videos. The application provides a user interface and has the ability to select and play news commentary videos.

[0174] 4. Generative AI Model Details

[0175] This system uses a generative AI model to generate explanatory text. The model is fed a prompt sentence in advance, and the explanatory text is generated based on the prompt sentence. The prompt sentence is appropriately set to suit the summary of the news information or a specific commentary perspective.

[0176] (Example)

[0177] News article: "COVID-19 vaccines begin to roll out"

[0178] Example prompt sentence:

[0179] "Government perspective: From the government's perspective, the rollout of COVID-19 vaccines is extremely important."

[0180] "Consumer perspective: For consumers, the rollout of this vaccine provides peace of mind."

[0181] "Expert perspective: Experts see this vaccine rollout as a major step in health protection."

[0182] This allows users to visually and aurally understand news commentary from multiple perspectives, improving the viewing experience.

[0183] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0184] Step 1:

[0185] The server obtains news information via the Internet.

[0186] Input: Request from News API

[0187] Data processing: Analyze news data (headlines, article content, etc.) obtained from the news API and store it in a database.

[0188] Output: Retrieved news information (news data in JSON format)

[0189] Step 2:

[0190] The news information obtained by the server is summarized.

[0191] Input: News article content

[0192] Data Computing: Using natural language processing techniques (e.g., spaCy), we extract key points and keywords from news articles and generate summaries.

[0193] Output: Summarized news information (short sentence format)

[0194] Step 3:

[0195] The server uses multiple commentary characters to generate text that explains the summarized news information from various perspectives.

[0196] Input: Summarized news information

[0197] Data computation: Use a generative AI model (e.g., GPT-3) to generate explanatory text based on different perspectives, using the prompt sentence as input.

[0198] Output: Explanatory text with different commentary characters (government perspective, consumer perspective, expert perspective, etc.)

[0199] Step 4:

[0200] Based on the explanatory text generated by the server, a video is generated by combining voice reading and animation for each character.

[0201] Input:Descriptive text

[0202] Data processing: Use a speech synthesis tool (e.g., gTTS) to convert the explanatory text into audio, and use a video editing library (e.g., moviepy) to combine the audio and animation to generate a video.

[0203] Output: Explainer video file

[0204] Step 5:

[0205] The server uploads the generated video to a dedicated video site and prepares it for distribution.

[0206] Input: Explanation video file

[0207] Data processing: Upload to video sites and set up ad insertion.

[0208] Output: Explanation video ready for distribution

[0209] Step 6:

[0210] The user's device accesses the video distribution platform and watches the news commentary video.

[0211] Input: URL of video streaming platform, dedicated application on user's device

[0212] Data manipulation: Video selection and playback through the user interface.

[0213] Output: Instructional video for users to watch

[0214] This will realize a system that links the server and user devices to generate and distribute news commentary videos from multiple perspectives, allowing users to easily view them.

[0215] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0216] MODE FOR CARRYING OUT THE INVENTION

[0217] The system for implementing the present invention includes a series of processing means for acquiring news information, summarizing it, generating videos in which multiple commentary characters provide commentary from various perspectives, and customizing the video content and advertisements by recognizing the user's emotions. Below, the processing steps of the server, terminal, and user are explained with concrete examples.

[0218] 1. Acquiring news information

[0219] The server retrieves news information from a news source. This is done, for example, by using the news site's API. The server sends a request to the news source at a specific time (for example, every morning at 8:00) to retrieve the latest news information. The retrieved news is then stored in a database.

[0220] Examples:

[0221] The server accesses a news API on the Internet and receives the latest news headlines and article content.

[0222] 2. News Summary

[0223] The server summarizes the news information it has acquired. It uses natural language processing technology to extract key points from news articles and summarize them in a concise format. The results of this summary are also stored in the database.

[0224] Examples:

[0225] The server analyzes the title, date, time, and text of the news article, extracts key keywords and key points, and generates a summary of about 100 characters.

[0226] 3. Generating explanatory text

[0227] The server loads multiple commentary characters and generates commentary text for each commentary character, allowing for commentary from a variety of perspectives, including government, consumer, and expert perspectives.

[0228] Examples:

[0229] The server generates text explaining the news from the perspective of character A, and then generates explanatory text from the perspective of character B.

[0230] 4. Customization with Emotion Engine

[0231] The emotion engine recognizes the user's emotions. The user's emotion data is acquired, for example, through real-time facial expression analysis using a camera or microphone, or through voice tone analysis. The emotion engine then customizes explanatory text and advertisements based on this data.

[0232] Examples:

[0233] If the emotion engine detects that the user is feeling stressed from their facial expression, it will change the explanatory text to something more relaxing and select an advertisement that will relieve stress.

[0234] 5. Generating explanatory videos

[0235] The server takes the generated commentary text and character information, and then uses the emotion engine to generate a video using the customized commentary text, including character voice-overs, animations, and related graphics and advertisements.

[0236] Examples:

[0237] The server uses an emotion engine to synthesize the customized commentary text, combines the voice with character animation to generate a video file, and also inserts advertisements selected to match the emotion into the video.

[0238] 6. Preparing and Streaming Videos

[0239] The server uploads the generated video to a dedicated video site, where advertisements selected by the emotion engine are inserted into the video, and the video is ready to be distributed free of charge through an advertising model. Users can then access the video site and watch the video.

[0240] Examples:

[0241] The server uploads the explanatory videos to the video site and configures the settings to insert advertisements selected by the emotion engine. Users access the site and select and watch news explanatory videos that interest them.

[0242] 7. User Video Viewing

[0243] Users access a dedicated video site and watch news commentary videos that are customized to their emotions, allowing them to receive information that is appropriate for their emotional state.

[0244] Examples:

[0245] Users access a video site and find videos they are interested in by category or from the latest news list. When they select a video and start playing it, the emotional engine displays customized content and advertisements, and they can watch news commentary.

[0246] This system not only allows users to receive news from multiple perspectives in an easy-to-understand format, but also provides a more personalized experience by providing information and advertisements that are appropriate to their emotional state.

[0247] The processing flow will be explained below.

[0248] MODE FOR CARRYING OUT THE INVENTION

[0249] Step 1: Get news information

[0250] 1.1. The server checks the news source's API endpoint.

[0251] 1.2. The server sends an API request on a specified schedule (e.g., every morning at 8:00).

[0252] 1.3. The server receives the news data and extracts information such as the title, text, date and time, and URL.

[0253] 1.4. The server stores the extracted news data in a database.

[0254] Specific behavior:

[0255] The server accesses the "News API" and requests the latest news. It analyzes the received news data, extracts the necessary information (title, text, date and time, URL), and records it in the database.

[0256] Step 2: Summarize the news

[0257] 2.1. The server loads the stored news article.

[0258] 2.2. The server parses the article using a natural language processing library.

[0259] 2.3. The server extracts key points and keywords and generates a news summary.

[0260] 2.4. The server associates the generated summaries with the original news data and stores them in a database.

[0261] Specific behavior:

[0262] The server applies natural language processing (NLP) techniques to summarize news articles, extracting key points from the articles, and then creates a concise summary based on the extracted information and stores it in a database.

[0263] Step 3: Generate explanatory text

[0264] 3.1. The server loads data for multiple commentary characters (e.g., character profiles and attributes).

[0265] 3.2. The server analyzes the news from each character's perspective.

[0266] 3.3. The server generates explanatory text from each character's perspective using a natural language generation model.

[0267] 3.4. The server stores the generated explanatory text in a database.

[0268] Specific behavior:

[0269] The server reads the attributes of the commentary characters (e.g., government officials, consumers, experts, etc.) and analyzes the news article from each character's perspective. Based on the analysis results, it uses natural language generation technology to create commentary text and records it in a database.

[0270] Step 4: Customization with Emotion Engine

[0271] 4.1. The server starts the emotion engine.

[0272] 4.2. The server acquires the user's emotion data.

[0273] 4.3. The server customizes the commentary text based on the emotion data.

[0274] 4.4. The server selects advertisements based on the emotion data.

[0275] Specific behavior:

[0276] The server activates the emotion engine to obtain the user's emotion data in real time, and uses facial expression recognition technology and voice tone analysis to identify the user's emotion, and then customizes the explanatory text and displayed advertisements accordingly.

[0277] Step 5: Generate an explainer video

[0278] 5.1. The server retrieves the customized description text and character information.

[0279] 5.2. The server launches a video generation engine (e.g., OpenCV, FFmpeg).

[0280] 5.3. The server generates the character's voice data and combines it with the animation.

[0281] 5.4. The server builds the explainer video and adds any necessary graphics and effects.

[0282] 5.5. The server saves the completed video file in storage.

[0283] Specific behavior:

[0284] The server converts the customized commentary text into audio data using speech synthesis technology, synchronizes it with the corresponding character animation, and also incorporates emotionally appropriate advertisements into the video, before saving the completed video file to storage.

[0285] Step 6: Prepare and stream your video

[0286] 6.1. The server accesses the management interface of the dedicated video site.

[0287] 6.2. The server uploads the generated video to the site.

[0288] 6.3. Configure the server to insert ads into the ad slots.

[0289] 6.4. The server sets the video to public and makes it available for viewing.

[0290] Specific behavior:

[0291] The server logs in to the video site and uploads the saved explanatory video. It then configures the site so that emotionally appropriate ads are displayed when the video is played, enabling free distribution under an advertising model. It also configures the site so that newly added videos are available for viewing when users access the site.

[0292] Step 7: User Watches Video

[0293] 7.1. The user accesses a dedicated video site.

[0294] 7.2. The user selects the news commentary video that interests them.

[0295] 7.3. The user plays the video and watches the news commentary.

[0296] Specific behavior:

[0297] Users access a video site and find videos they are interested in by category or from the latest news list. When they select a video and start playing it, an advertisement is displayed first, and then they can watch a news commentary video with content customized by the emotion engine.

[0298] The system allows users to receive news from multiple perspectives in an easy-to-understand format, and provides information and advertising tailored to the user's emotional state, enabling a more personalized experience.

[0299] Example 2

[0300] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0301] Conventional news distribution systems simply collect and distribute information, but are unable to respond to individual users' emotions and needs. This limits the news understanding and viewing experience, making it difficult to provide personalized information tailored to individual interests and emotions. Furthermore, there is a lack of efficient ways to summarize news or provide commentary from various perspectives, making it difficult to provide information in a way that is sufficiently useful to the recipient.

[0302] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0303] In this invention, the server includes means for acquiring news information, means for summarizing the acquired news information, means for generating text that uses multiple commentary characters to explain the summarized news information from various perspectives, means for recognizing a user's emotions using an emotion engine and customizing the commentary text and advertisements, means for generating videos using the character commentary texts, and means for delivering the generated videos. This makes it possible to provide personalized news information and customize information delivery and advertisements according to the user's emotions.

[0304] "News information" refers to information such as articles, headlines, publication dates, and authors obtained from various news sources.

[0305] A "summary" is a short sentence that extracts important points from news information and summarizes them in a concise form.

[0306] A "commentary character" is a fictional person or character whose role is to explain the news from different perspectives (e.g., government perspective, consumer perspective, expert perspective, etc.).

[0307] "Explanatory text" is a sentence generated by the commentary character to explain news information from various perspectives.

[0308] An "emotion engine" is a technology or system for recognizing a user's emotions and customizing explanatory text and advertisements based on those emotions.

[0309] The "means for generating video" refers to a technology that converts explanatory text into audio and then combines character animations and related graphics based on that audio to create a video.

[0310] "Means for distributing the generated video" refers to a technology or system for uploading the generated explanatory video to a video distribution platform on the Internet and making it available for public viewing.

[0311] MODE FOR CARRYING OUT THE INVENTION

[0312] The system for implementing the present invention includes a series of processing means for acquiring news information, summarizing it, generating videos in which multiple commentary characters provide commentary from various perspectives, and customizing the content of the videos and advertisements by recognizing the user's emotions. The processing flow of the server, terminal, and user is explained below with concrete examples.

[0313] Getting news information

[0314] The server retrieves news information from news sources. This retrieval is done using the news site's API (for example, Google News API). The server sends a request to the news source at a specific time (for example, every morning at 8:00) to retrieve the latest news information. The retrieved news is then stored in a database.

[0315] Examples:

[0316] The server accesses a news API on the Internet (e.g., Google News API) and receives the latest news headlines and article content.

[0317] News Summary

[0318] The server uses natural language processing technology (e.g., generative AI models such as BERT and GPT-3) to summarize the acquired news information. It extracts key points from the news article and summarizes them in a concise form. This summary is also stored in the database.

[0319] Examples:

[0320] The server analyzes the title, date, time, and text of the news article, extracts key keywords and key points, and generates a summary of about 100 characters. For example, a summary like "The world's first clean energy plant has begun operation, marking a new step forward in environmental protection" might be generated.

[0321] Generate explanatory text

[0322] The server reads multiple commentary characters and generates commentary text for the news based on each commentary character, thereby constructing commentary from various perspectives.

[0323] Examples:

[0324] The server generates text explaining the news from the perspective of character A (e.g., the government's perspective), and then generates explanatory text from the perspective of character B (e.g., the consumer's perspective).

[0325] Customization with Emotion Engine

[0326] The emotion engine recognizes the user's emotions. The user's emotion data is acquired in real time using the device's (PC or smartphone) camera and microphone. The emotion engine analyzes this data to determine the user's emotional state (e.g., joy, sadness, stress, etc.) and customizes explanatory text and advertisements accordingly.

[0327] Examples:

[0328] If the emotion engine detects that the user is feeling stressed from their facial expression, the explanatory text will be changed to something more relaxing, and an ad that will help relieve stress will be selected, such as an ad that says, "Take deep breaths while listening to relaxing music."

[0329] Generate explainer videos

[0330] The server generates a video based on the customized explanatory text. It uses text-to-speech technology (e.g., Google Text-to-Speech API) to convert the explanatory text into speech, and then combines the character animations and related graphics to create a video.

[0331] Examples:

[0332] The server uses an emotion engine to synthesize the customized commentary text, then combines the voice with character animation to generate a video file. Furthermore, advertisements selected to match the emotion are inserted into the video. For example, while a commentary character is explaining the news, an advertisement for relaxation goods is displayed on the side of the screen.

[0333] Preparing and streaming video

[0334] The server uploads the generated video to a dedicated video distribution platform (e.g., YouTube). When uploading, it sets up the insertion of advertisements selected by the emotion engine. After that, it makes the video available to users by setting it to public.

[0335] Examples:

[0336] The server uploads the explanatory videos to the video site, and the settings are configured to insert advertisements selected by the emotion engine. Users access the site and select and watch news explanatory videos that interest them.

[0337] User video viewing

[0338] Users access a dedicated video site from their PC, smartphone, or other device and select the news commentary video they want to watch. While the video is playing, customized commentary content and advertisements are displayed that are tailored to the user's emotional state. This allows users to receive information appropriate to their own emotions.

[0339] Examples:

[0340] Users access a video site and find videos of interest by category or from the latest news list. Once a video is selected and playback begins, the emotional engine displays customized content and advertisements, allowing users to watch news commentary. For example, a user feeling stressed can be provided with commentary in a relaxing atmosphere and related advertisements.

[0341] Prompt Sentence Examples

[0342] The following news information has been obtained from a news site. Please generate explanatory text from Character A (government perspective) and Character B (consumer perspective).

[0343] (News article)

[0344] Title: World's first clean energy plant begins operation

[0345] Body text: The world's first clean energy plant officially began operations yesterday. The plant uses renewable energy to significantly reduce carbon dioxide emissions compared to traditional power generation methods.

[0346] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0347] Step 1:

[0348] Getting news information

[0349] The server sends a request to the news source's API endpoint to retrieve the latest news information. Specific input includes parameters such as news category and region. The response from the API includes the news title, text, publication date, etc., and is stored in the database.

[0350] Input: API request parameters such as news category, region, etc.

[0351] Data processing: Extracting necessary information from API responses

[0352] Output: News information stored in a database

[0353] Step 2:

[0354] News Summary

[0355] The server analyzes the acquired news information and uses natural language processing technology to extract key points and generate summaries. Specifically, it extracts important keywords from the text of the news article and constructs a concise summary. The generated summary is then stored in a database.

[0356] Input: News information stored in a database

[0357] Data processing: Extract keywords from news articles and summarize key points

[0358] Output: Summary text stored in the database

[0359] Step 3:

[0360] Generate explanatory text

[0361] The server generates text that explains the summary of the acquired news information based on the perspective of the commentary character. Using a generative AI model, multiple explanatory texts suitable for each character are generated, enabling commentary from a variety of perspectives.

[0362] Input: Summary text stored in the database, character information for commentary

[0363] Data processing: Generate explanatory text using a generative AI model

[0364] Output: Descriptive text stored in the database

[0365] Step 4:

[0366] Customization with Emotion Engine

[0367] The device collects the user's emotional data (facial expressions and voice) in real time, which is then analyzed by the emotion engine. Based on the analysis results, explanatory text and displayed advertisements are customized.

[0368] Input: User's facial expression data, voice data

[0369] Data processing: Data analysis using sentiment analysis algorithms

[0370] Output: Customized explanatory text and advertising information

[0371] Step 5:

[0372] Generate explainer videos

[0373] The server generates a video based on the customized explanatory text, converts the explanatory text into audio using text-to-speech synthesis technology, and combines the audio with character animation and related graphics to create a video file.

[0374] Input: Custom description text, character animation data

[0375] Data processing: text-to-speech synthesis, video editing

[0376] Output: Generated video file

[0377] Step 6:

[0378] Preparing and streaming video

[0379] The server uploads the generated explanatory video to a video distribution platform. After uploading, advertisements selected by the emotion engine are inserted into the video, and the video is made available to users by setting it to public.

[0380] Input: Generated video file, ad data

[0381] Data processing: video upload, ad insertion, publishing settings

[0382] Output: Published instructional video

[0383] Step 7:

[0384] User video viewing

[0385] Users access a dedicated video site using their device and select the news commentary video they want to watch. When the video is played, customized commentary content and advertisements are displayed, allowing users to receive information tailored to their emotions.

[0386] Input: URL of video site, user selection information

[0387] Data processing: User interface operation, video playback

[0388] Output: A customized explainer video to be watched

[0389] (Application example 2)

[0390] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0391] In modern society, the Internet allows users to quickly access a wide variety of news, but it can be difficult for users to find information that is relevant to them from the vast amount of information available. Furthermore, news commentary tends to be biased toward a one-sided perspective, making it difficult for users to understand the content from multiple angles. Furthermore, advertisements included in news and commentary videos are often not relevant to users' interests or emotions, resulting in reduced advertising effectiveness. There is a need for a system that can solve these issues and enable users to receive more personalized news commentary and advertisements.

[0392] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0393] In this invention, the server includes means for acquiring news information, means for summarizing the acquired news information, means for generating text that explains the summarized news information from various perspectives using multiple commentary characters, means for recognizing a user's emotions and customizing the commentary text and advertisements based on the emotions, and means for delivering the generated videos. This not only enables users to receive news from multiple perspectives in an easy-to-understand format, but also enables a more personalized experience by providing information and advertisements that are appropriate for their emotional state.

[0394] "News information" refers to information about the latest happenings and events that is published on the Internet.

[0395] "Means of Acquisition" refers to the technology or process used to collect news information from specific websites or APIs.

[0396] The "summarization method" is a natural language processing technique that extracts the main points of acquired news information and summarizes them in a concise form.

[0397] A "commentary character" is a virtual character with specific characteristics and personality that is used to explain news information from multiple perspectives.

[0398] The "means for generating explanatory text" is a technology for creating text to explain news information from different perspectives for each commentary character.

[0399] "Emotion recognition and customization" refers to technology that analyzes users' emotions in real time and adjusts explanatory text and advertising content accordingly.

[0400] "Means for generating video" refers to a technology that creates a video file that combines audio, animation, and graphics based on explanatory text and character information.

[0401] "Means of distribution" refers to the technology and process for providing the generated video to users in a viewable format via the Internet, etc.

[0402] "Natural language processing technology" refers to technology that allows computers to analyze, understand, and generate human language.

[0403] The "database" is a digital information management system for organizing and storing acquired news information and generated commentary text.

[0404] "Advertisement" refers to information used to notify or solicit specific products or services to users.

[0405] In the system for implementing the present invention, the server, terminal, and user parts operate in cooperation with each other.

[0406] First, the server retrieves news information from a news API on the Internet. This is done using an HTTP request, and the retrieved news data is stored in a database. Next, the server summarizes the news information using natural language processing technology. The summarized news is concisely summarized to around 100 characters and stored again in the database.

[0407] The server generates commentary text based on the summarized news information from the perspective of each commentator. The commentary text is generated from various perspectives, such as the government perspective, expert perspective, and consumer perspective.

[0408] The device is then equipped with a camera and microphone, which capture the user's facial expressions and tone of voice in real time. The emotion engine analyzes this data to recognize the user's emotional state, and customises explanatory text and advertisements based on the results.

[0409] The server generates an explanatory video using the customized explanatory text and advertisements. The video generation combines character voice-overs, animations, and related graphics. The generated video includes advertisements customized according to the user's emotions.

[0410] Finally, the generated video is uploaded to a dedicated video site, where users can access and watch it. The video is displayed with content and advertisements optimized for the user's emotions, providing a more personalized experience for users.

[0411] The hardware required is a high-performance computer for processing on the server, and a device equipped with a camera and microphone for recognizing user emotions. The software used is Python and natural language processing libraries (e.g., Transformers) for acquiring and summarizing news information and generating explanatory text. The emotion engine performs voice and facial expression analysis, and a video distribution platform is used to distribute the generated explanatory videos.

[0412] As a concrete example, the server retrieves news information from a news API every morning at 8 a.m. and generates a summary using natural language processing technology. Multiple commentary characters then generate commentary text from different perspectives, and the system recognizes the user's emotions through the device's camera and microphone, customizing the commentary text and advertisements. For example, if the user is feeling stressed, it will display an advertisement for a relaxing hot spring, and if they are in a happy mood, it will display an advertisement for the latest gadgets.

[0413] An example prompt for a generative AI model might look like this:

[0414] Generate a news summary: "Summarize the following news article in 100 characters or less: [news article text]"

[0415] Sentiment Analysis: "Determine the user's current emotion based on speech and facial expression analysis."

[0416] This allows users to enjoyably understand the news from multiple perspectives, and also allows them to receive appropriate advertisements that match their emotions.

[0417] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0418] Step 1:

[0419] The server sends a request to the news API to retrieve the latest news information. The input includes the news API endpoint and request parameters. The retrieved news data is received in JSON format and stored in a database. Specifically, the HTTP request is executed using the Python Requests library.

[0420] Step 2:

[0421] The server extracts the text of news information retrieved from the database and summarizes it using natural language processing technology. The input includes the text of the retrieved news article. The summary results are then saved back into the database. Specifically, the Transformers library is used to extract key points from the news article and summarize it to about 100 characters.

[0422] Step 3:

[0423] The server generates commentary text from the perspectives of multiple commentary characters. The input includes summarized news information and each character's perspective information. As output, commentary text from each perspective is generated and stored in a database. Specifically, customized commentary text is created using a template corresponding to each character's perspective.

[0424] Step 4:

[0425] The device uses a camera and microphone to capture the user's facial expressions and voice tone. Input includes real-time video and audio data. This data is sent to an emotion engine that analyzes the user's emotional state. This can be done using the Emotion API or similar emotion analysis tools.

[0426] Step 5:

[0427] The emotion engine customizes explanatory text and advertisements based on the user's emotional state. The input includes the emotional data obtained in step 4. The output is customized explanatory text and advertisements appropriate for the emotion. Specifically, if the user is feeling stressed, the content is changed to something that will help them relax, and an appropriate advertisement is selected.

[0428] Step 6:

[0429] The server generates a video by combining character voice-overs and animation based on customized explanatory text and advertisements. The input includes customized explanatory text and advertisement information. The output is the generated video file. Specifically, the system converts text into speech using a TTS (Text-to-Speech) engine and generates a video using an animation tool.

[0430] Step 7:

[0431] The server uploads the generated video to a dedicated video site and prepares it for distribution so that users can access it. The input includes the generated video file. The output is published on the video site and available for users to view. Specifically, the video file is uploaded to the server using an API or FTP and the publishing settings are configured.

[0432] In this way, a series of processes are realized, from acquiring news information to summarizing it, generating explanatory text, customizing it based on the user's emotions, generating videos, and finally delivering it.

[0433] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0434] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0435] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0436] [Second embodiment]

[0437] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0438] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0439] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0440] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0441] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0442] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0443] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0444] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0445] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0446] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0447] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0448] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0449] MODE FOR CARRYING OUT THE INVENTION

[0450] The system for implementing the present invention includes a series of processing means for acquiring news information, summarizing it, generating videos in which multiple commentary characters provide commentary from various perspectives, and finally distributing the videos. Below, the processing steps of the server, terminal, and user are explained with concrete examples.

[0451] 1. Acquiring news information

[0452] The server retrieves news information from a news source. This is done, for example, by using the news site's API. The server sends a request to the news source at a specific time (for example, every morning at 8:00) to retrieve the latest news information. The retrieved news is then stored in a database.

[0453] Examples:

[0454] The server accesses a news API on the Internet and receives the latest news headlines and article content.

[0455] 2. News Summary

[0456] The server summarizes the news information it has acquired. It uses natural language processing technology to extract key points from news articles and summarize them in a concise format. The results of this summary are also stored in the database.

[0457] Examples:

[0458] The server analyzes the title, date, time, and text of the news article, extracts key keywords and key points, and generates a summary of about 100 characters.

[0459] 3. Generating explanatory text

[0460] The server loads multiple commentary characters and generates commentary text for each commentary character, allowing for commentary from a variety of perspectives, including government, consumer, and expert perspectives.

[0461] Examples:

[0462] The server generates text explaining the news from the perspective of character A, and then generates explanatory text from the perspective of character B.

[0463] 4. Generating explanatory videos

[0464] The server generates a video based on the generated explanatory text, combining voice-overs and animations of each character. Related graphics and images are also added to the video, making it visually easy to understand.

[0465] Examples:

[0466] The server synthesizes the text of commentary character A into voice and combines the voice with the character's animation to generate a video file.

[0467] 5. Preparing and Streaming Videos

[0468] The server uploads the generated video to a dedicated video site, where advertisements are inserted and the video is ready to be distributed for free under an advertising model. Users can access the video site and watch the video.

[0469] Examples:

[0470] The server uploads the explanatory video to the video site and configures it to insert an advertisement in the first five seconds. Users access the site and select and watch the news explanatory video that interests them.

[0471] This system allows users to receive news from multiple perspectives in an easy-to-understand format. The server periodically retrieves news and quickly provides the latest information, ensuring that fresh news commentary videos are always delivered.

[0472] The processing flow will be explained below.

[0473] Specific processing steps of the program

[0474] Step 1: Get news information

[0475] 1.1. The server checks the news source's API endpoint.

[0476] 1.2. The server sends an API request on a specified schedule (e.g., every morning at 8:00).

[0477] 1.3. The server receives the news data and extracts information such as the title, text, date and time, and URL.

[0478] 1.4. The server stores the extracted news data in a database.

[0479] Specific behavior:

[0480] The server accesses the "News API" and requests the latest news. It analyzes the received news data, extracts the necessary information (title, text, date and time, URL), and records it in the database.

[0481] Step 2: Summarize the news

[0482] 2.1. The server loads the stored news article.

[0483] 2.2. The server parses the article using a natural language processing library.

[0484] 2.3. The server extracts key points and keywords and generates a news summary.

[0485] 2.4. The server associates the generated summaries with the original news data and stores them in a database.

[0486] Specific behavior:

[0487] The server applies natural language processing (NLP) techniques to summarize news articles, extracting key points from the articles, and then creates a concise summary based on the extracted information and stores it in a database.

[0488] Step 3: Generate explanatory text

[0489] 3.1. The server loads data for multiple commentary characters (e.g., character profiles and attributes).

[0490] 3.2. The server analyzes the news from each character's perspective.

[0491] 3.3. The server generates explanatory text from each character's perspective using a natural language generation model.

[0492] 3.4. The server stores the generated explanatory text in a database.

[0493] Specific behavior:

[0494] The server reads the attributes of the commentary characters (e.g., government officials, consumers, experts, etc.) and analyzes the news article from each character's perspective. Based on the analysis results, it uses natural language generation technology to create commentary text and records it in a database.

[0495] Step 4: Generate an explainer video

[0496] 4.1. The server retrieves the generated description text and character information.

[0497] 4.2. The server launches a video generation engine (e.g., OpenCV, FFmpeg).

[0498] 4.3. The server generates the character's voice data and combines it with the animation.

[0499] 4.4. The server builds the explainer video and adds any necessary graphics and effects.

[0500] 4.5. The server saves the completed video file in storage.

[0501] Specific behavior:

[0502] The server converts the explanatory text into audio data using speech synthesis technology and synchronizes it with the corresponding character animation. It also incorporates images and graphs related to the explanatory content into the video, generating the final explanatory video and saving it in storage.

[0503] Step 5: Prepare and stream your video

[0504] 5.1. The server accesses the management interface of the dedicated video site.

[0505] 5.2. The server uploads the generated video to the site.

[0506] 5.3. The server configures the insertion of ads into the ad slots.

[0507] 5.4. The server sets the video to public and makes it available for viewing.

[0508] Specific behavior:

[0509] The server logs in to the video site and uploads the saved instructional video. It then sets up the video so that ads are displayed when the video is played, allowing for free distribution under an advertising model. It also sets up the public settings so that new videos can be viewed when users access the site.

[0510] Step 6: User Watches Video

[0511] 6.1. The user accesses a dedicated video site.

[0512] 6.2. The user selects the news commentary video that interests them.

[0513] 6.3. The user plays the video and watches the news commentary.

[0514] Specific behavior:

[0515] A user visits a video site and finds a video they are interested in by category or from the latest news list. When they select a video and start playing it, an advertisement is displayed first, and then they can watch a news commentary video.

[0516] Example 1

[0517] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0518] Conventional news distribution systems are limited to providing information from a single perspective, making it difficult for users to understand the news from a variety of perspectives. Furthermore, there is a lack of automated means for quickly and effectively delivering the latest news information, and periodic updates are often performed manually. This can result in delayed or biased information being provided to users.

[0519] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0520] In this invention, the server includes means for acquiring news information, means for summarizing the news information, means for generating text that uses multiple commentary characters to explain the summarized news information from various perspectives, means for generating videos using the character commentary texts, and means for distributing the generated videos, thereby enabling users to understand the news from multiple perspectives and always receive the latest news commentary videos quickly.

[0521] "Means for obtaining news information" refers to a function for automatically obtaining the latest news information from news sources or APIs on the Internet.

[0522] The "means for summarizing news information" is a function that uses natural language processing technology to extract important points from retrieved news articles and summarize them in a concise form.

[0523] "A means for generating text that explains summarized news information from various perspectives using multiple commentary characters" is a function that uses a generative AI model to generate explanatory text to explain news information from various perspectives, such as the government perspective, consumer perspective, and expert perspective.

[0524] "Means for generating videos using character explanatory text" is a function that combines voice synthesis technology and animation based on the generated explanatory text to create videos that are easy to understand visually and aurally.

[0525] The "means for distributing the generated video" is a function for uploading the generated video to a dedicated video site and distributing the video so that users can view it.

[0526] MODE FOR CARRYING OUT THE INVENTION

[0527] The system for implementing this invention includes a series of processing means for acquiring news information, summarizing it, generating videos in which multiple commentary characters provide commentary from various perspectives, and finally distributing the videos. This system functions through the following processing steps by the server, terminal, and user.

[0528] The server sends a request to the API of a news site on the Internet and automatically obtains the latest news information. For example, an HTTP GET request is made to the news site's API endpoint, and the obtained news data is stored in a database. This ensures that the latest news information is always recorded.

[0529] The server then uses natural language processing (NLP) techniques to summarize the retrieved news information. For example, NLP models such as BERT are used to extract key points from news articles and summarize them in a concise format. The summary results are also stored in a database for further processing.

[0530] Next, the server loads multiple commentary characters and generates news commentary text from each character's perspective. Generative AI models used here include (for example) GPT-2. Users can receive news commentary from different commentary characters' perspectives, such as the government perspective, consumer perspective, and expert perspective.

[0531] The server generates videos based on the generated explanatory text, combining each character's voice reading with animation. The Google Text-to-Speech API is used for voice synthesis, and FFmpeg is used as video editing software to generate the video. Related graphics and images are also added to the generated video, making it visually easy to understand.

[0532] As a final step, the server uploads the generated video to a dedicated video site. For example, the video can be uploaded using the YouTube API and configured to insert advertisements. Users can then access this video site and select and watch news commentary videos of their interest.

[0533] Example prompt sentence:

[0534] Below is a major news article. Use this article as a starting point to generate commentary from the government's perspective, the consumer's perspective, and the expert's perspective.

[0535] News article title: {News title}

[0536] News article content: {News content}

[0537] This system allows users to receive news from multiple perspectives in an easy-to-understand format. The server periodically retrieves news and quickly provides the latest information, ensuring that fresh news commentary videos are always delivered.

[0538] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0539] Step 1:

[0540] The server sends a request to the API of a news site on the Internet to obtain the latest news information. The input is the API endpoint URL, and the output is news data in JSON format. Specifically, the server sends an HTTP GET request according to a regular schedule (for example, every morning at 8:00), and the obtained news data is stored in a database.

[0541] Input: News API endpoint URL

[0542] Output: News data (JSON format)

[0543] Specific operation: The server sends an HTTP GET request and receives news data.

[0544] Step 2:

[0545] The server summarizes the news information it obtains. The input is news data, and the output is summarized news text. Natural language processing (NLP) technology is used to extract key points from news articles and summarize them in a concise form. The generated summary is also stored in a database.

[0546] Input: News data

[0547] Output: Summary text

[0548] What it does: The server uses an NLP model to analyze a news article and generate a summary.

[0549] Step 3:

[0550] The server loads multiple commentary characters and generates commentary text from each character's perspective based on the summarized news information. The input is the summary text, and the output is the commentary text. A generative AI model is used to create commentary from various perspectives.

[0551] Input: Summary text

[0552] Output:Descriptive text

[0553] Specific operation: The server inputs a prompt sentence into the generative AI model and generates explanatory text from various perspectives.

[0554] Example prompt sentence:

[0555] Below is a major news article. Use this article as a starting point to generate commentary from the government's perspective, the consumer's perspective, and the expert's perspective.

[0556] News article title: {News title}

[0557] News article content: {News content}

[0558] Step 4:

[0559] The server generates a video that combines voice reading and animation based on the generated explanatory text. The input is the explanatory text, and the output is a completed video file. It integrates speech synthesis technology (e.g., Google Text-to-Speech API) and animation to create a visual video.

[0560] Input:Descriptive text

[0561] Output: Video file

[0562] Specific operation: The server synthesizes explanatory text into speech and combines it with character animation to generate a video.

[0563] Step 5:

[0564] The server uploads the generated video to a dedicated video site and prepares it for distribution. The input is a video file, and the output is the published video content. The video is uploaded using the API of the video distribution platform (for example, YouTube API).

[0565] Input: Video file

[0566] Output: Published video content

[0567] Specific operation: The server uploads the video file to the video site and performs settings such as inserting advertisements.

[0568] (Application example 1)

[0569] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0570] In today's world, news information is extremely diverse, making it difficult for users to efficiently understand it. Furthermore, while commentary on news topics requires commentary from multiple perspectives, there is a lack of easy ways to view such information. Furthermore, systems that provide commentary on news topics from multiple perspectives, rather than a single one, would be effective in helping users gain a deeper understanding of specific news topics, but such systems are currently limited.

[0571] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0572] In this invention, the server includes means for acquiring news information, means for summarizing the acquired news information, means for generating text that uses multiple commentary characters to explain the summarized news information from various perspectives, means for generating a video using the character's commentary text, means for distributing the generated video, means for playing the news commentary video on a smart device, means for generating commentary text using a generative AI model, and means for using prompt sentences as input to the generative AI model, thereby enabling users to easily watch news commentary from multiple perspectives in video format.

[0573] "News information" is information about current events and happenings obtained from media such as newspapers, television, and websites.

[0574] A "summary" is a short sentence that succinctly summarizes the content of the original news article.

[0575] A "commentary character" is a fictional character with a particular perspective or expertise who provides multifaceted commentary on news information.

[0576] "Explanatory text" is a sentence used by the commentary character to explain the news information.

[0577] "Video" refers to a video file created based on the explanatory text of the commentary character.

[0578] A "smart device" is an electronic device with internet connectivity, such as a smartphone, tablet, or smart glasses.

[0579] A "generative AI model" is an algorithm that uses artificial intelligence techniques to generate output data of a specific format from input data.

[0580] A "prompt sentence" is initial input data that is input to a generative AI model to elicit a specific output.

[0581] A system embodying the present invention combines various types of hardware and software to provide users with video news commentary from multiple viewpoints.

[0582] 1. System Programming

[0583] We use a program that acquires news information, summarizes it, generates explanatory text by multiple commentators, creates a video based on that text, and performs a series of processes to distribute that video. This program has the following components:

[0584] 2. Server processing

[0585] The server retrieves the latest news information from a news API via the Internet and stores it in a database. Next, it uses natural language processing technology to summarize the news article, and generates commentary text from a different perspective by a commentary character based on the summary. Based on this commentary text, a video is created using speech synthesis technology (e.g., gTTS) and video generation technology (e.g., moviepy). The generated video is then uploaded to a video distribution platform.

[0586] Specific hardware and software usage examples:

[0587] Hardware: Servers, smart devices (smartphones, tablets, etc.)

[0588] Software: News APIs (e.g., NewsAPI), natural language processing libraries (e.g., spaCy), speech synthesis tools (e.g., gTTS), video editing libraries (e.g., moviepy)

[0589] 3. User terminal processing

[0590] Users access the video distribution platform through a dedicated application on their smart devices to watch news commentary videos. The application provides a user interface and has the ability to select and play news commentary videos.

[0591] 4. Generative AI Model Details

[0592] This system uses a generative AI model to generate explanatory text. The model is fed a prompt sentence in advance, and the explanatory text is generated based on the prompt sentence. The prompt sentence is appropriately set to suit the summary of the news information or a specific commentary perspective.

[0593] (Example)

[0594] News article: "COVID-19 vaccines begin to roll out"

[0595] Example prompt sentence:

[0596] "Government perspective: From the government's perspective, the rollout of COVID-19 vaccines is extremely important."

[0597] "Consumer perspective: For consumers, the rollout of this vaccine provides peace of mind."

[0598] "Expert perspective: Experts see this vaccine rollout as a major step in health protection."

[0599] This allows users to visually and aurally understand news commentary from multiple perspectives, improving the viewing experience.

[0600] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0601] Step 1:

[0602] The server obtains news information via the Internet.

[0603] Input: Request from News API

[0604] Data processing: Analyze news data (headlines, article content, etc.) obtained from the news API and store it in a database.

[0605] Output: Retrieved news information (news data in JSON format)

[0606] Step 2:

[0607] The news information obtained by the server is summarized.

[0608] Input: News article content

[0609] Data Computing: Using natural language processing techniques (e.g., spaCy), we extract key points and keywords from news articles and generate summaries.

[0610] Output: Summarized news information (short sentence format)

[0611] Step 3:

[0612] The server uses multiple commentary characters to generate text that explains the summarized news information from various perspectives.

[0613] Input: Summarized news information

[0614] Data computation: Use a generative AI model (e.g., GPT-3) to generate explanatory text based on different perspectives, using the prompt sentence as input.

[0615] Output: Explanatory text with different commentary characters (government perspective, consumer perspective, expert perspective, etc.)

[0616] Step 4:

[0617] Based on the explanatory text generated by the server, a video is generated by combining voice reading and animation for each character.

[0618] Input:Descriptive text

[0619] Data processing: Use a speech synthesis tool (e.g., gTTS) to convert the explanatory text into audio, and use a video editing library (e.g., moviepy) to combine the audio and animation to generate a video.

[0620] Output: Explainer video file

[0621] Step 5:

[0622] The server uploads the generated video to a dedicated video site and prepares it for distribution.

[0623] Input: Explanation video file

[0624] Data processing: Upload to video sites and set up ad insertion.

[0625] Output: Explanation video ready for distribution

[0626] Step 6:

[0627] The user's device accesses the video distribution platform and watches the news commentary video.

[0628] Input: URL of video streaming platform, dedicated application on user's device

[0629] Data manipulation: Video selection and playback through the user interface.

[0630] Output: Instructional video for users to watch

[0631] This will realize a system that links the server and user devices to generate and distribute news commentary videos from multiple perspectives, allowing users to easily view them.

[0632] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0633] MODE FOR CARRYING OUT THE INVENTION

[0634] The system for implementing the present invention includes a series of processing means for acquiring news information, summarizing it, generating videos in which multiple commentary characters provide commentary from various perspectives, and customizing the video content and advertisements by recognizing the user's emotions. Below, the processing steps of the server, terminal, and user are explained with concrete examples.

[0635] 1. Acquiring news information

[0636] The server retrieves news information from a news source. This is done, for example, by using the news site's API. The server sends a request to the news source at a specific time (for example, every morning at 8:00) to retrieve the latest news information. The retrieved news is then stored in a database.

[0637] Examples:

[0638] The server accesses a news API on the Internet and receives the latest news headlines and article content.

[0639] 2. News Summary

[0640] The server summarizes the news information it has acquired. It uses natural language processing technology to extract key points from news articles and summarize them in a concise format. The results of this summary are also stored in the database.

[0641] Examples:

[0642] The server analyzes the title, date, time, and text of the news article, extracts key keywords and key points, and generates a summary of about 100 characters.

[0643] 3. Generating explanatory text

[0644] The server loads multiple commentary characters and generates commentary text for each commentary character, allowing for commentary from a variety of perspectives, including government, consumer, and expert perspectives.

[0645] Examples:

[0646] The server generates text explaining the news from the perspective of character A, and then generates explanatory text from the perspective of character B.

[0647] 4. Customization with Emotion Engine

[0648] The emotion engine recognizes the user's emotions. The user's emotion data is acquired, for example, through real-time facial expression analysis using a camera or microphone, or through voice tone analysis. The emotion engine then customizes explanatory text and advertisements based on this data.

[0649] Examples:

[0650] If the emotion engine detects that the user is feeling stressed from their facial expression, it will change the explanatory text to something more relaxing and select an advertisement that will relieve stress.

[0651] 5. Generating explanatory videos

[0652] The server takes the generated commentary text and character information, and then uses the emotion engine to generate a video using the customized commentary text, including character voice-overs, animations, and related graphics and advertisements.

[0653] Examples:

[0654] The server uses an emotion engine to synthesize the customized commentary text, combines the voice with character animation to generate a video file, and also inserts advertisements selected to match the emotion into the video.

[0655] 6. Preparing and Streaming Videos

[0656] The server uploads the generated video to a dedicated video site, where advertisements selected by the emotion engine are inserted into the video, and the video is ready to be distributed free of charge through an advertising model. Users can then access the video site and watch the video.

[0657] Examples:

[0658] The server uploads the explanatory videos to the video site and configures the settings to insert advertisements selected by the emotion engine. Users access the site and select and watch news explanatory videos that interest them.

[0659] 7. User Video Viewing

[0660] Users access a dedicated video site and watch news commentary videos that are customized to their emotions, allowing them to receive information that is appropriate for their emotional state.

[0661] Examples:

[0662] Users access a video site and find videos they are interested in by category or from the latest news list. When they select a video and start playing it, the emotional engine displays customized content and advertisements, and they can watch news commentary.

[0663] This system not only allows users to receive news from multiple perspectives in an easy-to-understand format, but also provides a more personalized experience by providing information and advertisements that are appropriate to their emotional state.

[0664] The processing flow will be explained below.

[0665] MODE FOR CARRYING OUT THE INVENTION

[0666] Step 1: Get news information

[0667] 1.1. The server checks the news source's API endpoint.

[0668] 1.2. The server sends an API request on a specified schedule (e.g., every morning at 8:00).

[0669] 1.3. The server receives the news data and extracts information such as the title, text, date and time, and URL.

[0670] 1.4. The server stores the extracted news data in a database.

[0671] Specific behavior:

[0672] The server accesses the "News API" and requests the latest news. It analyzes the received news data, extracts the necessary information (title, text, date and time, URL), and records it in the database.

[0673] Step 2: Summarize the news

[0674] 2.1. The server loads the stored news article.

[0675] 2.2. The server parses the article using a natural language processing library.

[0676] 2.3. The server extracts key points and keywords and generates a news summary.

[0677] 2.4. The server associates the generated summaries with the original news data and stores them in a database.

[0678] Specific behavior:

[0679] The server applies natural language processing (NLP) techniques to summarize news articles, extracting key points from the articles, and then creates a concise summary based on the extracted information and stores it in a database.

[0680] Step 3: Generate explanatory text

[0681] 3.1. The server loads data for multiple commentary characters (e.g., character profiles and attributes).

[0682] 3.2. The server analyzes the news from each character's perspective.

[0683] 3.3. The server generates explanatory text from each character's perspective using a natural language generation model.

[0684] 3.4. The server stores the generated explanatory text in a database.

[0685] Specific behavior:

[0686] The server reads the attributes of the commentary characters (e.g., government officials, consumers, experts, etc.) and analyzes the news article from each character's perspective. Based on the analysis results, it uses natural language generation technology to create commentary text and records it in a database.

[0687] Step 4: Customization with Emotion Engine

[0688] 4.1. The server starts the emotion engine.

[0689] 4.2. The server acquires the user's emotion data.

[0690] 4.3. The server customizes the commentary text based on the emotion data.

[0691] 4.4. The server selects advertisements based on the emotion data.

[0692] Specific behavior:

[0693] The server activates the emotion engine to obtain the user's emotion data in real time, and uses facial expression recognition technology and voice tone analysis to identify the user's emotion, and then customizes the explanatory text and displayed advertisements accordingly.

[0694] Step 5: Generate an explainer video

[0695] 5.1. The server retrieves the customized description text and character information.

[0696] 5.2. The server launches a video generation engine (e.g., OpenCV, FFmpeg).

[0697] 5.3. The server generates the character's voice data and combines it with the animation.

[0698] 5.4. The server builds the explainer video and adds any necessary graphics and effects.

[0699] 5.5. The server saves the completed video file in storage.

[0700] Specific behavior:

[0701] The server converts the customized commentary text into audio data using speech synthesis technology, synchronizes it with the corresponding character animation, and also incorporates emotionally appropriate advertisements into the video, before saving the completed video file to storage.

[0702] Step 6: Prepare and stream your video

[0703] 6.1. The server accesses the management interface of the dedicated video site.

[0704] 6.2. The server uploads the generated video to the site.

[0705] 6.3. Configure the server to insert ads into the ad slots.

[0706] 6.4. The server sets the video to public and makes it available for viewing.

[0707] Specific behavior:

[0708] The server logs in to the video site and uploads the saved explanatory video. It then configures the site so that emotionally appropriate ads are displayed when the video is played, enabling free distribution under an advertising model. It also configures the site so that newly added videos are available for viewing when users access the site.

[0709] Step 7: User Watches Video

[0710] 7.1. The user accesses a dedicated video site.

[0711] 7.2. The user selects the news commentary video that interests them.

[0712] 7.3. The user plays the video and watches the news commentary.

[0713] Specific behavior:

[0714] Users access a video site and find videos they are interested in by category or from the latest news list. When they select a video and start playing it, an advertisement is displayed first, and then they can watch a news commentary video with content customized by the emotion engine.

[0715] The system allows users to receive news from multiple perspectives in an easy-to-understand format, and provides information and advertising tailored to the user's emotional state, enabling a more personalized experience.

[0716] Example 2

[0717] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0718] Conventional news distribution systems simply collect and distribute information, but are unable to respond to individual users' emotions and needs. This limits the news understanding and viewing experience, making it difficult to provide personalized information tailored to individual interests and emotions. Furthermore, there is a lack of efficient ways to summarize news or provide commentary from various perspectives, making it difficult to provide information in a way that is sufficiently useful to the recipient.

[0719] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0720] In this invention, the server includes means for acquiring news information, means for summarizing the acquired news information, means for generating text that uses multiple commentary characters to explain the summarized news information from various perspectives, means for recognizing a user's emotions using an emotion engine and customizing the commentary text and advertisements, means for generating videos using the character commentary texts, and means for delivering the generated videos. This makes it possible to provide personalized news information and customize information delivery and advertisements according to the user's emotions.

[0721] "News information" refers to information such as articles, headlines, publication dates, and authors obtained from various news sources.

[0722] A "summary" is a short sentence that extracts important points from news information and summarizes them in a concise form.

[0723] A "commentary character" is a fictional person or character whose role is to explain the news from different perspectives (e.g., government perspective, consumer perspective, expert perspective, etc.).

[0724] "Explanatory text" is a sentence generated by the commentary character to explain news information from various perspectives.

[0725] An "emotion engine" is a technology or system for recognizing a user's emotions and customizing explanatory text and advertisements based on those emotions.

[0726] The "means for generating video" refers to a technology that converts explanatory text into audio and then combines character animations and related graphics based on that audio to create a video.

[0727] "Means for distributing the generated video" refers to a technology or system for uploading the generated explanatory video to a video distribution platform on the Internet and making it available for public viewing.

[0728] MODE FOR CARRYING OUT THE INVENTION

[0729] The system for implementing the present invention includes a series of processing means for acquiring news information, summarizing it, generating videos in which multiple commentary characters provide commentary from various perspectives, and customizing the content of the videos and advertisements by recognizing the user's emotions. The processing flow of the server, terminal, and user is explained below with concrete examples.

[0730] Getting news information

[0731] The server retrieves news information from news sources. This retrieval is done using the news site's API (for example, Google News API). The server sends a request to the news source at a specific time (for example, every morning at 8:00) to retrieve the latest news information. The retrieved news is then stored in a database.

[0732] Examples:

[0733] The server accesses a news API on the Internet (e.g., Google News API) and receives the latest news headlines and article content.

[0734] News Summary

[0735] The server uses natural language processing technology (e.g., generative AI models such as BERT and GPT-3) to summarize the acquired news information. It extracts key points from the news article and summarizes them in a concise form. This summary is also stored in the database.

[0736] Examples:

[0737] The server analyzes the title, date, time, and text of the news article, extracts key keywords and key points, and generates a summary of about 100 characters. For example, a summary like "The world's first clean energy plant has begun operation, marking a new step forward in environmental protection" might be generated.

[0738] Generate explanatory text

[0739] The server reads multiple commentary characters and generates commentary text for the news based on each commentary character, thereby constructing commentary from various perspectives.

[0740] Examples:

[0741] The server generates text explaining the news from the perspective of character A (e.g., the government's perspective), and then generates explanatory text from the perspective of character B (e.g., the consumer's perspective).

[0742] Customization with Emotion Engine

[0743] The emotion engine recognizes the user's emotions. The user's emotion data is acquired in real time using the device's (PC or smartphone) camera and microphone. The emotion engine analyzes this data to determine the user's emotional state (e.g., joy, sadness, stress, etc.) and customizes explanatory text and advertisements accordingly.

[0744] Examples:

[0745] If the emotion engine detects that the user is feeling stressed from their facial expression, the explanatory text will be changed to something more relaxing, and an ad that will help relieve stress will be selected, such as an ad that says, "Take deep breaths while listening to relaxing music."

[0746] Generate explainer videos

[0747] The server generates a video based on the customized explanatory text. It uses text-to-speech technology (e.g., Google Text-to-Speech API) to convert the explanatory text into speech, and then combines the character animations and related graphics to create a video.

[0748] Examples:

[0749] The server uses an emotion engine to synthesize the customized commentary text, then combines the voice with character animation to generate a video file. Furthermore, advertisements selected to match the emotion are inserted into the video. For example, while a commentary character is explaining the news, an advertisement for relaxation goods is displayed on the side of the screen.

[0750] Preparing and streaming video

[0751] The server uploads the generated video to a dedicated video distribution platform (e.g., YouTube). When uploading, it sets up the insertion of advertisements selected by the emotion engine. After that, it makes the video available to users by setting it to public.

[0752] Examples:

[0753] The server uploads the explanatory videos to the video site, and the settings are configured to insert advertisements selected by the emotion engine. Users access the site and select and watch news explanatory videos that interest them.

[0754] User video viewing

[0755] Users access a dedicated video site from their PC, smartphone, or other device and select the news commentary video they want to watch. While the video is playing, customized commentary content and advertisements are displayed that are tailored to the user's emotional state. This allows users to receive information appropriate to their own emotions.

[0756] Examples:

[0757] Users access a video site and find videos of interest by category or from the latest news list. Once a video is selected and playback begins, the emotional engine displays customized content and advertisements, allowing users to watch news commentary. For example, a user feeling stressed can be provided with commentary in a relaxing atmosphere and related advertisements.

[0758] Prompt Sentence Examples

[0759] The following news information has been obtained from a news site. Please generate explanatory text from Character A (government perspective) and Character B (consumer perspective).

[0760] (News article)

[0761] Title: World's first clean energy plant begins operation

[0762] Body text: The world's first clean energy plant officially began operations yesterday. The plant uses renewable energy to significantly reduce carbon dioxide emissions compared to traditional power generation methods.

[0763] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0764] Step 1:

[0765] Getting news information

[0766] The server sends a request to the news source's API endpoint to retrieve the latest news information. Specific input includes parameters such as news category and region. The response from the API includes the news title, text, publication date, etc., and is stored in the database.

[0767] Input: API request parameters such as news category, region, etc.

[0768] Data processing: Extracting necessary information from API responses

[0769] Output: News information stored in a database

[0770] Step 2:

[0771] News Summary

[0772] The server analyzes the acquired news information and uses natural language processing technology to extract key points and generate summaries. Specifically, it extracts important keywords from the text of the news article and constructs a concise summary. The generated summary is then stored in a database.

[0773] Input: News information stored in a database

[0774] Data processing: Extract keywords from news articles and summarize key points

[0775] Output: Summary text stored in the database

[0776] Step 3:

[0777] Generate explanatory text

[0778] The server generates text that explains the summary of the acquired news information based on the perspective of the commentary character. Using a generative AI model, multiple explanatory texts suitable for each character are generated, enabling commentary from a variety of perspectives.

[0779] Input: Summary text stored in the database, character information for commentary

[0780] Data processing: Generate explanatory text using a generative AI model

[0781] Output: Descriptive text stored in the database

[0782] Step 4:

[0783] Customization with Emotion Engine

[0784] The device collects the user's emotional data (facial expressions and voice) in real time, which is then analyzed by the emotion engine. Based on the analysis results, explanatory text and displayed advertisements are customized.

[0785] Input: User's facial expression data, voice data

[0786] Data processing: Data analysis using sentiment analysis algorithms

[0787] Output: Customized explanatory text and advertising information

[0788] Step 5:

[0789] Generate explainer videos

[0790] The server generates a video based on the customized explanatory text, converts the explanatory text into audio using text-to-speech synthesis technology, and combines the audio with character animation and related graphics to create a video file.

[0791] Input: Custom description text, character animation data

[0792] Data processing: text-to-speech synthesis, video editing

[0793] Output: Generated video file

[0794] Step 6:

[0795] Preparing and streaming video

[0796] The server uploads the generated explanatory video to a video distribution platform. After uploading, advertisements selected by the emotion engine are inserted into the video, and the video is made available to users by setting it to public.

[0797] Input: Generated video file, ad data

[0798] Data processing: video upload, ad insertion, publishing settings

[0799] Output: Published instructional video

[0800] Step 7:

[0801] User video viewing

[0802] Users access a dedicated video site using their device and select the news commentary video they want to watch. When the video is played, customized commentary content and advertisements are displayed, allowing users to receive information tailored to their emotions.

[0803] Input: URL of video site, user selection information

[0804] Data processing: User interface operation, video playback

[0805] Output: A customized explainer video to be watched

[0806] (Application example 2)

[0807] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0808] In modern society, the Internet allows users to quickly access a wide variety of news, but it can be difficult for users to find information that is relevant to them from the vast amount of information available. Furthermore, news commentary tends to be biased toward a one-sided perspective, making it difficult for users to understand the content from multiple angles. Furthermore, advertisements included in news and commentary videos are often not relevant to users' interests or emotions, resulting in reduced advertising effectiveness. There is a need for a system that can solve these issues and enable users to receive more personalized news commentary and advertisements.

[0809] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0810] In this invention, the server includes means for acquiring news information, means for summarizing the acquired news information, means for generating text that explains the summarized news information from various perspectives using multiple commentary characters, means for recognizing a user's emotions and customizing the commentary text and advertisements based on the emotions, and means for delivering the generated videos. This not only enables users to receive news from multiple perspectives in an easy-to-understand format, but also enables a more personalized experience by providing information and advertisements that are appropriate for their emotional state.

[0811] "News information" refers to information about the latest happenings and events that is published on the Internet.

[0812] "Means of Acquisition" refers to the technology or process used to collect news information from specific websites or APIs.

[0813] The "summarization method" is a natural language processing technique that extracts the main points of acquired news information and summarizes them in a concise form.

[0814] A "commentary character" is a virtual character with specific characteristics and personality that is used to explain news information from multiple perspectives.

[0815] The "means for generating explanatory text" is a technology for creating text to explain news information from different perspectives for each commentary character.

[0816] "Emotion recognition and customization" refers to technology that analyzes users' emotions in real time and adjusts explanatory text and advertising content accordingly.

[0817] "Means for generating video" refers to a technology that creates a video file that combines audio, animation, and graphics based on explanatory text and character information.

[0818] "Means of distribution" refers to the technology and process for providing the generated video to users in a viewable format via the Internet, etc.

[0819] "Natural language processing technology" refers to technology that allows computers to analyze, understand, and generate human language.

[0820] The "database" is a digital information management system for organizing and storing acquired news information and generated commentary text.

[0821] "Advertisement" refers to information used to notify or solicit specific products or services to users.

[0822] In the system for implementing the present invention, the server, terminal, and user parts operate in cooperation with each other.

[0823] First, the server retrieves news information from a news API on the Internet. This is done using an HTTP request, and the retrieved news data is stored in a database. Next, the server summarizes the news information using natural language processing technology. The summarized news is concisely summarized to around 100 characters and stored again in the database.

[0824] The server generates commentary text based on the summarized news information from the perspective of each commentator. The commentary text is generated from various perspectives, such as the government perspective, expert perspective, and consumer perspective.

[0825] The device is then equipped with a camera and microphone, which capture the user's facial expressions and tone of voice in real time. The emotion engine analyzes this data to recognize the user's emotional state, and customises explanatory text and advertisements based on the results.

[0826] The server generates an explanatory video using the customized explanatory text and advertisements. The video generation combines character voice-overs, animations, and related graphics. The generated video includes advertisements customized according to the user's emotions.

[0827] Finally, the generated video is uploaded to a dedicated video site, where users can access and watch it. The video is displayed with content and advertisements optimized for the user's emotions, providing a more personalized experience for users.

[0828] The hardware required is a high-performance computer for processing on the server, and a device equipped with a camera and microphone for recognizing user emotions. The software used is Python and natural language processing libraries (e.g., Transformers) for acquiring and summarizing news information and generating explanatory text. The emotion engine performs voice and facial expression analysis, and a video distribution platform is used to distribute the generated explanatory videos.

[0829] As a concrete example, the server retrieves news information from a news API every morning at 8 a.m. and generates a summary using natural language processing technology. Multiple commentary characters then generate commentary text from different perspectives, and the system recognizes the user's emotions through the device's camera and microphone, customizing the commentary text and advertisements. For example, if the user is feeling stressed, it will display an advertisement for a relaxing hot spring, and if they are in a happy mood, it will display an advertisement for the latest gadgets.

[0830] An example prompt for a generative AI model might look like this:

[0831] Generate a news summary: "Summarize the following news article in 100 characters or less: [news article text]"

[0832] Sentiment Analysis: "Determine the user's current emotion based on speech and facial expression analysis."

[0833] This allows users to enjoyably understand the news from multiple perspectives, and also allows them to receive appropriate advertisements that match their emotions.

[0834] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0835] Step 1:

[0836] The server sends a request to the news API to retrieve the latest news information. The input includes the news API endpoint and request parameters. The retrieved news data is received in JSON format and stored in a database. Specifically, the HTTP request is executed using the Python Requests library.

[0837] Step 2:

[0838] The server extracts the text of news information retrieved from the database and summarizes it using natural language processing technology. The input includes the text of the retrieved news article. The summary results are then saved back into the database. Specifically, the Transformers library is used to extract key points from the news article and summarize it to about 100 characters.

[0839] Step 3:

[0840] The server generates commentary text from the perspectives of multiple commentary characters. The input includes summarized news information and each character's perspective information. As output, commentary text from each perspective is generated and stored in a database. Specifically, customized commentary text is created using a template corresponding to each character's perspective.

[0841] Step 4:

[0842] The device uses a camera and microphone to capture the user's facial expressions and voice tone. Input includes real-time video and audio data. This data is sent to an emotion engine that analyzes the user's emotional state. This can be done using the Emotion API or similar emotion analysis tools.

[0843] Step 5:

[0844] The emotion engine customizes explanatory text and advertisements based on the user's emotional state. The input includes the emotional data obtained in step 4. The output is customized explanatory text and advertisements appropriate for the emotion. Specifically, if the user is feeling stressed, the content is changed to something that will help them relax, and an appropriate advertisement is selected.

[0845] Step 6:

[0846] The server generates a video by combining character voice-overs and animation based on customized explanatory text and advertisements. The input includes customized explanatory text and advertisement information. The output is the generated video file. Specifically, the system converts text into speech using a TTS (Text-to-Speech) engine and generates a video using an animation tool.

[0847] Step 7:

[0848] The server uploads the generated video to a dedicated video site and prepares it for distribution so that users can access it. The input includes the generated video file. The output is published on the video site and available for users to view. Specifically, the video file is uploaded to the server using an API or FTP and the publishing settings are configured.

[0849] In this way, a series of processes are realized, from acquiring news information to summarizing it, generating explanatory text, customizing it based on the user's emotions, generating videos, and finally delivering it.

[0850] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0851] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0852] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0853] [Third embodiment]

[0854] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0855] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[0856] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0857] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0858] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0859] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0860] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0861] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0862] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0863] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0864] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0865] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0866] MODE FOR CARRYING OUT THE INVENTION

[0867] The system for implementing the present invention includes a series of processing means for acquiring news information, summarizing it, generating videos in which multiple commentary characters provide commentary from various perspectives, and finally distributing the videos. Below, the processing steps of the server, terminal, and user are explained with concrete examples.

[0868] 1. Acquiring news information

[0869] The server retrieves news information from a news source. This is done, for example, by using the news site's API. The server sends a request to the news source at a specific time (for example, every morning at 8:00) to retrieve the latest news information. The retrieved news is then stored in a database.

[0870] Examples:

[0871] The server accesses a news API on the Internet and receives the latest news headlines and article content.

[0872] 2. News Summary

[0873] The server summarizes the news information it has acquired. It uses natural language processing technology to extract key points from news articles and summarize them in a concise format. The results of this summary are also stored in the database.

[0874] Examples:

[0875] The server analyzes the title, date, time, and text of the news article, extracts key keywords and key points, and generates a summary of about 100 characters.

[0876] 3. Generating explanatory text

[0877] The server loads multiple commentary characters and generates commentary text for each commentary character, allowing for commentary from a variety of perspectives, including government, consumer, and expert perspectives.

[0878] Examples:

[0879] The server generates text explaining the news from the perspective of character A, and then generates explanatory text from the perspective of character B.

[0880] 4. Generating explanatory videos

[0881] The server generates a video based on the generated explanatory text, combining voice-overs and animations of each character. Related graphics and images are also added to the video, making it visually easy to understand.

[0882] Examples:

[0883] The server synthesizes the text of commentary character A into voice and combines the voice with the character's animation to generate a video file.

[0884] 5. Preparing and Streaming Videos

[0885] The server uploads the generated video to a dedicated video site, where advertisements are inserted and the video is ready to be distributed for free under an advertising model. Users can access the video site and watch the video.

[0886] Examples:

[0887] The server uploads the explanatory video to the video site and configures it to insert an advertisement in the first five seconds. Users access the site and select and watch the news explanatory video that interests them.

[0888] This system allows users to receive news from multiple perspectives in an easy-to-understand format. The server periodically retrieves news and quickly provides the latest information, ensuring that fresh news commentary videos are always delivered.

[0889] The processing flow will be explained below.

[0890] Specific processing steps of the program

[0891] Step 1: Get news information

[0892] 1.1. The server checks the news source's API endpoint.

[0893] 1.2. The server sends an API request on a specified schedule (e.g., every morning at 8:00).

[0894] 1.3. The server receives the news data and extracts information such as the title, text, date and time, and URL.

[0895] 1.4. The server stores the extracted news data in a database.

[0896] Specific behavior:

[0897] The server accesses the "News API" and requests the latest news. It analyzes the received news data, extracts the necessary information (title, text, date and time, URL), and records it in the database.

[0898] Step 2: Summarize the news

[0899] 2.1. The server loads the stored news article.

[0900] 2.2. The server parses the article using a natural language processing library.

[0901] 2.3. The server extracts key points and keywords and generates a news summary.

[0902] 2.4. The server associates the generated summaries with the original news data and stores them in a database.

[0903] Specific behavior:

[0904] The server applies natural language processing (NLP) techniques to summarize news articles, extracting key points from the articles, and then creates a concise summary based on the extracted information and stores it in a database.

[0905] Step 3: Generate explanatory text

[0906] 3.1. The server loads data for multiple commentary characters (e.g., character profiles and attributes).

[0907] 3.2. The server analyzes the news from each character's perspective.

[0908] 3.3. The server generates explanatory text from each character's perspective using a natural language generation model.

[0909] 3.4. The server stores the generated explanatory text in a database.

[0910] Specific behavior:

[0911] The server reads the attributes of the commentary characters (e.g., government officials, consumers, experts, etc.) and analyzes the news article from each character's perspective. Based on the analysis results, it uses natural language generation technology to create commentary text and records it in a database.

[0912] Step 4: Generate an explainer video

[0913] 4.1. The server retrieves the generated description text and character information.

[0914] 4.2. The server launches a video generation engine (e.g., OpenCV, FFmpeg).

[0915] 4.3. The server generates the character's voice data and combines it with the animation.

[0916] 4.4. The server builds the explainer video and adds any necessary graphics and effects.

[0917] 4.5. The server saves the completed video file in storage.

[0918] Specific behavior:

[0919] The server converts the explanatory text into audio data using speech synthesis technology and synchronizes it with the corresponding character animation. It also incorporates images and graphs related to the explanatory content into the video, generating the final explanatory video and saving it in storage.

[0920] Step 5: Prepare and stream your video

[0921] 5.1. The server accesses the management interface of the dedicated video site.

[0922] 5.2. The server uploads the generated video to the site.

[0923] 5.3. The server configures the insertion of ads into the ad slots.

[0924] 5.4. The server sets the video to public and makes it available for viewing.

[0925] Specific behavior:

[0926] The server logs in to the video site and uploads the saved instructional video. It then sets up the video so that ads are displayed when the video is played, allowing for free distribution under an advertising model. It also sets up the public settings so that new videos can be viewed when users access the site.

[0927] Step 6: User Watches Video

[0928] 6.1. The user accesses a dedicated video site.

[0929] 6.2. The user selects the news commentary video that interests them.

[0930] 6.3. The user plays the video and watches the news commentary.

[0931] Specific behavior:

[0932] A user visits a video site and finds a video they are interested in by category or from the latest news list. When they select a video and start playing it, an advertisement is displayed first, and then they can watch a news commentary video.

[0933] Example 1

[0934] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0935] Conventional news distribution systems are limited to providing information from a single perspective, making it difficult for users to understand the news from a variety of perspectives. Furthermore, there is a lack of automated means for quickly and effectively delivering the latest news information, and periodic updates are often performed manually. This can result in delayed or biased information being provided to users.

[0936] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0937] In this invention, the server includes means for acquiring news information, means for summarizing the news information, means for generating text that uses multiple commentary characters to explain the summarized news information from various perspectives, means for generating videos using the character commentary texts, and means for distributing the generated videos, thereby enabling users to understand the news from multiple perspectives and always receive the latest news commentary videos quickly.

[0938] "Means for obtaining news information" refers to a function for automatically obtaining the latest news information from news sources or APIs on the Internet.

[0939] The "means for summarizing news information" is a function that uses natural language processing technology to extract important points from retrieved news articles and summarize them in a concise form.

[0940] "A means for generating text that explains summarized news information from various perspectives using multiple commentary characters" is a function that uses a generative AI model to generate explanatory text to explain news information from various perspectives, such as the government perspective, consumer perspective, and expert perspective.

[0941] "Means for generating videos using character explanatory text" is a function that combines voice synthesis technology and animation based on the generated explanatory text to create videos that are easy to understand visually and aurally.

[0942] The "means for distributing the generated video" is a function for uploading the generated video to a dedicated video site and distributing the video so that users can view it.

[0943] MODE FOR CARRYING OUT THE INVENTION

[0944] The system for implementing this invention includes a series of processing means for acquiring news information, summarizing it, generating videos in which multiple commentary characters provide commentary from various perspectives, and finally distributing the videos. This system functions through the following processing steps by the server, terminal, and user.

[0945] The server sends a request to the API of a news site on the Internet and automatically obtains the latest news information. For example, an HTTP GET request is made to the news site's API endpoint, and the obtained news data is stored in a database. This ensures that the latest news information is always recorded.

[0946] The server then uses natural language processing (NLP) techniques to summarize the retrieved news information. For example, NLP models such as BERT are used to extract key points from news articles and summarize them in a concise format. The summary results are also stored in a database for further processing.

[0947] Next, the server loads multiple commentary characters and generates news commentary text from each character's perspective. Generative AI models used here include (for example) GPT-2. Users can receive news commentary from different commentary characters' perspectives, such as the government perspective, consumer perspective, and expert perspective.

[0948] The server generates videos based on the generated explanatory text, combining each character's voice reading with animation. The Google Text-to-Speech API is used for voice synthesis, and FFmpeg is used as video editing software to generate the video. Related graphics and images are also added to the generated video, making it visually easy to understand.

[0949] As a final step, the server uploads the generated video to a dedicated video site. For example, the video can be uploaded using the YouTube API and configured to insert advertisements. Users can then access this video site and select and watch news commentary videos of their interest.

[0950] Example prompt sentence:

[0951] Below is a major news article. Use this article as a starting point to generate commentary from the government's perspective, the consumer's perspective, and the expert's perspective.

[0952] News article title: {News title}

[0953] News article content: {News content}

[0954] This system allows users to receive news from multiple perspectives in an easy-to-understand format. The server periodically retrieves news and quickly provides the latest information, ensuring that fresh news commentary videos are always delivered.

[0955] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0956] Step 1:

[0957] The server sends a request to the API of a news site on the Internet to obtain the latest news information. The input is the API endpoint URL, and the output is news data in JSON format. Specifically, the server sends an HTTP GET request according to a regular schedule (for example, every morning at 8:00), and the obtained news data is stored in a database.

[0958] Input: News API endpoint URL

[0959] Output: News data (JSON format)

[0960] Specific operation: The server sends an HTTP GET request and receives news data.

[0961] Step 2:

[0962] The server summarizes the news information it obtains. The input is news data, and the output is summarized news text. Natural language processing (NLP) technology is used to extract key points from news articles and summarize them in a concise form. The generated summary is also stored in a database.

[0963] Input: News data

[0964] Output: Summary text

[0965] What it does: The server uses an NLP model to analyze a news article and generate a summary.

[0966] Step 3:

[0967] The server loads multiple commentary characters and generates commentary text from each character's perspective based on the summarized news information. The input is the summary text, and the output is the commentary text. A generative AI model is used to create commentary from various perspectives.

[0968] Input: Summary text

[0969] Output:Descriptive text

[0970] Specific operation: The server inputs a prompt sentence into the generative AI model and generates explanatory text from various perspectives.

[0971] Example prompt sentence:

[0972] Below is a major news article. Use this article as a starting point to generate commentary from the government's perspective, the consumer's perspective, and the expert's perspective.

[0973] News article title: {News title}

[0974] News article content: {News content}

[0975] Step 4:

[0976] The server generates a video that combines voice reading and animation based on the generated explanatory text. The input is the explanatory text, and the output is a completed video file. It integrates speech synthesis technology (e.g., Google Text-to-Speech API) and animation to create a visual video.

[0977] Input:Descriptive text

[0978] Output: Video file

[0979] Specific operation: The server synthesizes explanatory text into speech and combines it with character animation to generate a video.

[0980] Step 5:

[0981] The server uploads the generated video to a dedicated video site and prepares it for distribution. The input is a video file, and the output is the published video content. The video is uploaded using the API of the video distribution platform (for example, YouTube API).

[0982] Input: Video file

[0983] Output: Published video content

[0984] Specific operation: The server uploads the video file to the video site and performs settings such as inserting advertisements.

[0985] (Application example 1)

[0986] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0987] In today's world, news information is extremely diverse, making it difficult for users to efficiently understand it. Furthermore, while commentary on news topics requires commentary from multiple perspectives, there is a lack of easy ways to view such information. Furthermore, systems that provide commentary on news topics from multiple perspectives, rather than a single one, would be effective in helping users gain a deeper understanding of specific news topics, but such systems are currently limited.

[0988] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0989] In this invention, the server includes means for acquiring news information, means for summarizing the acquired news information, means for generating text that uses multiple commentary characters to explain the summarized news information from various perspectives, means for generating a video using the character's commentary text, means for distributing the generated video, means for playing the news commentary video on a smart device, means for generating commentary text using a generative AI model, and means for using prompt sentences as input to the generative AI model, thereby enabling users to easily watch news commentary from multiple perspectives in video format.

[0990] "News information" is information about current events and happenings obtained from media such as newspapers, television, and websites.

[0991] A "summary" is a short sentence that succinctly summarizes the content of the original news article.

[0992] A "commentary character" is a fictional character with a particular perspective or expertise who provides multifaceted commentary on news information.

[0993] "Explanatory text" is a sentence used by the commentary character to explain the news information.

[0994] "Video" refers to a video file created based on the explanatory text of the commentary character.

[0995] A "smart device" is an electronic device with internet connectivity, such as a smartphone, tablet, or smart glasses.

[0996] A "generative AI model" is an algorithm that uses artificial intelligence techniques to generate output data of a specific format from input data.

[0997] A "prompt sentence" is initial input data that is input to a generative AI model to elicit a specific output.

[0998] A system embodying the present invention combines various types of hardware and software to provide users with video news commentary from multiple viewpoints.

[0999] 1. System Programming

[1000] We use a program that acquires news information, summarizes it, generates explanatory text by multiple commentators, creates a video based on that text, and performs a series of processes to distribute that video. This program has the following components:

[1001] 2. Server processing

[1002] The server retrieves the latest news information from a news API via the Internet and stores it in a database. Next, it uses natural language processing technology to summarize the news article, and generates commentary text from a different perspective by a commentary character based on the summary. Based on this commentary text, a video is created using speech synthesis technology (e.g., gTTS) and video generation technology (e.g., moviepy). The generated video is then uploaded to a video distribution platform.

[1003] Specific hardware and software usage examples:

[1004] Hardware: Servers, smart devices (smartphones, tablets, etc.)

[1005] Software: News APIs (e.g., NewsAPI), natural language processing libraries (e.g., spaCy), speech synthesis tools (e.g., gTTS), video editing libraries (e.g., moviepy)

[1006] 3. User terminal processing

[1007] Users access the video distribution platform through a dedicated application on their smart devices to watch news commentary videos. The application provides a user interface and has the ability to select and play news commentary videos.

[1008] 4. Generative AI Model Details

[1009] This system uses a generative AI model to generate explanatory text. The model is fed a prompt sentence in advance, and the explanatory text is generated based on the prompt sentence. The prompt sentence is appropriately set to suit the summary of the news information or a specific commentary perspective.

[1010] (Example)

[1011] News article: "COVID-19 vaccines begin to roll out"

[1012] Example prompt sentence:

[1013] "Government perspective: From the government's perspective, the rollout of COVID-19 vaccines is extremely important."

[1014] "Consumer perspective: For consumers, the rollout of this vaccine provides peace of mind."

[1015] "Expert perspective: Experts see this vaccine rollout as a major step in health protection."

[1016] This allows users to visually and aurally understand news commentary from multiple perspectives, improving the viewing experience.

[1017] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1018] Step 1:

[1019] The server obtains news information via the Internet.

[1020] Input: Request from News API

[1021] Data processing: Analyze news data (headlines, article content, etc.) obtained from the news API and store it in a database.

[1022] Output: Retrieved news information (news data in JSON format)

[1023] Step 2:

[1024] The news information obtained by the server is summarized.

[1025] Input: News article content

[1026] Data Computing: Using natural language processing techniques (e.g., spaCy), we extract key points and keywords from news articles and generate summaries.

[1027] Output: Summarized news information (short sentence format)

[1028] Step 3:

[1029] The server uses multiple commentary characters to generate text that explains the summarized news information from various perspectives.

[1030] Input: Summarized news information

[1031] Data computation: Use a generative AI model (e.g., GPT-3) to generate explanatory text based on different perspectives, using the prompt sentence as input.

[1032] Output: Explanatory text with different commentary characters (government perspective, consumer perspective, expert perspective, etc.)

[1033] Step 4:

[1034] Based on the explanatory text generated by the server, a video is generated by combining voice reading and animation for each character.

[1035] Input:Descriptive text

[1036] Data processing: Use a speech synthesis tool (e.g., gTTS) to convert the explanatory text into audio, and use a video editing library (e.g., moviepy) to combine the audio and animation to generate a video.

[1037] Output: Explainer video file

[1038] Step 5:

[1039] The server uploads the generated video to a dedicated video site and prepares it for distribution.

[1040] Input: Explanation video file

[1041] Data processing: Upload to video sites and set up ad insertion.

[1042] Output: Explanation video ready for distribution

[1043] Step 6:

[1044] The user's device accesses the video distribution platform and watches the news commentary video.

[1045] Input: URL of video streaming platform, dedicated application on user's device

[1046] Data manipulation: Video selection and playback through the user interface.

[1047] Output: Instructional video for users to watch

[1048] This will realize a system that links the server and user devices to generate and distribute news commentary videos from multiple perspectives, allowing users to easily view them.

[1049] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1050] MODE FOR CARRYING OUT THE INVENTION

[1051] The system for implementing the present invention includes a series of processing means for acquiring news information, summarizing it, generating videos in which multiple commentary characters provide commentary from various perspectives, and customizing the video content and advertisements by recognizing the user's emotions. Below, the processing steps of the server, terminal, and user are explained with concrete examples.

[1052] 1. Acquiring news information

[1053] The server retrieves news information from a news source. This is done, for example, by using the news site's API. The server sends a request to the news source at a specific time (for example, every morning at 8:00) to retrieve the latest news information. The retrieved news is then stored in a database.

[1054] Examples:

[1055] The server accesses a news API on the Internet and receives the latest news headlines and article content.

[1056] 2. News Summary

[1057] The server summarizes the news information it has acquired. It uses natural language processing technology to extract key points from news articles and summarize them in a concise format. The results of this summary are also stored in the database.

[1058] Examples:

[1059] The server analyzes the title, date, time, and text of the news article, extracts key keywords and key points, and generates a summary of about 100 characters.

[1060] 3. Generating explanatory text

[1061] The server loads multiple commentary characters and generates commentary text for each commentary character, allowing for commentary from a variety of perspectives, including government, consumer, and expert perspectives.

[1062] Examples:

[1063] The server generates text explaining the news from the perspective of character A, and then generates explanatory text from the perspective of character B.

[1064] 4. Customization with Emotion Engine

[1065] The emotion engine recognizes the user's emotions. The user's emotion data is acquired, for example, through real-time facial expression analysis using a camera or microphone, or through voice tone analysis. The emotion engine then customizes explanatory text and advertisements based on this data.

[1066] Examples:

[1067] If the emotion engine detects that the user is feeling stressed from their facial expression, it will change the explanatory text to something more relaxing and select an advertisement that will relieve stress.

[1068] 5. Generating explanatory videos

[1069] The server takes the generated commentary text and character information, and then uses the emotion engine to generate a video using the customized commentary text, including character voice-overs, animations, and related graphics and advertisements.

[1070] Examples:

[1071] The server uses an emotion engine to synthesize the customized commentary text, combines the voice with character animation to generate a video file, and also inserts advertisements selected to match the emotion into the video.

[1072] 6. Preparing and Streaming Videos

[1073] The server uploads the generated video to a dedicated video site, where advertisements selected by the emotion engine are inserted into the video, and the video is ready to be distributed free of charge through an advertising model. Users can then access the video site and watch the video.

[1074] Examples:

[1075] The server uploads the explanatory videos to the video site and configures the settings to insert advertisements selected by the emotion engine. Users access the site and select and watch news explanatory videos that interest them.

[1076] 7. User Video Viewing

[1077] Users access a dedicated video site and watch news commentary videos that are customized to their emotions, allowing them to receive information that is appropriate for their emotional state.

[1078] Examples:

[1079] Users access a video site and find videos they are interested in by category or from the latest news list. When they select a video and start playing it, the emotional engine displays customized content and advertisements, and they can watch news commentary.

[1080] This system not only allows users to receive news from multiple perspectives in an easy-to-understand format, but also provides a more personalized experience by providing information and advertisements that are appropriate to their emotional state.

[1081] The processing flow will be explained below.

[1082] MODE FOR CARRYING OUT THE INVENTION

[1083] Step 1: Get news information

[1084] 1.1. The server checks the news source's API endpoint.

[1085] 1.2. The server sends an API request on a specified schedule (e.g., every morning at 8:00).

[1086] 1.3. The server receives the news data and extracts information such as the title, text, date and time, and URL.

[1087] 1.4. The server stores the extracted news data in a database.

[1088] Specific behavior:

[1089] The server accesses the "News API" and requests the latest news. It analyzes the received news data, extracts the necessary information (title, text, date and time, URL), and records it in the database.

[1090] Step 2: Summarize the news

[1091] 2.1. The server loads the stored news article.

[1092] 2.2. The server parses the article using a natural language processing library.

[1093] 2.3. The server extracts key points and keywords and generates a news summary.

[1094] 2.4. The server associates the generated summaries with the original news data and stores them in a database.

[1095] Specific behavior:

[1096] The server applies natural language processing (NLP) techniques to summarize news articles, extracting key points from the articles, and then creates a concise summary based on the extracted information and stores it in a database.

[1097] Step 3: Generate explanatory text

[1098] 3.1. The server loads data for multiple commentary characters (e.g., character profiles and attributes).

[1099] 3.2. The server analyzes the news from each character's perspective.

[1100] 3.3. The server generates explanatory text from each character's perspective using a natural language generation model.

[1101] 3.4. The server stores the generated explanatory text in a database.

[1102] Specific behavior:

[1103] The server reads the attributes of the commentary characters (e.g., government officials, consumers, experts, etc.) and analyzes the news article from each character's perspective. Based on the analysis results, it uses natural language generation technology to create commentary text and records it in a database.

[1104] Step 4: Customization with Emotion Engine

[1105] 4.1. The server starts the emotion engine.

[1106] 4.2. The server acquires the user's emotion data.

[1107] 4.3. The server customizes the commentary text based on the emotion data.

[1108] 4.4. The server selects advertisements based on the emotion data.

[1109] Specific behavior:

[1110] The server activates the emotion engine to obtain the user's emotion data in real time, and uses facial expression recognition technology and voice tone analysis to identify the user's emotion, and then customizes the explanatory text and displayed advertisements accordingly.

[1111] Step 5: Generate an explainer video

[1112] 5.1. The server retrieves the customized description text and character information.

[1113] 5.2. The server launches a video generation engine (e.g., OpenCV, FFmpeg).

[1114] 5.3. The server generates the character's voice data and combines it with the animation.

[1115] 5.4. The server builds the explainer video and adds any necessary graphics and effects.

[1116] 5.5. The server saves the completed video file in storage.

[1117] Specific behavior:

[1118] The server converts the customized commentary text into audio data using speech synthesis technology, synchronizes it with the corresponding character animation, and also incorporates emotionally appropriate advertisements into the video, before saving the completed video file to storage.

[1119] Step 6: Prepare and stream your video

[1120] 6.1. The server accesses the management interface of the dedicated video site.

[1121] 6.2. The server uploads the generated video to the site.

[1122] 6.3. Configure the server to insert ads into the ad slots.

[1123] 6.4. The server sets the video to public and makes it available for viewing.

[1124] Specific behavior:

[1125] The server logs in to the video site and uploads the saved explanatory video. It then configures the site so that emotionally appropriate ads are displayed when the video is played, enabling free distribution under an advertising model. It also configures the site so that newly added videos are available for viewing when users access the site.

[1126] Step 7: User Watches Video

[1127] 7.1. The user accesses a dedicated video site.

[1128] 7.2. The user selects the news commentary video that interests them.

[1129] 7.3. The user plays the video and watches the news commentary.

[1130] Specific behavior:

[1131] Users access a video site and find videos they are interested in by category or from the latest news list. When they select a video and start playing it, an advertisement is displayed first, and then they can watch a news commentary video with content customized by the emotion engine.

[1132] The system allows users to receive news from multiple perspectives in an easy-to-understand format, and provides information and advertising tailored to the user's emotional state, enabling a more personalized experience.

[1133] Example 2

[1134] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1135] Conventional news distribution systems simply collect and distribute information, but are unable to respond to individual users' emotions and needs. This limits the news understanding and viewing experience, making it difficult to provide personalized information tailored to individual interests and emotions. Furthermore, there is a lack of efficient ways to summarize news or provide commentary from various perspectives, making it difficult to provide information in a way that is sufficiently useful to the recipient.

[1136] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1137] In this invention, the server includes means for acquiring news information, means for summarizing the acquired news information, means for generating text that uses multiple commentary characters to explain the summarized news information from various perspectives, means for recognizing a user's emotions using an emotion engine and customizing the commentary text and advertisements, means for generating videos using the character commentary texts, and means for delivering the generated videos. This makes it possible to provide personalized news information and customize information delivery and advertisements according to the user's emotions.

[1138] "News information" refers to information such as articles, headlines, publication dates, and authors obtained from various news sources.

[1139] A "summary" is a short sentence that extracts important points from news information and summarizes them in a concise form.

[1140] A "commentary character" is a fictional person or character whose role is to explain the news from different perspectives (e.g., government perspective, consumer perspective, expert perspective, etc.).

[1141] "Explanatory text" is a sentence generated by the commentary character to explain news information from various perspectives.

[1142] An "emotion engine" is a technology or system for recognizing a user's emotions and customizing explanatory text and advertisements based on those emotions.

[1143] The "means for generating video" refers to a technology that converts explanatory text into audio and then combines character animations and related graphics based on that audio to create a video.

[1144] "Means for distributing the generated video" refers to a technology or system for uploading the generated explanatory video to a video distribution platform on the Internet and making it available for public viewing.

[1145] MODE FOR CARRYING OUT THE INVENTION

[1146] The system for implementing the present invention includes a series of processing means for acquiring news information, summarizing it, generating videos in which multiple commentary characters provide commentary from various perspectives, and customizing the content of the videos and advertisements by recognizing the user's emotions. The processing flow of the server, terminal, and user is explained below with concrete examples.

[1147] Getting news information

[1148] The server retrieves news information from news sources. This retrieval is done using the news site's API (for example, Google News API). The server sends a request to the news source at a specific time (for example, every morning at 8:00) to retrieve the latest news information. The retrieved news is then stored in a database.

[1149] Examples:

[1150] The server accesses a news API on the Internet (e.g., Google News API) and receives the latest news headlines and article content.

[1151] News Summary

[1152] The server uses natural language processing technology (e.g., generative AI models such as BERT and GPT-3) to summarize the acquired news information. It extracts key points from the news article and summarizes them in a concise form. This summary is also stored in the database.

[1153] Examples:

[1154] The server analyzes the title, date, time, and text of the news article, extracts key keywords and key points, and generates a summary of about 100 characters. For example, a summary like "The world's first clean energy plant has begun operation, marking a new step forward in environmental protection" might be generated.

[1155] Generate explanatory text

[1156] The server reads multiple commentary characters and generates commentary text for the news based on each commentary character, thereby constructing commentary from various perspectives.

[1157] Examples:

[1158] The server generates text explaining the news from the perspective of character A (e.g., the government's perspective), and then generates explanatory text from the perspective of character B (e.g., the consumer's perspective).

[1159] Customization with Emotion Engine

[1160] The emotion engine recognizes the user's emotions. The user's emotion data is acquired in real time using the device's (PC or smartphone) camera and microphone. The emotion engine analyzes this data to determine the user's emotional state (e.g., joy, sadness, stress, etc.) and customizes explanatory text and advertisements accordingly.

[1161] Examples:

[1162] If the emotion engine detects that the user is feeling stressed from their facial expression, the explanatory text will be changed to something more relaxing, and an ad that will help relieve stress will be selected, such as an ad that says, "Take deep breaths while listening to relaxing music."

[1163] Generate explainer videos

[1164] The server generates a video based on the customized explanatory text. It uses text-to-speech technology (e.g., Google Text-to-Speech API) to convert the explanatory text into speech, and then combines the character animations and related graphics to create a video.

[1165] Examples:

[1166] The server uses an emotion engine to synthesize the customized commentary text, then combines the voice with character animation to generate a video file. Furthermore, advertisements selected to match the emotion are inserted into the video. For example, while a commentary character is explaining the news, an advertisement for relaxation goods is displayed on the side of the screen.

[1167] Preparing and streaming video

[1168] The server uploads the generated video to a dedicated video distribution platform (e.g., YouTube). When uploading, it sets up the insertion of advertisements selected by the emotion engine. After that, it makes the video available to users by setting it to public.

[1169] Examples:

[1170] The server uploads the explanatory videos to the video site, and the settings are configured to insert advertisements selected by the emotion engine. Users access the site and select and watch news explanatory videos that interest them.

[1171] User video viewing

[1172] Users access a dedicated video site from their PC, smartphone, or other device and select the news commentary video they want to watch. While the video is playing, customized commentary content and advertisements are displayed that are tailored to the user's emotional state. This allows users to receive information appropriate to their own emotions.

[1173] Examples:

[1174] Users access a video site and find videos of interest by category or from the latest news list. Once a video is selected and playback begins, the emotional engine displays customized content and advertisements, allowing users to watch news commentary. For example, a user feeling stressed can be provided with commentary in a relaxing atmosphere and related advertisements.

[1175] Prompt Sentence Examples

[1176] The following news information has been obtained from a news site. Please generate explanatory text from Character A (government perspective) and Character B (consumer perspective).

[1177] (News article)

[1178] Title: World's first clean energy plant begins operation

[1179] Body text: The world's first clean energy plant officially began operations yesterday. The plant uses renewable energy to significantly reduce carbon dioxide emissions compared to traditional power generation methods.

[1180] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1181] Step 1:

[1182] Getting news information

[1183] The server sends a request to the news source's API endpoint to retrieve the latest news information. Specific input includes parameters such as news category and region. The response from the API includes the news title, text, publication date, etc., and is stored in the database.

[1184] Input: API request parameters such as news category, region, etc.

[1185] Data processing: Extracting necessary information from API responses

[1186] Output: News information stored in a database

[1187] Step 2:

[1188] News Summary

[1189] The server analyzes the acquired news information and uses natural language processing technology to extract key points and generate summaries. Specifically, it extracts important keywords from the text of the news article and constructs a concise summary. The generated summary is then stored in a database.

[1190] Input: News information stored in a database

[1191] Data processing: Extract keywords from news articles and summarize key points

[1192] Output: Summary text stored in the database

[1193] Step 3:

[1194] Generate explanatory text

[1195] The server generates text that explains the summary of the acquired news information based on the perspective of the commentary character. Using a generative AI model, multiple explanatory texts suitable for each character are generated, enabling commentary from a variety of perspectives.

[1196] Input: Summary text stored in the database, character information for commentary

[1197] Data processing: Generate explanatory text using a generative AI model

[1198] Output: Descriptive text stored in the database

[1199] Step 4:

[1200] Customization with Emotion Engine

[1201] The device collects the user's emotional data (facial expressions and voice) in real time, which is then analyzed by the emotion engine. Based on the analysis results, explanatory text and displayed advertisements are customized.

[1202] Input: User's facial expression data, voice data

[1203] Data processing: Data analysis using sentiment analysis algorithms

[1204] Output: Customized explanatory text and advertising information

[1205] Step 5:

[1206] Generate explainer videos

[1207] The server generates a video based on the customized explanatory text, converts the explanatory text into audio using text-to-speech synthesis technology, and combines the audio with character animation and related graphics to create a video file.

[1208] Input: Custom description text, character animation data

[1209] Data processing: text-to-speech synthesis, video editing

[1210] Output: Generated video file

[1211] Step 6:

[1212] Preparing and streaming video

[1213] The server uploads the generated explanatory video to a video distribution platform. After uploading, advertisements selected by the emotion engine are inserted into the video, and the video is made available to users by setting it to public.

[1214] Input: Generated video file, ad data

[1215] Data processing: video upload, ad insertion, publishing settings

[1216] Output: Published instructional video

[1217] Step 7:

[1218] User video viewing

[1219] Users access a dedicated video site using their device and select the news commentary video they want to watch. When the video is played, customized commentary content and advertisements are displayed, allowing users to receive information tailored to their emotions.

[1220] Input: URL of video site, user selection information

[1221] Data processing: User interface operation, video playback

[1222] Output: A customized explainer video to be watched

[1223] (Application example 2)

[1224] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1225] In modern society, the Internet allows users to quickly access a wide variety of news, but it can be difficult for users to find information that is relevant to them from the vast amount of information available. Furthermore, news commentary tends to be biased toward a one-sided perspective, making it difficult for users to understand the content from multiple angles. Furthermore, advertisements included in news and commentary videos are often not relevant to users' interests or emotions, resulting in reduced advertising effectiveness. There is a need for a system that can solve these issues and enable users to receive more personalized news commentary and advertisements.

[1226] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1227] In this invention, the server includes means for acquiring news information, means for summarizing the acquired news information, means for generating text that explains the summarized news information from various perspectives using multiple commentary characters, means for recognizing a user's emotions and customizing the commentary text and advertisements based on the emotions, and means for delivering the generated videos. This not only enables users to receive news from multiple perspectives in an easy-to-understand format, but also enables a more personalized experience by providing information and advertisements that are appropriate for their emotional state.

[1228] "News information" refers to information about the latest happenings and events that is published on the Internet.

[1229] "Means of Acquisition" refers to the technology or process used to collect news information from specific websites or APIs.

[1230] The "summarization method" is a natural language processing technique that extracts the main points of acquired news information and summarizes them in a concise form.

[1231] A "commentary character" is a virtual character with specific characteristics and personality that is used to explain news information from multiple perspectives.

[1232] The "means for generating explanatory text" is a technology for creating text to explain news information from different perspectives for each commentary character.

[1233] "Emotion recognition and customization" refers to technology that analyzes users' emotions in real time and adjusts explanatory text and advertising content accordingly.

[1234] "Means for generating video" refers to a technology that creates a video file that combines audio, animation, and graphics based on explanatory text and character information.

[1235] "Means of distribution" refers to the technology and process for providing the generated video to users in a viewable format via the Internet, etc.

[1236] "Natural language processing technology" refers to technology that allows computers to analyze, understand, and generate human language.

[1237] The "database" is a digital information management system for organizing and storing acquired news information and generated commentary text.

[1238] "Advertisement" refers to information used to notify or solicit specific products or services to users.

[1239] In the system for implementing the present invention, the server, terminal, and user parts operate in cooperation with each other.

[1240] First, the server retrieves news information from a news API on the Internet. This is done using an HTTP request, and the retrieved news data is stored in a database. Next, the server summarizes the news information using natural language processing technology. The summarized news is concisely summarized to around 100 characters and stored again in the database.

[1241] The server generates commentary text based on the summarized news information from the perspective of each commentator. The commentary text is generated from various perspectives, such as the government perspective, expert perspective, and consumer perspective.

[1242] The device is then equipped with a camera and microphone, which capture the user's facial expressions and tone of voice in real time. The emotion engine analyzes this data to recognize the user's emotional state, and customises explanatory text and advertisements based on the results.

[1243] The server generates an explanatory video using the customized explanatory text and advertisements. The video generation combines character voice-overs, animations, and related graphics. The generated video includes advertisements customized according to the user's emotions.

[1244] Finally, the generated video is uploaded to a dedicated video site, where users can access and watch it. The video is displayed with content and advertisements optimized for the user's emotions, providing a more personalized experience for users.

[1245] The hardware required is a high-performance computer for processing on the server, and a device equipped with a camera and microphone for recognizing user emotions. The software used is Python and natural language processing libraries (e.g., Transformers) for acquiring and summarizing news information and generating explanatory text. The emotion engine performs voice and facial expression analysis, and a video distribution platform is used to distribute the generated explanatory videos.

[1246] As a concrete example, the server retrieves news information from a news API every morning at 8 a.m. and generates a summary using natural language processing technology. Multiple commentary characters then generate commentary text from different perspectives, and the system recognizes the user's emotions through the device's camera and microphone, customizing the commentary text and advertisements. For example, if the user is feeling stressed, it will display an advertisement for a relaxing hot spring, and if they are in a happy mood, it will display an advertisement for the latest gadgets.

[1247] An example prompt for a generative AI model might look like this:

[1248] Generate a news summary: "Summarize the following news article in 100 characters or less: [news article text]"

[1249] Sentiment Analysis: "Determine the user's current emotion based on speech and facial expression analysis."

[1250] This allows users to enjoyably understand the news from multiple perspectives, and also allows them to receive appropriate advertisements that match their emotions.

[1251] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1252] Step 1:

[1253] The server sends a request to the news API to retrieve the latest news information. The input includes the news API endpoint and request parameters. The retrieved news data is received in JSON format and stored in a database. Specifically, the HTTP request is executed using the Python Requests library.

[1254] Step 2:

[1255] The server extracts the text of news information retrieved from the database and summarizes it using natural language processing technology. The input includes the text of the retrieved news article. The summary results are then saved back into the database. Specifically, the Transformers library is used to extract key points from the news article and summarize it to about 100 characters.

[1256] Step 3:

[1257] The server generates commentary text from the perspectives of multiple commentary characters. The input includes summarized news information and each character's perspective information. As output, commentary text from each perspective is generated and stored in a database. Specifically, customized commentary text is created using a template corresponding to each character's perspective.

[1258] Step 4:

[1259] The device uses a camera and microphone to capture the user's facial expressions and voice tone. Input includes real-time video and audio data. This data is sent to an emotion engine that analyzes the user's emotional state. This can be done using the Emotion API or similar emotion analysis tools.

[1260] Step 5:

[1261] The emotion engine customizes explanatory text and advertisements based on the user's emotional state. The input includes the emotional data obtained in step 4. The output is customized explanatory text and advertisements appropriate for the emotion. Specifically, if the user is feeling stressed, the content is changed to something that will help them relax, and an appropriate advertisement is selected.

[1262] Step 6:

[1263] The server generates a video by combining character voice-overs and animation based on customized explanatory text and advertisements. The input includes customized explanatory text and advertisement information. The output is the generated video file. Specifically, the system converts text into speech using a TTS (Text-to-Speech) engine and generates a video using an animation tool.

[1264] Step 7:

[1265] The server uploads the generated video to a dedicated video site and prepares it for distribution so that users can access it. The input includes the generated video file. The output is published on the video site and available for users to view. Specifically, the video file is uploaded to the server using an API or FTP and the publishing settings are configured.

[1266] In this way, a series of processes are realized, from acquiring news information to summarizing it, generating explanatory text, customizing it based on the user's emotions, generating videos, and finally delivering it.

[1267] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1268] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1269] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1270] [Fourth embodiment]

[1271] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1272] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1273] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1274] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1275] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1276] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1277] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1278] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1279] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1280] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1281] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1282] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1283] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1284] MODE FOR CARRYING OUT THE INVENTION

[1285] The system for implementing the present invention includes a series of processing means for acquiring news information, summarizing it, generating videos in which multiple commentary characters provide commentary from various perspectives, and finally distributing the videos. Below, the processing steps of the server, terminal, and user are explained with concrete examples.

[1286] 1. Acquiring news information

[1287] The server retrieves news information from a news source. This is done, for example, by using the news site's API. The server sends a request to the news source at a specific time (for example, every morning at 8:00) to retrieve the latest news information. The retrieved news is then stored in a database.

[1288] Examples:

[1289] The server accesses a news API on the Internet and receives the latest news headlines and article content.

[1290] 2. News Summary

[1291] The server summarizes the news information it has acquired. It uses natural language processing technology to extract key points from news articles and summarize them in a concise format. The results of this summary are also stored in the database.

[1292] Examples:

[1293] The server analyzes the title, date, time, and text of the news article, extracts key keywords and key points, and generates a summary of about 100 characters.

[1294] 3. Generating explanatory text

[1295] The server loads multiple commentary characters and generates commentary text for each commentary character, allowing for commentary from a variety of perspectives, including government, consumer, and expert perspectives.

[1296] Examples:

[1297] The server generates text explaining the news from the perspective of character A, and then generates explanatory text from the perspective of character B.

[1298] 4. Generating explanatory videos

[1299] The server generates a video based on the generated explanatory text, combining voice-overs and animations of each character. Related graphics and images are also added to the video, making it visually easy to understand.

[1300] Examples:

[1301] The server synthesizes the text of commentary character A into voice and combines the voice with the character's animation to generate a video file.

[1302] 5. Preparing and Streaming Videos

[1303] The server uploads the generated video to a dedicated video site, where advertisements are inserted and the video is ready to be distributed for free under an advertising model. Users can access the video site and watch the video.

[1304] Examples:

[1305] The server uploads the explanatory video to the video site and configures it to insert an advertisement in the first five seconds. Users access the site and select and watch the news explanatory video that interests them.

[1306] This system allows users to receive news from multiple perspectives in an easy-to-understand format. The server periodically retrieves news and quickly provides the latest information, ensuring that fresh news commentary videos are always delivered.

[1307] The processing flow will be explained below.

[1308] Specific processing steps of the program

[1309] Step 1: Get news information

[1310] 1.1. The server checks the news source's API endpoint.

[1311] 1.2. The server sends an API request on a specified schedule (e.g., every morning at 8:00).

[1312] 1.3. The server receives the news data and extracts information such as the title, text, date and time, and URL.

[1313] 1.4. The server stores the extracted news data in a database.

[1314] Specific behavior:

[1315] The server accesses the "News API" and requests the latest news. It analyzes the received news data, extracts the necessary information (title, text, date and time, URL), and records it in the database.

[1316] Step 2: Summarize the news

[1317] 2.1. The server loads the stored news article.

[1318] 2.2. The server parses the article using a natural language processing library.

[1319] 2.3. The server extracts key points and keywords and generates a news summary.

[1320] 2.4. The server associates the generated summaries with the original news data and stores them in a database.

[1321] Specific behavior:

[1322] The server applies natural language processing (NLP) techniques to summarize news articles, extracting key points from the articles, and then creates a concise summary based on the extracted information and stores it in a database.

[1323] Step 3: Generate explanatory text

[1324] 3.1. The server loads data for multiple commentary characters (e.g., character profiles and attributes).

[1325] 3.2. The server analyzes the news from each character's perspective.

[1326] 3.3. The server generates explanatory text from each character's perspective using a natural language generation model.

[1327] 3.4. The server stores the generated explanatory text in a database.

[1328] Specific behavior:

[1329] The server reads the attributes of the commentary characters (e.g., government officials, consumers, experts, etc.) and analyzes the news article from each character's perspective. Based on the analysis results, it uses natural language generation technology to create commentary text and records it in a database.

[1330] Step 4: Generate an explainer video

[1331] 4.1. The server retrieves the generated description text and character information.

[1332] 4.2. The server launches a video generation engine (e.g., OpenCV, FFmpeg).

[1333] 4.3. The server generates the character's voice data and combines it with the animation.

[1334] 4.4. The server builds the explainer video and adds any necessary graphics and effects.

[1335] 4.5. The server saves the completed video file in storage.

[1336] Specific behavior:

[1337] The server converts the explanatory text into audio data using speech synthesis technology and synchronizes it with the corresponding character animation. It also incorporates images and graphs related to the explanatory content into the video, generating the final explanatory video and saving it in storage.

[1338] Step 5: Prepare and stream your video

[1339] 5.1. The server accesses the management interface of the dedicated video site.

[1340] 5.2. The server uploads the generated video to the site.

[1341] 5.3. The server configures the insertion of ads into the ad slots.

[1342] 5.4. The server sets the video to public and makes it available for viewing.

[1343] Specific behavior:

[1344] The server logs in to the video site and uploads the saved instructional video. It then sets up the video so that ads are displayed when the video is played, allowing for free distribution under an advertising model. It also sets up the public settings so that new videos can be viewed when users access the site.

[1345] Step 6: User Watches Video

[1346] 6.1. The user accesses a dedicated video site.

[1347] 6.2. The user selects the news commentary video that interests them.

[1348] 6.3. The user plays the video and watches the news commentary.

[1349] Specific behavior:

[1350] A user visits a video site and finds a video they are interested in by category or from the latest news list. When they select a video and start playing it, an advertisement is displayed first, and then they can watch a news commentary video.

[1351] Example 1

[1352] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1353] Conventional news distribution systems are limited to providing information from a single perspective, making it difficult for users to understand the news from a variety of perspectives. Furthermore, there is a lack of automated means for quickly and effectively delivering the latest news information, and periodic updates are often performed manually. This can result in delayed or biased information being provided to users.

[1354] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1355] In this invention, the server includes means for acquiring news information, means for summarizing the news information, means for generating text that uses multiple commentary characters to explain the summarized news information from various perspectives, means for generating videos using the character commentary texts, and means for distributing the generated videos, thereby enabling users to understand the news from multiple perspectives and always receive the latest news commentary videos quickly.

[1356] "Means for obtaining news information" refers to a function for automatically obtaining the latest news information from news sources or APIs on the Internet.

[1357] The "means for summarizing news information" is a function that uses natural language processing technology to extract important points from retrieved news articles and summarize them in a concise form.

[1358] "A means for generating text that explains summarized news information from various perspectives using multiple commentary characters" is a function that uses a generative AI model to generate explanatory text to explain news information from various perspectives, such as the government perspective, consumer perspective, and expert perspective.

[1359] "Means for generating videos using character explanatory text" is a function that combines voice synthesis technology and animation based on the generated explanatory text to create videos that are easy to understand visually and aurally.

[1360] The "means for distributing the generated video" is a function for uploading the generated video to a dedicated video site and distributing the video so that users can view it.

[1361] MODE FOR CARRYING OUT THE INVENTION

[1362] The system for implementing this invention includes a series of processing means for acquiring news information, summarizing it, generating videos in which multiple commentary characters provide commentary from various perspectives, and finally distributing the videos. This system functions through the following processing steps by the server, terminal, and user.

[1363] The server sends a request to the API of a news site on the Internet and automatically obtains the latest news information. For example, an HTTP GET request is made to the news site's API endpoint, and the obtained news data is stored in a database. This ensures that the latest news information is always recorded.

[1364] The server then uses natural language processing (NLP) techniques to summarize the retrieved news information. For example, NLP models such as BERT are used to extract key points from news articles and summarize them in a concise format. The summary results are also stored in a database for further processing.

[1365] Next, the server loads multiple commentary characters and generates news commentary text from each character's perspective. Generative AI models used here include (for example) GPT-2. Users can receive news commentary from different commentary characters' perspectives, such as the government perspective, consumer perspective, and expert perspective.

[1366] The server generates videos based on the generated explanatory text, combining each character's voice reading with animation. The Google Text-to-Speech API is used for voice synthesis, and FFmpeg is used as video editing software to generate the video. Related graphics and images are also added to the generated video, making it visually easy to understand.

[1367] As a final step, the server uploads the generated video to a dedicated video site. For example, the video can be uploaded using the YouTube API and configured to insert advertisements. Users can then access this video site and select and watch news commentary videos of their interest.

[1368] Example prompt sentence:

[1369] Below is a major news article. Use this article as a starting point to generate commentary from the government's perspective, the consumer's perspective, and the expert's perspective.

[1370] News article title: {News title}

[1371] News article content: {News content}

[1372] This system allows users to receive news from multiple perspectives in an easy-to-understand format. The server periodically retrieves news and quickly provides the latest information, ensuring that fresh news commentary videos are always delivered.

[1373] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1374] Step 1:

[1375] The server sends a request to the API of a news site on the Internet to obtain the latest news information. The input is the API endpoint URL, and the output is news data in JSON format. Specifically, the server sends an HTTP GET request according to a regular schedule (for example, every morning at 8:00), and the obtained news data is stored in a database.

[1376] Input: News API endpoint URL

[1377] Output: News data (JSON format)

[1378] Specific operation: The server sends an HTTP GET request and receives news data.

[1379] Step 2:

[1380] The server summarizes the news information it obtains. The input is news data, and the output is summarized news text. Natural language processing (NLP) technology is used to extract key points from news articles and summarize them in a concise form. The generated summary is also stored in a database.

[1381] Input: News data

[1382] Output: Summary text

[1383] What it does: The server uses an NLP model to analyze a news article and generate a summary.

[1384] Step 3:

[1385] The server loads multiple commentary characters and generates commentary text from each character's perspective based on the summarized news information. The input is the summary text, and the output is the commentary text. A generative AI model is used to create commentary from various perspectives.

[1386] Input: Summary text

[1387] Output:Descriptive text

[1388] Specific operation: The server inputs a prompt sentence into the generative AI model and generates explanatory text from various perspectives.

[1389] Example prompt sentence:

[1390] Below is a major news article. Use this article as a starting point to generate commentary from the government's perspective, the consumer's perspective, and the expert's perspective.

[1391] News article title: {News title}

[1392] News article content: {News content}

[1393] Step 4:

[1394] The server generates a video that combines voice reading and animation based on the generated explanatory text. The input is the explanatory text, and the output is a completed video file. It integrates speech synthesis technology (e.g., Google Text-to-Speech API) and animation to create a visual video.

[1395] Input:Descriptive text

[1396] Output: Video file

[1397] Specific operation: The server synthesizes explanatory text into speech and combines it with character animation to generate a video.

[1398] Step 5:

[1399] The server uploads the generated video to a dedicated video site and prepares it for distribution. The input is a video file, and the output is the published video content. The video is uploaded using the API of the video distribution platform (for example, YouTube API).

[1400] Input: Video file

[1401] Output: Published video content

[1402] Specific operation: The server uploads the video file to the video site and performs settings such as inserting advertisements.

[1403] (Application example 1)

[1404] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1405] In today's world, news information is extremely diverse, making it difficult for users to efficiently understand it. Furthermore, while commentary on news topics requires commentary from multiple perspectives, there is a lack of easy ways to view such information. Furthermore, systems that provide commentary on news topics from multiple perspectives, rather than a single one, would be effective in helping users gain a deeper understanding of specific news topics, but such systems are currently limited.

[1406] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1407] In this invention, the server includes means for acquiring news information, means for summarizing the acquired news information, means for generating text that uses multiple commentary characters to explain the summarized news information from various perspectives, means for generating a video using the character's commentary text, means for distributing the generated video, means for playing the news commentary video on a smart device, means for generating commentary text using a generative AI model, and means for using prompt sentences as input to the generative AI model, thereby enabling users to easily watch news commentary from multiple perspectives in video format.

[1408] "News information" is information about current events and happenings obtained from media such as newspapers, television, and websites.

[1409] A "summary" is a short sentence that succinctly summarizes the content of the original news article.

[1410] A "commentary character" is a fictional character with a particular perspective or expertise who provides multifaceted commentary on news information.

[1411] "Explanatory text" is a sentence used by the commentary character to explain the news information.

[1412] "Video" refers to a video file created based on the explanatory text of the commentary character.

[1413] A "smart device" is an electronic device with internet connectivity, such as a smartphone, tablet, or smart glasses.

[1414] A "generative AI model" is an algorithm that uses artificial intelligence techniques to generate output data of a specific format from input data.

[1415] A "prompt sentence" is initial input data that is input to a generative AI model to elicit a specific output.

[1416] A system embodying the present invention combines various types of hardware and software to provide users with video news commentary from multiple viewpoints.

[1417] 1. System Programming

[1418] We use a program that acquires news information, summarizes it, generates explanatory text by multiple commentators, creates a video based on that text, and performs a series of processes to distribute that video. This program has the following components:

[1419] 2. Server processing

[1420] The server retrieves the latest news information from a news API via the Internet and stores it in a database. Next, it uses natural language processing technology to summarize the news article, and generates commentary text from a different perspective by a commentary character based on the summary. Based on this commentary text, a video is created using speech synthesis technology (e.g., gTTS) and video generation technology (e.g., moviepy). The generated video is then uploaded to a video distribution platform.

[1421] Specific hardware and software usage examples:

[1422] Hardware: Servers, smart devices (smartphones, tablets, etc.)

[1423] Software: News APIs (e.g., NewsAPI), natural language processing libraries (e.g., spaCy), speech synthesis tools (e.g., gTTS), video editing libraries (e.g., moviepy)

[1424] 3. User terminal processing

[1425] Users access the video distribution platform through a dedicated application on their smart devices to watch news commentary videos. The application provides a user interface and has the ability to select and play news commentary videos.

[1426] 4. Generative AI Model Details

[1427] This system uses a generative AI model to generate explanatory text. The model is fed a prompt sentence in advance, and the explanatory text is generated based on the prompt sentence. The prompt sentence is appropriately set to suit the summary of the news information or a specific commentary perspective.

[1428] (Example)

[1429] News article: "COVID-19 vaccines begin to roll out"

[1430] Example prompt sentence:

[1431] "Government perspective: From the government's perspective, the rollout of COVID-19 vaccines is extremely important."

[1432] "Consumer perspective: For consumers, the rollout of this vaccine provides peace of mind."

[1433] "Expert perspective: Experts see this vaccine rollout as a major step in health protection."

[1434] This allows users to visually and aurally understand news commentary from multiple perspectives, improving the viewing experience.

[1435] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1436] Step 1:

[1437] The server obtains news information via the Internet.

[1438] Input: Request from News API

[1439] Data processing: Analyze news data (headlines, article content, etc.) obtained from the news API and store it in a database.

[1440] Output: Retrieved news information (news data in JSON format)

[1441] Step 2:

[1442] The news information obtained by the server is summarized.

[1443] Input: News article content

[1444] Data Computing: Using natural language processing techniques (e.g., spaCy), we extract key points and keywords from news articles and generate summaries.

[1445] Output: Summarized news information (short sentence format)

[1446] Step 3:

[1447] The server uses multiple commentary characters to generate text that explains the summarized news information from various perspectives.

[1448] Input: Summarized news information

[1449] Data computation: Use a generative AI model (e.g., GPT-3) to generate explanatory text based on different perspectives, using the prompt sentence as input.

[1450] Output: Explanatory text with different commentary characters (government perspective, consumer perspective, expert perspective, etc.)

[1451] Step 4:

[1452] Based on the explanatory text generated by the server, a video is generated by combining voice reading and animation for each character.

[1453] Input:Descriptive text

[1454] Data processing: Use a speech synthesis tool (e.g., gTTS) to convert the explanatory text into audio, and use a video editing library (e.g., moviepy) to combine the audio and animation to generate a video.

[1455] Output: Explainer video file

[1456] Step 5:

[1457] The server uploads the generated video to a dedicated video site and prepares it for distribution.

[1458] Input: Explanation video file

[1459] Data processing: Upload to video sites and set up ad insertion.

[1460] Output: Explanation video ready for distribution

[1461] Step 6:

[1462] The user's device accesses the video distribution platform and watches the news commentary video.

[1463] Input: URL of video streaming platform, dedicated application on user's device

[1464] Data manipulation: Video selection and playback through the user interface.

[1465] Output: Instructional video for users to watch

[1466] This will realize a system that links the server and user devices to generate and distribute news commentary videos from multiple perspectives, allowing users to easily view them.

[1467] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1468] MODE FOR CARRYING OUT THE INVENTION

[1469] The system for implementing the present invention includes a series of processing means for acquiring news information, summarizing it, generating videos in which multiple commentary characters provide commentary from various perspectives, and customizing the video content and advertisements by recognizing the user's emotions. Below, the processing steps of the server, terminal, and user are explained with concrete examples.

[1470] 1. Acquiring news information

[1471] The server retrieves news information from a news source. This is done, for example, by using the news site's API. The server sends a request to the news source at a specific time (for example, every morning at 8:00) to retrieve the latest news information. The retrieved news is then stored in a database.

[1472] Examples:

[1473] The server accesses a news API on the Internet and receives the latest news headlines and article content.

[1474] 2. News Summary

[1475] The server summarizes the news information it has acquired. It uses natural language processing technology to extract key points from news articles and summarize them in a concise format. The results of this summary are also stored in the database.

[1476] Examples:

[1477] The server analyzes the title, date, time, and text of the news article, extracts key keywords and key points, and generates a summary of about 100 characters.

[1478] 3. Generating explanatory text

[1479] The server loads multiple commentary characters and generates commentary text for each commentary character, allowing for commentary from a variety of perspectives, including government, consumer, and expert perspectives.

[1480] Examples:

[1481] The server generates text explaining the news from the perspective of character A, and then generates explanatory text from the perspective of character B.

[1482] 4. Customization with Emotion Engine

[1483] The emotion engine recognizes the user's emotions. The user's emotion data is acquired, for example, through real-time facial expression analysis using a camera or microphone, or through voice tone analysis. The emotion engine then customizes explanatory text and advertisements based on this data.

[1484] Examples:

[1485] If the emotion engine detects that the user is feeling stressed from their facial expression, it will change the explanatory text to something more relaxing and select an advertisement that will relieve stress.

[1486] 5. Generating explanatory videos

[1487] The server takes the generated commentary text and character information, and then uses the emotion engine to generate a video using the customized commentary text, including character voice-overs, animations, and related graphics and advertisements.

[1488] Examples:

[1489] The server uses an emotion engine to synthesize the customized commentary text, combines the voice with character animation to generate a video file, and also inserts advertisements selected to match the emotion into the video.

[1490] 6. Preparing and Streaming Videos

[1491] The server uploads the generated video to a dedicated video site, where advertisements selected by the emotion engine are inserted into the video, and the video is ready to be distributed free of charge through an advertising model. Users can then access the video site and watch the video.

[1492] Examples:

[1493] The server uploads the explanatory videos to the video site and configures the settings to insert advertisements selected by the emotion engine. Users access the site and select and watch news explanatory videos that interest them.

[1494] 7. User Video Viewing

[1495] Users access a dedicated video site and watch news commentary videos that are customized to their emotions, allowing them to receive information that is appropriate for their emotional state.

[1496] Examples:

[1497] Users access a video site and find videos they are interested in by category or from the latest news list. When they select a video and start playing it, the emotional engine displays customized content and advertisements, and they can watch news commentary.

[1498] This system not only allows users to receive news from multiple perspectives in an easy-to-understand format, but also provides a more personalized experience by providing information and advertisements that are appropriate to their emotional state.

[1499] The processing flow will be explained below.

[1500] MODE FOR CARRYING OUT THE INVENTION

[1501] Step 1: Get news information

[1502] 1.1. The server checks the news source's API endpoint.

[1503] 1.2. The server sends an API request on a specified schedule (e.g., every morning at 8:00).

[1504] 1.3. The server receives the news data and extracts information such as the title, text, date and time, and URL.

[1505] 1.4. The server stores the extracted news data in a database.

[1506] Specific behavior:

[1507] The server accesses the "News API" and requests the latest news. It analyzes the received news data, extracts the necessary information (title, text, date and time, URL), and records it in the database.

[1508] Step 2: Summarize the news

[1509] 2.1. The server loads the stored news article.

[1510] 2.2. The server parses the article using a natural language processing library.

[1511] 2.3. The server extracts key points and keywords and generates a news summary.

[1512] 2.4. The server associates the generated summaries with the original news data and stores them in a database.

[1513] Specific behavior:

[1514] The server applies natural language processing (NLP) techniques to summarize news articles, extracting key points from the articles, and then creates a concise summary based on the extracted information and stores it in a database.

[1515] Step 3: Generate explanatory text

[1516] 3.1. The server loads data for multiple commentary characters (e.g., character profiles and attributes).

[1517] 3.2. The server analyzes the news from each character's perspective.

[1518] 3.3. The server generates explanatory text from each character's perspective using a natural language generation model.

[1519] 3.4. The server stores the generated explanatory text in a database.

[1520] Specific behavior:

[1521] The server reads the attributes of the commentary characters (e.g., government officials, consumers, experts, etc.) and analyzes the news article from each character's perspective. Based on the analysis results, it uses natural language generation technology to create commentary text and records it in a database.

[1522] Step 4: Customization with Emotion Engine

[1523] 4.1. The server starts the emotion engine.

[1524] 4.2. The server acquires the user's emotion data.

[1525] 4.3. The server customizes the commentary text based on the emotion data.

[1526] 4.4. The server selects advertisements based on the emotion data.

[1527] Specific behavior:

[1528] The server activates the emotion engine to obtain the user's emotion data in real time, and uses facial expression recognition technology and voice tone analysis to identify the user's emotion, and then customizes the explanatory text and displayed advertisements accordingly.

[1529] Step 5: Generate an explainer video

[1530] 5.1. The server retrieves the customized description text and character information.

[1531] 5.2. The server launches a video generation engine (e.g., OpenCV, FFmpeg).

[1532] 5.3. The server generates the character's voice data and combines it with the animation.

[1533] 5.4. The server builds the explainer video and adds any necessary graphics and effects.

[1534] 5.5. The server saves the completed video file in storage.

[1535] Specific behavior:

[1536] The server converts the customized commentary text into audio data using speech synthesis technology, synchronizes it with the corresponding character animation, and also incorporates emotionally appropriate advertisements into the video, before saving the completed video file to storage.

[1537] Step 6: Prepare and stream your video

[1538] 6.1. The server accesses the management interface of the dedicated video site.

[1539] 6.2. The server uploads the generated video to the site.

[1540] 6.3. Configure the server to insert ads into the ad slots.

[1541] 6.4. The server sets the video to public and makes it available for viewing.

[1542] Specific behavior:

[1543] The server logs in to the video site and uploads the saved explanatory video. It then configures the site so that emotionally appropriate ads are displayed when the video is played, enabling free distribution under an advertising model. It also configures the site so that newly added videos are available for viewing when users access the site.

[1544] Step 7: User Watches Video

[1545] 7.1. The user accesses a dedicated video site.

[1546] 7.2. The user selects the news commentary video that interests them.

[1547] 7.3. The user plays the video and watches the news commentary.

[1548] Specific behavior:

[1549] Users access a video site and find videos they are interested in by category or from the latest news list. When they select a video and start playing it, an advertisement is displayed first, and then they can watch a news commentary video with content customized by the emotion engine.

[1550] The system allows users to receive news from multiple perspectives in an easy-to-understand format, and provides information and advertising tailored to the user's emotional state, enabling a more personalized experience.

[1551] Example 2

[1552] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1553] Conventional news distribution systems simply collect and distribute information, but are unable to respond to individual users' emotions and needs. This limits the news understanding and viewing experience, making it difficult to provide personalized information tailored to individual interests and emotions. Furthermore, there is a lack of efficient ways to summarize news or provide commentary from various perspectives, making it difficult to provide information in a way that is sufficiently useful to the recipient.

[1554] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1555] In this invention, the server includes means for acquiring news information, means for summarizing the acquired news information, means for generating text that uses multiple commentary characters to explain the summarized news information from various perspectives, means for recognizing a user's emotions using an emotion engine and customizing the commentary text and advertisements, means for generating videos using the character commentary texts, and means for delivering the generated videos. This makes it possible to provide personalized news information and customize information delivery and advertisements according to the user's emotions.

[1556] "News information" refers to information such as articles, headlines, publication dates, and authors obtained from various news sources.

[1557] A "summary" is a short sentence that extracts important points from news information and summarizes them in a concise form.

[1558] A "commentary character" is a fictional person or character whose role is to explain the news from different perspectives (e.g., government perspective, consumer perspective, expert perspective, etc.).

[1559] "Explanatory text" is a sentence generated by the commentary character to explain news information from various perspectives.

[1560] An "emotion engine" is a technology or system for recognizing a user's emotions and customizing explanatory text and advertisements based on those emotions.

[1561] The "means for generating video" refers to a technology that converts explanatory text into audio and then combines character animations and related graphics based on that audio to create a video.

[1562] "Means for distributing the generated video" refers to a technology or system for uploading the generated explanatory video to a video distribution platform on the Internet and making it available for public viewing.

[1563] MODE FOR CARRYING OUT THE INVENTION

[1564] The system for implementing the present invention includes a series of processing means for acquiring news information, summarizing it, generating videos in which multiple commentary characters provide commentary from various perspectives, and customizing the content of the videos and advertisements by recognizing the user's emotions. The processing flow of the server, terminal, and user is explained below with concrete examples.

[1565] Getting news information

[1566] The server retrieves news information from news sources. This retrieval is done using the news site's API (for example, Google News API). The server sends a request to the news source at a specific time (for example, every morning at 8:00) to retrieve the latest news information. The retrieved news is then stored in a database.

[1567] Examples:

[1568] The server accesses a news API on the Internet (e.g., Google News API) and receives the latest news headlines and article content.

[1569] News Summary

[1570] The server uses natural language processing technology (e.g., generative AI models such as BERT and GPT-3) to summarize the acquired news information. It extracts key points from the news article and summarizes them in a concise form. This summary is also stored in the database.

[1571] Examples:

[1572] The server analyzes the title, date, time, and text of the news article, extracts key keywords and key points, and generates a summary of about 100 characters. For example, a summary like "The world's first clean energy plant has begun operation, marking a new step forward in environmental protection" might be generated.

[1573] Generate explanatory text

[1574] The server reads multiple commentary characters and generates commentary text for the news based on each commentary character, thereby constructing commentary from various perspectives.

[1575] Examples:

[1576] The server generates text explaining the news from the perspective of character A (e.g., the government's perspective), and then generates explanatory text from the perspective of character B (e.g., the consumer's perspective).

[1577] Customization with Emotion Engine

[1578] The emotion engine recognizes the user's emotions. The user's emotion data is acquired in real time using the device's (PC or smartphone) camera and microphone. The emotion engine analyzes this data to determine the user's emotional state (e.g., joy, sadness, stress, etc.) and customizes explanatory text and advertisements accordingly.

[1579] Examples:

[1580] If the emotion engine detects that the user is feeling stressed from their facial expression, the explanatory text will be changed to something more relaxing, and an ad that will help relieve stress will be selected, such as an ad that says, "Take deep breaths while listening to relaxing music."

[1581] Generate explainer videos

[1582] The server generates a video based on the customized explanatory text. It uses text-to-speech technology (e.g., Google Text-to-Speech API) to convert the explanatory text into speech, and then combines the character animations and related graphics to create a video.

[1583] Examples:

[1584] The server uses an emotion engine to synthesize the customized commentary text, then combines the voice with character animation to generate a video file. Furthermore, advertisements selected to match the emotion are inserted into the video. For example, while a commentary character is explaining the news, an advertisement for relaxation goods is displayed on the side of the screen.

[1585] Preparing and streaming video

[1586] The server uploads the generated video to a dedicated video distribution platform (e.g., YouTube). When uploading, it sets up the insertion of advertisements selected by the emotion engine. After that, it makes the video available to users by setting it to public.

[1587] Examples:

[1588] The server uploads the explanatory videos to the video site, and the settings are configured to insert advertisements selected by the emotion engine. Users access the site and select and watch news explanatory videos that interest them.

[1589] User video viewing

[1590] Users access a dedicated video site from their PC, smartphone, or other device and select the news commentary video they want to watch. While the video is playing, customized commentary content and advertisements are displayed that are tailored to the user's emotional state. This allows users to receive information appropriate to their own emotions.

[1591] Examples:

[1592] Users access a video site and find videos of interest by category or from the latest news list. Once a video is selected and playback begins, the emotional engine displays customized content and advertisements, allowing users to watch news commentary. For example, a user feeling stressed can be provided with commentary in a relaxing atmosphere and related advertisements.

[1593] Prompt Sentence Examples

[1594] The following news information has been obtained from a news site. Please generate explanatory text from Character A (government perspective) and Character B (consumer perspective).

[1595] (News article)

[1596] Title: World's first clean energy plant begins operation

[1597] Body text: The world's first clean energy plant officially began operations yesterday. The plant uses renewable energy to significantly reduce carbon dioxide emissions compared to traditional power generation methods.

[1598] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1599] Step 1:

[1600] Getting news information

[1601] The server sends a request to the news source's API endpoint to retrieve the latest news information. Specific input includes parameters such as news category and region. The response from the API includes the news title, text, publication date, etc., and is stored in the database.

[1602] Input: API request parameters such as news category, region, etc.

[1603] Data processing: Extracting necessary information from API responses

[1604] Output: News information stored in a database

[1605] Step 2:

[1606] News Summary

[1607] The server analyzes the acquired news information and uses natural language processing technology to extract key points and generate summaries. Specifically, it extracts important keywords from the text of the news article and constructs a concise summary. The generated summary is then stored in a database.

[1608] Input: News information stored in a database

[1609] Data processing: Extract keywords from news articles and summarize key points

[1610] Output: Summary text stored in the database

[1611] Step 3:

[1612] Generate explanatory text

[1613] The server generates text that explains the summary of the acquired news information based on the perspective of the commentary character. Using a generative AI model, multiple explanatory texts suitable for each character are generated, enabling commentary from a variety of perspectives.

[1614] Input: Summary text stored in the database, character information for commentary

[1615] Data processing: Generate explanatory text using a generative AI model

[1616] Output: Descriptive text stored in the database

[1617] Step 4:

[1618] Customization with Emotion Engine

[1619] The device collects the user's emotional data (facial expressions and voice) in real time, which is then analyzed by the emotion engine. Based on the analysis results, explanatory text and displayed advertisements are customized.

[1620] Input: User's facial expression data, voice data

[1621] Data processing: Data analysis using sentiment analysis algorithms

[1622] Output: Customized explanatory text and advertising information

[1623] Step 5:

[1624] Generate explainer videos

[1625] The server generates a video based on the customized explanatory text, converts the explanatory text into audio using text-to-speech synthesis technology, and combines the audio with character animation and related graphics to create a video file.

[1626] Input: Custom description text, character animation data

[1627] Data processing: text-to-speech synthesis, video editing

[1628] Output: Generated video file

[1629] Step 6:

[1630] Preparing and streaming video

[1631] The server uploads the generated explanatory video to a video distribution platform. After uploading, advertisements selected by the emotion engine are inserted into the video, and the video is made available to users by setting it to public.

[1632] Input: Generated video file, ad data

[1633] Data processing: video upload, ad insertion, publishing settings

[1634] Output: Published instructional video

[1635] Step 7:

[1636] User video viewing

[1637] Users access a dedicated video site using their device and select the news commentary video they want to watch. When the video is played, customized commentary content and advertisements are displayed, allowing users to receive information tailored to their emotions.

[1638] Input: URL of video site, user selection information

[1639] Data processing: User interface operation, video playback

[1640] Output: A customized explainer video to be watched

[1641] (Application example 2)

[1642] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1643] In modern society, the Internet allows users to quickly access a wide variety of news, but it can be difficult for users to find information that is relevant to them from the vast amount of information available. Furthermore, news commentary tends to be biased toward a one-sided perspective, making it difficult for users to understand the content from multiple angles. Furthermore, advertisements included in news and commentary videos are often not relevant to users' interests or emotions, resulting in reduced advertising effectiveness. There is a need for a system that can solve these issues and enable users to receive more personalized news commentary and advertisements.

[1644] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1645] In this invention, the server includes means for acquiring news information, means for summarizing the acquired news information, means for generating text that explains the summarized news information from various perspectives using multiple commentary characters, means for recognizing a user's emotions and customizing the commentary text and advertisements based on the emotions, and means for delivering the generated videos. This not only enables users to receive news from multiple perspectives in an easy-to-understand format, but also enables a more personalized experience by providing information and advertisements that are appropriate for their emotional state.

[1646] "News information" refers to information about the latest happenings and events that is published on the Internet.

[1647] "Means of Acquisition" refers to the technology or process used to collect news information from specific websites or APIs.

[1648] The "summarization method" is a natural language processing technique that extracts the main points of acquired news information and summarizes them in a concise form.

[1649] A "commentary character" is a virtual character with specific characteristics and personality that is used to explain news information from multiple perspectives.

[1650] The "means for generating explanatory text" is a technology for creating text to explain news information from different perspectives for each commentary character.

[1651] "Emotion recognition and customization" refers to technology that analyzes users' emotions in real time and adjusts explanatory text and advertising content accordingly.

[1652] "Means for generating video" refers to a technology that creates a video file that combines audio, animation, and graphics based on explanatory text and character information.

[1653] "Means of distribution" refers to the technology and process for providing the generated video to users in a viewable format via the Internet, etc.

[1654] "Natural language processing technology" refers to technology that allows computers to analyze, understand, and generate human language.

[1655] The "database" is a digital information management system for organizing and storing acquired news information and generated commentary text.

[1656] "Advertisement" refers to information used to notify or solicit specific products or services to users.

[1657] In the system for implementing the present invention, the server, terminal, and user parts operate in cooperation with each other.

[1658] First, the server retrieves news information from a news API on the Internet. This is done using an HTTP request, and the retrieved news data is stored in a database. Next, the server summarizes the news information using natural language processing technology. The summarized news is concisely summarized to around 100 characters and stored again in the database.

[1659] The server generates commentary text based on the summarized news information from the perspective of each commentator. The commentary text is generated from various perspectives, such as the government perspective, expert perspective, and consumer perspective.

[1660] The device is then equipped with a camera and microphone, which capture the user's facial expressions and tone of voice in real time. The emotion engine analyzes this data to recognize the user's emotional state, and customises explanatory text and advertisements based on the results.

[1661] The server generates an explanatory video using the customized explanatory text and advertisements. The video generation combines character voice-overs, animations, and related graphics. The generated video includes advertisements customized according to the user's emotions.

[1662] Finally, the generated video is uploaded to a dedicated video site, where users can access and watch it. The video is displayed with content and advertisements optimized for the user's emotions, providing a more personalized experience for users.

[1663] The hardware required is a high-performance computer for processing on the server, and a device equipped with a camera and microphone for recognizing user emotions. The software used is Python and natural language processing libraries (e.g., Transformers) for acquiring and summarizing news information and generating explanatory text. The emotion engine performs voice and facial expression analysis, and a video distribution platform is used to distribute the generated explanatory videos.

[1664] As a concrete example, the server retrieves news information from a news API every morning at 8 a.m. and generates a summary using natural language processing technology. Multiple commentary characters then generate commentary text from different perspectives, and the system recognizes the user's emotions through the device's camera and microphone, customizing the commentary text and advertisements. For example, if the user is feeling stressed, it will display an advertisement for a relaxing hot spring, and if they are in a happy mood, it will display an advertisement for the latest gadgets.

[1665] An example prompt for a generative AI model might look like this:

[1666] Generate a news summary: "Summarize the following news article in 100 characters or less: [news article text]"

[1667] Sentiment Analysis: "Determine the user's current emotion based on speech and facial expression analysis."

[1668] This allows users to enjoyably understand the news from multiple perspectives, and also allows them to receive appropriate advertisements that match their emotions.

[1669] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1670] Step 1:

[1671] The server sends a request to the news API to retrieve the latest news information. The input includes the news API endpoint and request parameters. The retrieved news data is received in JSON format and stored in a database. Specifically, the HTTP request is executed using the Python Requests library.

[1672] Step 2:

[1673] The server extracts the text of news information retrieved from the database and summarizes it using natural language processing technology. The input includes the text of the retrieved news article. The summary results are then saved back into the database. Specifically, the Transformers library is used to extract key points from the news article and summarize it to about 100 characters.

[1674] Step 3:

[1675] The server generates commentary text from the perspectives of multiple commentary characters. The input includes summarized news information and each character's perspective information. As output, commentary text from each perspective is generated and stored in a database. Specifically, customized commentary text is created using a template corresponding to each character's perspective.

[1676] Step 4:

[1677] The device uses a camera and microphone to capture the user's facial expressions and voice tone. Input includes real-time video and audio data. This data is sent to an emotion engine that analyzes the user's emotional state. This can be done using the Emotion API or similar emotion analysis tools.

[1678] Step 5:

[1679] The emotion engine customizes explanatory text and advertisements based on the user's emotional state. The input includes the emotional data obtained in step 4. The output is customized explanatory text and advertisements appropriate for the emotion. Specifically, if the user is feeling stressed, the content is changed to something that will help them relax, and an appropriate advertisement is selected.

[1680] Step 6:

[1681] The server generates a video by combining character voice-overs and animation based on customized explanatory text and advertisements. The input includes customized explanatory text and advertisement information. The output is the generated video file. Specifically, the system converts text into speech using a TTS (Text-to-Speech) engine and generates a video using an animation tool.

[1682] Step 7:

[1683] The server uploads the generated video to a dedicated video site and prepares it for distribution so that users can access it. The input includes the generated video file. The output is published on the video site and available for users to view. Specifically, the video file is uploaded to the server using an API or FTP and the publishing settings are configured.

[1684] In this way, a series of processes are realized, from acquiring news information to summarizing it, generating explanatory text, customizing it based on the user's emotions, generating videos, and finally delivering it.

[1685] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1686] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1687] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1688] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1689] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1690] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1691] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1692] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1693] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1694] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1695] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1696] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1697] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1698] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1699] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1700] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1701] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1702] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1703] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1704] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1705] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1706] The following is further disclosed regarding the above embodiment.

[1707] (Claim 1)

[1708] a means for obtaining news information;

[1709] means for summarizing the retrieved news information;

[1710] A means for generating text that explains summarized news information from various perspectives using multiple commentary characters;

[1711] A means for generating a video using explanatory text of a character;

[1712] The system includes a means for delivering the generated video.

[1713] (Claim 2)

[1714] 10. The system of claim 1, further comprising means for automatically obtaining news information at a specific time and storing it in a database.

[1715] (Claim 3)

[1716] 10. The system of claim 1, further comprising means for generating a summary from the news information using natural language processing techniques.

[1717] (Claim 4)

[1718] 10. The system of claim 1, further comprising means for inserting advertisements into the generated video and distributing it free of charge under an advertising model.

[1719] (Claim 5)

[1720] 10. The system of claim 1, further comprising means for uploading the video to a dedicated video site accessible to the user.

[1721] (Claim 6)

[1722] 2. The system according to claim 1, further comprising means for providing commentary from a variety of perspectives using profiles and attributes of a plurality of commentary characters.

[1723] (Claim 7)

[1724] 10. The system of claim 1, further comprising means for generating video using the video generation engine, the video including voice-over and animation of the character.

[1725] (Claim 8)

[1726] 10. The system of claim 1, further comprising means for organizing the news commentary videos into categories to allow a viewer to easily find videos of interest.

[1727] "Example 1"

[1728] (Claim 1)

[1729] a means for obtaining news information;

[1730] means for summarizing the retrieved news information;

[1731] A means for generating text that explains summarized news information from various perspectives using multiple commentary characters;

[1732] A means for generating a video using explanatory text of a character;

[1733] The system includes means for delivering the generated video.

[1734] (Claim 2)

[1735] 10. The system of claim 1, further comprising means for automatically obtaining news information at a specific time and storing it in a database.

[1736] (Claim 3)

[1737] 10. The system of claim 1, further comprising means for generating a summary from the news information using natural language processing techniques.

[1738] "Application Example 1"

[1739] (Claim 1)

[1740] a means for obtaining news information;

[1741] means for summarizing the retrieved news information;

[1742] A means for generating text that explains summarized news information from various perspectives using multiple commentary characters;

[1743] A means for generating a video using explanatory text of a character;

[1744] a means for distributing the generated video;

[1745] A means to play news commentary videos on smart devices,

[1746] a means for generating explanatory text using a generative AI model;

[1747] A system including means for using a prompt sentence as input to a generative AI model.

[1748] (Claim 2)

[1749] 10. The system of claim 1, further comprising means for automatically obtaining news information at a specific time and storing it in a database.

[1750] (Claim 3)

[1751] 10. The system of claim 1, further comprising means for generating a summary from the news information using natural language processing techniques.

[1752] "Example 2: Combining Emotion Engines"

[1753] (Claim 1)

[1754] a means for obtaining news information;

[1755] means for summarizing the retrieved news information;

[1756] A means for generating text that explains summarized news information from various perspectives using multiple commentary characters;

[1757] A means for recognizing user emotions using an emotion engine and customizing explanatory text and advertisements;

[1758] A means for generating a video using explanatory text of a character;

[1759] The system includes a means for delivering the generated video.

[1760] (Claim 2)

[1761] 10. The system of claim 1, further comprising means for automatically obtaining news information at a specific time and storing it in a database.

[1762] (Claim 3)

[1763] 10. The system of claim 1, further comprising means for generating a summary from the news information using natural language processing techniques.

[1764] "Application example 2 when combining emotion engines"

[1765] (Claim 1)

[1766] a means for obtaining news information;

[1767] means for summarizing the retrieved news information;

[1768] A means for generating text that explains summarized news information from various perspectives using multiple commentary characters;

[1769] A means for generating a video using explanatory text of a character;

[1770] a means for recognizing user emotions and customizing explanatory text and advertisements accordingly;

[1771] The system includes a means for delivering the generated video.

[1772] (Claim 2)

[1773] 10. The system of claim 1, further comprising means for automatically obtaining news information at a specific time and storing it in a database.

[1774] (Claim 3)

[1775] 10. The system of claim 1, further comprising means for generating a summary from the news information using natural language processing techniques. [Explanation of symbols]

[1776] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. a means for obtaining news information; means for summarizing the retrieved news information; A means for generating text that explains summarized news information from various perspectives using multiple commentary characters; A means for generating a video using explanatory text of a character; The system includes a means for delivering the generated video.

2. 2. The system of claim 1, further comprising means for automatically obtaining news information at a specific time and storing the information in a database.

3. The system of claim 1 further comprising means for generating summaries from the news information using natural language processing techniques.

4. The system according to claim 1, further comprising means for inserting advertisements into the generated video and distributing it free of charge under an advertising model.

5. 10. The system of claim 1, further comprising means for uploading the video to a dedicated video site accessible to the user.

6. The system according to claim 1, further comprising means for providing commentary from a variety of viewpoints using profiles and attributes of a plurality of commentary characters.

7. The system of claim 1 further comprising means for generating animations including character text-to-speech and animations using an animation generation engine.

8. 10. The system of claim 1, further comprising means for organizing the news commentary videos by category to allow a viewer to easily find a video of interest.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A