System

The system converts news articles into four-panel comics and videos using AI to address the digital native generation's preference for visual content, engaging them effectively and efficiently.

JP2026019217APending Publication Date: 2026-02-05SOFTBANK GROUP CORP

Patent Information

Application Number
JP2024120626
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-25
Publication Date
2026-02-05

AI Technical Summary

Technical Problem

The digital native generation, particularly teens to early 30s, shows little interest in text-based news content and prefers visual, intuitive content, making traditional news sites unattractive. Providing visual news content is time-consuming and costly.

Method used

A system that converts news articles into visually intuitive four-panel comics and videos using generative AI and image generation AI, extracting key points, generating scenarios, and sharing them on video-sharing platforms to promote access.

Benefits of technology

Efficiently generates visual and intuitive news content at low cost, effectively engaging the digital native generation and promoting access to news sites.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026019217000001_ABST
    Figure 2026019217000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: Means for acquiring a news article, means for analyzing the acquired news article and extracting an important keyword or point, means for generating a scenario of a four frame comic using a generation AI based on the extracted keyword or point, means for generating the four frame comic using an image generation AI based on the generated scenario, means for generating an animation based on the generated four frame comic and creating a four frame moving image, means for displaying the generated four frame comic and four frame moving image, and means for posting the generated four frame moving image to a moving image sharing platform or a social networking service, means for facilitating access to the original article.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] In recent years, the digital native generation, especially young people in their teens to early 30s, have become increasingly accustomed to using smartphones and social networking services on a daily basis, primarily gathering information from these platforms. However, this demographic shows little interest in text-based news content and tends to prefer visual, intuitive content, making traditional news sites unattractive to them. Another problem is the time and cost required to provide visual news content to this demographic. The present invention aims to solve these issues and provide an efficient method for attracting the interest of young people and promoting the use of news sites. [Means for solving the problem]

[0005] This invention provides a means for acquiring and analyzing news articles and extracting important keywords and key points. It also provides a means for generating a four-panel comic scenario using a generation AI based on the extracted keywords and key points. It also provides a means for generating a four-panel comic based on the generated scenario using an image generation AI, generating animation based on the generated four-panel comic, and creating a four-panel video. It proposes a system that includes a means for displaying the generated four-panel comic and four-panel video, and a means for posting the generated four-panel video on a video-sharing platform or social networking service to promote access to the original article. This makes it possible to efficiently generate visual and intuitive news content at low cost and effectively provide information to the digital native generation.

[0006] A "news article" is text-based information content provided by an online news site or other information platform.

[0007] "Analysis" is the process of examining data or information in detail and extracting its components and key points.

[0008] "Keywords" are key words or phrases that represent the content of a news article and play an important role in understanding the essence of the article.

[0009] "Points" are the key facts or pieces of information in a news story that form the heart of the story.

[0010] "Generative AI" refers to software or algorithms that use machine learning and artificial intelligence techniques to generate new content.

[0011] "Four-panel comics" are a comic format that expresses a series of stories or concepts in four consecutive frames.

[0012] A "scenario" is an explanatory text or plot that defines the situation and storyline in each panel of a four-panel comic.

[0013] "Image generation AI" refers to software or algorithms that use artificial intelligence technology to generate illustrations or images based on text descriptions.

[0014] "Animation" is a visual expression technique that conveys movement by displaying multiple still images in succession.

[0015] A "four-panel video" is a short video that animates each frame of a four-panel comic.

[0016] A "video sharing platform" is an online platform that enables users to upload, share and view videos.

[0017] A "social networking service" is an online service that allows users to interact with each other and share information. [Brief explanation of the drawings]

[0018] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7]FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0019] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0020] First, the terms used in the following description will be explained.

[0021] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0022] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0023] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0024] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0025] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0026] [First embodiment]

[0027] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0028] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0029] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0030] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0031] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0032] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0033] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0034] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0035] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0036] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0037] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0038] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0039] overview

[0040] This invention is a system that converts news articles into visually intuitive four-panel comics and four-panel videos. This system utilizes generative AI and image generation AI technologies to extract the key points of news articles and present them in an easy-to-understand format for users.

[0041] System Configuration

[0042] The system of the present invention consists of the following main parts:

[0043] 1. News article acquisition module

[0044] 2. Text Analysis Module

[0045] 3. Four-panel comic generation module

[0046] 4. 4-frame video generation module

[0047] 5. Content Display and Sharing Module

[0048] News article acquisition module

[0049] The server retrieves the latest news articles through the news site's API or RSS feed. This retrieval operation can be set to occur periodically, and the latest articles are automatically imported into the system. The retrieved news articles are stored in a database.

[0050] Examples:

[0051] The server accesses the "News API" to retrieve the title, body text, and URL of the latest news article, then stores this information in a database.

[0052] Text Analysis Module

[0053] The server uses natural language processing techniques to analyze the text of the retrieved news articles and extract key keywords and points, a process that clarifies the core of the article.

[0054] Examples:

[0055] The server analyzes the news article "Impacts of Climate Change" and extracts the main keywords "climate change," "abnormal weather," and "impact."

[0056] Four-panel comic generation module

[0057] The server uses a generation AI to create a four-panel comic scenario based on the extracted keywords and points. This scenario includes the situation and storyline of each panel. The generated scenario is then passed to an image generation AI, which generates illustrations for each panel.

[0058] Examples:

[0059] The server creates the scenario as follows:

[0060] 1. Frame 1: A character is thinking about "climate change."

[0061] 2. Frame 2: A scene where abnormal weather is occurring

[0062] 3. Panel 3: People are surprised by the abnormal weather.

[0063] 4. Panel 4: A scene of characters thinking about climate change countermeasures

[0064] Based on this, the image generation AI generates four illustrations.

[0065] 4-frame video generation module

[0066] The server then creates an animation based on the four-panel comic, generating a four-panel video. During this process, the still images in each frame are animated and presented as a continuous video.

[0067] Examples:

[0068] The server breaks down a four-panel comic and animates each scene. For example, it animates the character's thoughts in the first panel, then moves to show the occurrence of extreme weather.

[0069] Content Viewing and Sharing Module

[0070] The server displays the generated four-panel comics and videos on the homepage of the news site, and also posts these videos on video sharing platforms and social networking services to promote access to the news articles.

[0071] Examples:

[0072] The server places the generated four-panel comic on the homepage of the news site, and then uploads the four-panel video to YouTube and Twitter, with a link to the article.

[0073] Example

[0074] When users visit a news site, they can click on a thumbnail of a four-panel comic that appears on the homepage to read a detailed news article. Users who are interested in a four-panel comic that has been shared on a video sharing platform or social networking service can click on the link to access the original article.

[0075] example:

[0076] A user clicks on a four-panel cartoon on the homepage of a news site to read a detailed article about the impacts of climate change. They also watch a four-panel video on YouTube and access the news site via a link in the video description.

[0077] As described above, the system of the present invention converts news articles into a visual and intuitive format, attracting the interest of the digital native generation and promoting access to news sites.

[0078] The processing flow will be explained below.

[0079] Step 1:

[0080] The server retrieves the latest news articles using the news site's API or RSS feed. It periodically sends requests using the API key to collect news data. The retrieved news articles are then stored in a database.

[0081] Specific behavior:

[0082] The server accesses the "News API" and sends a request to retrieve the latest news articles.

[0083] The title, text, and URL of the news article received as a response are extracted and stored in a database.

[0084] Step 2:

[0085] The server invokes a natural language processing (NLP) engine to analyze the text of the stored news articles, thereby extracting important keywords and key points.

[0086] Specific behavior:

[0087] The server inputs the text of the news article into the NLP engine.

[0088] The NLP engine analyzes the text and extracts key keywords (e.g., "climate change," "extreme weather," "impact") and key points.

[0089] The extracted keywords and points are stored in a database.

[0090] Step 3:

[0091] The server uses generative AI to generate a four-panel comic scenario based on the extracted keywords and points. The generated scenario defines the situation and storyline of each panel.

[0092] Specific behavior:

[0093] The server inputs keywords and points into the generation AI and requests it to generate a four-panel scenario.

[0094] The generation AI creates a scenario and returns a description of each frame (e.g., "Frame 1: A character thinks about climate change," "Frame 2: Extreme weather occurs," etc.).

[0095] The generated scenario is saved in the database.

[0096] Step 4:

[0097] The server uses image generation AI to create a four-panel comic based on the generated scenario. The image generation AI generates illustrations based on the text description of each panel.

[0098] Specific behavior:

[0099] The server inputs the scenario text into the image generation AI and requests it to generate an illustration.

[0100] The image generation AI generates an illustration corresponding to each frame and returns it to the server.

[0101] The generated illustrations are arranged in the form of a four-panel comic.

[0102] Step 5:

[0103] The server generates animations based on each frame of the four-panel comic, creating a four-panel video. It displays still images in succession and adds movement to each scene.

[0104] Specific behavior:

[0105] The server provides the AI ​​with a script to animate the illustrations in each panel of the four-panel comic.

[0106] The generative AI generates the movement of each frame and creates a video file as an animation.

[0107] The generated video file is saved on the server.

[0108] Step 6:

[0109] The server places the generated four-panel comics and videos on the front page of the news site, and also posts them on video sharing platforms and social networking services to promote access.

[0110] Specific behavior:

[0111] The server generates HTML code to embed the four-panel comic on the top page of the news site.

[0112] The generated four-panel comic is displayed on the top page.

[0113] Upload your four-frame video to platforms such as YouTube and Twitter.

[0114] Link to the news article in the video description.

[0115] Step 7:

[0116] Users can click on thumbnails of four-panel comics or videos displayed on the homepage of news sites to read detailed news articles, or watch videos on social media or video-sharing platforms and access the original articles.

[0117] Specific behavior:

[0118] A user visits the homepage of a news site and clicks on a thumbnail of a four-panel comic strip.

[0119] Clicking on the thumbnail will take you to the news article details page.

[0120] Users watch four-frame videos on YouTube or Twitter and click on a link in the video description to access a news site.

[0121] The above are the specific processing steps of the present invention. By executing these steps, it is possible to provide visual and intuitive news content to the digital native generation and promote access to news sites.

[0122] Example 1

[0123] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0124] Traditional news content is often expressed in text format, making it difficult to provide a visual and intuitive understanding, especially for the digital native generation. Furthermore, the process of visualizing and animating news articles is time-consuming, and automation is needed.

[0125] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0126] In this invention, the server includes a means for acquiring news content, a means for analyzing the acquired news content and extracting important keywords and information, and a means for generating a scenario for story-style content using a generative AI model based on the extracted keywords and information, thereby automating the task of visually and intuitively presenting news articles.

[0127] "News content" refers to information such as articles and reports obtained from news sites and other information platforms.

[0128] "Important keywords and information" refers to content and key points that are particularly noteworthy extracted from news content.

[0129] A "generative AI model" refers to an artificial intelligence model that understands natural language and generates new sentences and scenarios based on that information.

[0130] "Scenario for narrative content" refers to a storyline for constructing a continuous narrative, such as a four-panel comic strip, generated based on keywords and information extracted from news content.

[0131] An "image generation model" refers to an artificial intelligence model for generating images based on text information.

[0132] "Visual content" refers to visually presented information, such as four-panel comic strips, created using generative AI models and image generation models.

[0133] "Animation" refers to moving image content created by sequentially displaying visual content and adding movement.

[0134] "Video Content" refers to continuous, moving images generated from visual content.

[0135] A "video sharing platform" refers to a service that allows users to share videos over the Internet, such as YouTube.

[0136] "Social media" refers to online platforms such as Twitter and Facebook that allow users to communicate with each other and share content.

[0137] overview

[0138] This invention is a system that converts news content into visually intuitive four-panel comics and four-panel videos. This system utilizes generative AI and image generation model technologies to extract key points from news content and present them in an easy-to-understand format for users.

[0139] System Configuration

[0140] The system of the present invention consists of the following main modules:

[0141] 1. News article acquisition module

[0142] 2. Text Analysis Module

[0143] 3. Four-panel comic generation module

[0144] 4. 4-frame video generation module

[0145] 5. Content Display and Sharing Module

[0146] News article acquisition module

[0147] The server retrieves the latest news content using the news site's API or RSS feed. This retrieval operation can be performed periodically, and the latest news content is automatically imported into the system. The retrieved content is then stored in a database.

[0148] Examples:

[0149] The server accesses the "News API" to retrieve the title, body text, and URL of the latest news article, then stores this information in a database.

[0150] Text Analysis Module

[0151] The server uses natural language processing techniques to analyze the text of the retrieved news content and extract key keywords and information, thereby clarifying the core of the article.

[0152] Examples:

[0153] The server analyzes the news article "Impacts of Climate Change" and extracts the main keywords "climate change," "abnormal weather," and "impact."

[0154] Four-panel comic generation module

[0155] The server uses a generative AI model to create a four-panel comic scenario based on the extracted keywords and information. This scenario includes the situation and storyline of each panel. The generated scenario is then passed to an image generation model, which generates illustrations for each panel.

[0156] Examples:

[0157] The server creates the scenario as follows:

[0158] 1. Frame 1: A character is thinking about "climate change."

[0159] 2. Frame 2: A scene where abnormal weather is occurring

[0160] 3. Panel 3: People are surprised by the abnormal weather.

[0161] 4. Panel 4: A scene of characters thinking about climate change countermeasures

[0162] Based on this, an image generation model generates four illustrations.

[0163] 4-frame video generation module

[0164] The server then creates an animation based on the four-panel comic, generating a four-panel video. During this process, the still images in each frame are animated and presented as a continuous video.

[0165] Examples:

[0166] The server breaks down a four-panel comic and animates each scene. For example, it animates the character's thoughts in the first panel, then moves to show the occurrence of extreme weather.

[0167] Content Viewing and Sharing Module

[0168] The server displays the generated four-panel comics and videos on the homepage of the news site, and also posts these videos on video sharing platforms and social media to promote access to the news content.

[0169] Examples:

[0170] The server places the generated four-panel comic on the homepage of the news site, and also uploads the four-panel video to YouTube and Twitter, with a link to the article.

[0171] Examples of prompt statements

[0172] Below are some example prompts for using generative AI models:

[0173] Example prompt sentence:

[0174] 1. Prompts to extract necessary keywords from news articles:

[0175] Extract key keywords from the news article "Impacts of Climate Change."

[0176] 2. Prompts for generating a four-panel comic scenario:

[0177] Please create a four-panel comic scenario based on the extracted keywords "climate change," "extreme weather," and "impact." Please be specific and include a storyline for each panel.

[0178] 3. Prompt to pass to the image generation model:

[0179] Please generate an illustration for each frame based on the following scenario.

[0180] 1. Frame 1: A character is thinking about "climate change."

[0181] 2. Frame 2: A scene where abnormal weather is occurring

[0182] 3. Panel 3: People are surprised by the abnormal weather.

[0183] 4. Panel 4: A scene of characters thinking about climate change countermeasures

[0184] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0185] Step 1:

[0186] The server retrieves news content. It uses the news site's API or RSS feed as input. The output is news article data (title, body text, URL, etc.), which is stored in a database.

[0187] Specific behavior:

[0188] The server sends a request to the "News API" to receive the latest news article data. The received data (e.g., article data on the impact of climate change) is stored in the corresponding table in the database.

[0189] Step 2:

[0190] The server reads news content from the database and performs text analysis. The input is the text data of the news article obtained in step 1, and the output is a list of important keywords and information.

[0191] Specific behavior:

[0192] The server uses a natural language processing library (e.g., SpaCy or NLTK) to analyze the text of a news article (e.g., "impacts of climate change") and extracts the main keywords "climate change," "extreme weather," and "impact."

[0193] Step 3:

[0194] The server uses a generative AI model based on the extracted keywords and information to create a four-panel comic scenario. The input is the keywords and information extracted in step 2, and the output is the four-panel comic scenario.

[0195] Specific behavior:

[0196] The server sends a prompt to the generation AI model, for example, "Create a four-panel comic scenario based on the following keywords: Keywords: climate change, extreme weather, impact," and generates a scenario.

[0197] Example scenario:

[0198] 1. Frame 1: A character is thinking about "climate change."

[0199] 2. Frame 2: A scene where abnormal weather is occurring

[0200] 3. Panel 3: People are surprised by the abnormal weather.

[0201] 4. Panel 4: A scene of characters thinking about climate change countermeasures

[0202] Step 4:

[0203] Based on the generated scenario, the server uses an image generation model to generate illustrations for each frame of the four-panel comic. The input is the scenario created in step 3, and the output is a visual illustration for each frame.

[0204] Specific behavior:

[0205] The server sends prompts to the image generation model (e.g., DALL-E). For example, it uses the image generation prompt for the first frame, "A scene in which a character is thinking about climate change," to generate illustrations for each frame.

[0206] Step 5:

[0207] The server creates an animation based on the generated four-panel comic and generates a four-panel video. The input is the visual illustration generated in step 4, and the output is a four-panel video.

[0208] Specific behavior:

[0209] The server uses video editing software or libraries (e.g., OpenCV) to animate the images in each frame. For example, you could animate a character thinking in the first frame, then smoothly transition to an extreme weather scene.

[0210] Step 6:

[0211] The server displays the generated comic strips and videos on the homepage of the news site, and also posts these videos to video sharing platforms and social media. The input is the comic strips and videos generated in step 5, and the output is the web page where the information is displayed and the posted content.

[0212] Specific behavior:

[0213] The server places thumbnails of the four-panel comics on the homepage of the news site, and when users click on them, detailed news content is displayed. The server also uploads the four-panel comics to YouTube and Twitter, automatically linking them to the articles.

[0214] (Application example 1)

[0215] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0216] News articles are primarily text-based, making it difficult to understand their content quickly. Furthermore, digital natives, in particular, are demanding content in a visual, intuitive format. To meet these modern needs, a new way of presenting news articles in a visual, easy-to-understand format is needed.

[0217] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0218] In this invention, the server includes means for acquiring news articles, means for analyzing the acquired news articles and extracting important keywords and key points, means for generating a four-panel cartoon scenario using a generation AI based on the extracted keywords and key points, means for generating a four-panel cartoon using an image generation AI based on the generated scenario, means for generating animation based on the generated four-panel cartoon and creating a four-panel video, means for displaying the generated four-panel cartoon and four-panel video, means for posting the generated four-panel video on a video sharing platform or social networking service to promote access to the original article, and means for displaying and sharing the generated four-panel cartoon and four-panel video on a smart device, thereby enabling users to visually understand the content of the news article in a short amount of time.

[0219] A "means for acquiring news articles" is a system or module for automatically acquiring the latest news articles from news sites or RSS feeds on the Internet.

[0220] "Means for analyzing acquired news articles and extracting important keywords and key points" refers to a system or module that has the function of analyzing the text of news articles using natural language processing technology and extracting the subject matter and important information of the articles.

[0221] "Means for generating a four-panel comic scenario using generative AI" refers to a system or module that uses generative AI to construct a storyline for a four-panel comic and the content of each panel based on keywords and key points extracted from news articles.

[0222] "Means for generating four-panel comics using image generation AI" refers to a system or module that uses image generation AI to create illustrations for four-panel comics based on a generated scenario.

[0223] "Means for generating animation based on the generated four-panel comic and creating a four-panel video" refers to a system or module that animates the still images of each frame and presents them as a continuous four-panel video.

[0224] "Means for displaying the generated four-panel comics and four-panel videos" refers to a system or module for displaying the generated four-panel comics and four-panel videos on a user interface.

[0225] "Means for posting the generated four-panel videos on video sharing platforms or social networking services to promote access to the original article" refers to a system or module that uploads the generated four-panel videos to video sharing platforms or social networking services and provides users with a link that allows them to access the original news article.

[0226] "Means for displaying and sharing generated four-panel comics and four-panel videos on a smart device" refers to a system or module that allows users to display four-panel comics and four-panel videos generated on a smart device such as a smartphone or tablet, and share the content with other users via a link or share button.

[0227] "Natural language processing technology" is a computer technology for analyzing text data and extracting or summarizing important information.

[0228] A "content distribution service" is an online service that provides users with various types of digital content via the Internet.

[0229] Overall system overview

[0230] This system aims to provide news articles as visually intuitive four-panel comics and four-panel videos. Specifically, it supports processes including news article acquisition, analysis, scenario creation using generative AI, comic generation using image generation AI, animation, display, and sharing. The main components of this system are a news article acquisition module, text analysis module, four-panel comic generation module, four-panel video generation module, and content display and sharing module.

[0231] Hardware and Software

[0232] The hardware used includes servers and smart devices (e.g., smartphones). The main software includes:

[0233] Python 3.x: Programming Language

[0234] Requests: A library for making HTTP requests

[0235] BeautifulSoup: HTML parsing library

[0236] Spacy: A natural language processing library

[0237] Transformers: A library for using generative AI

[0238] Image generation API (hypothetical API)

[0239] Processing flow

[0240] The server first retrieves news articles from news sites and RSS feeds. Next, it analyzes the retrieved news articles and extracts important keywords and key points using natural language processing technology. Based on the extracted keywords and key points, it uses generative AI to generate a four-panel comic scenario. It then uses image generation AI to create four-panel comic illustrations based on the scenario. It then animates the generated four-panel comic to create a four-panel video. Finally, the generated four-panel comic and video can be displayed and shared on a smart device.

[0241] Specific examples

[0242] As a concrete example, we take the news article "The spread of the new coronavirus" and extract the key keywords "new coronavirus," "social distancing," and "vaccine." The following scenario is generated:

[0243] 1. A scene where a character is thinking about the new coronavirus

[0244] 2. Scenes where social distancing is being implemented

[0245] 3. Scenes of people maintaining social distancing

[0246] 4. A scene where a character thinks about getting vaccinated

[0247] An example of a prompt passed to the generative AI model is:

[0248] Prompt: Transform the contents of a news article into a four-panel comic book scenario.

[0249] Key Word 1: Novel Coronavirus

[0250] Key word 2: Social distancing

[0251] Key word 3: Vaccine

[0252] The generated four-panel comics and videos can be displayed on smart devices such as smartphones and can be easily shared via social media, allowing users to quickly and visually understand the content of news articles.

[0253] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0254] Step 1:

[0255] The server periodically retrieves news articles from news sites and RSS feeds on the Internet. Specifically, the server accesses a news API, retrieves the title, body, and URL of the latest news article, and stores them in a database. The URL of the news API is provided as input, and the JSON data of the news article is obtained as output.

[0256] Step 2:

[0257] The server analyzes the retrieved news articles and extracts important keywords and key points. In this process, it uses natural language processing technology (Spacy) to analyze the article text and extract key keywords and noun phrases. The body of the news article is provided as input, and a list of important keywords is obtained as output.

[0258] Step 3:

[0259] The server uses a generative AI model to generate a four-panel comic scenario based on the extracted keywords and key points. Specifically, a prompt is input to the generative AI, which outputs four scenario blocks. The extracted keywords and prompt are provided as input, and the scenario for each panel is generated as output.

[0260] Step 4:

[0261] The server uses image generation AI to generate a four-panel comic based on the generated scenario. The generated scenario is passed to an image generation API, which generates illustrations for each panel. The scenario is provided as input, and four illustrations are generated as output.

[0262] Step 5:

[0263] The server generates an animation based on the generated four-panel comic, creating a four-panel video. The illustrations for each frame are imported into animation software (e.g., Adobe After Effects) to create a continuous animation. Four illustrations are provided as input, and a four-panel video is generated as output.

[0264] Step 6:

[0265] The server displays the generated four-panel comic and four-panel video. In this step, the generated content is displayed on a user interface so that the user can view it. The four-panel comic and four-panel video are provided as input, and are displayed on the user interface as output.

[0266] Step 7:

[0267] The server posts the generated four-frame comics to video sharing platforms and social networking services to promote access to the original article. Specifically, it uploads the videos using APIs such as YouTube and Twitter and sets up links to the articles. The four-frame comics and the URLs of the original articles are provided as input, and the videos are published on the video platforms as output.

[0268] Step 8:

[0269] The server displays and shares the generated four-panel comics and four-panel videos on smart devices. Users view the generated content on their smartphones or tablets and click share links on social media to share the content with other users. Four-panel comics and four-panel videos are provided as input, and can be displayed and shared on smart devices as output.

[0270] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0271] overview

[0272] This invention combines an emotion engine with a system that converts news articles into four-panel comics and videos to customize the content based on the user's emotions, making it possible to provide individual news content that matches the user's interests and emotions.

[0273] System Configuration

[0274] The system of the present invention consists of the following main modules:

[0275] 1. News article acquisition module

[0276] 2. Text Analysis Module

[0277] 3. Four-panel comic generation module

[0278] 4. 4-frame video generation module

[0279] 5. Content Display and Sharing Module

[0280] 6. Emotion Recognition and Response Module

[0281] News article acquisition module

[0282] The server retrieves the latest news articles using the news site's API or RSS feed, and stores the retrieved news articles in a database.

[0283] Examples:

[0284] The server accesses the "News API" to retrieve the latest news articles, then stores the article title, text, and URL in a database.

[0285] Text Analysis Module

[0286] The server analyzes the text of the retrieved news articles and utilizes natural language processing (NLP) techniques to extract important keywords and key points.

[0287] Examples:

[0288] The server analyzes news articles and extracts the keywords "climate change," "abnormal weather," and "impact."

[0289] Four-panel comic generation module

[0290] The server uses a generation AI to automatically generate a four-panel comic scenario based on the extracted keywords and key points. This scenario includes the situation and storyline of each panel. The generated scenario is then passed to an image generation AI, which generates illustrations for each panel.

[0291] Examples:

[0292] The server creates the scenario as follows:

[0293] 1. Frame 1: A scene where a character thinks about "climate change"

[0294] 2. Frame 2: A scene where abnormal weather is occurring

[0295] 3. Panel 3: People are surprised by the abnormal weather.

[0296] 4. Panel 4: A scene of characters thinking about climate change countermeasures

[0297] Based on this, the image generation AI generates four illustrations.

[0298] 4-frame video generation module

[0299] The server generates an animation based on the generated four-panel comic, creating a four-panel video by displaying each still image in succession and adding movement.

[0300] Examples:

[0301] The server animates each frame of the four-panel comic to create a short video.

[0302] Content Viewing and Sharing Module

[0303] The server places the four-panel comics and videos on news sites and also posts them on video sharing platforms and social networking services (SNS).

[0304] Examples:

[0305] The server places the generated four-panel comic on the homepage of the news site, uploads the four-panel video to YouTube or Twitter, and sets up a link to the news article.

[0306] Emotion Recognition and Response Module

[0307] The server utilizes an emotion engine to recognize the user's emotions. The emotion engine analyzes the user's camera footage and voice input to recognize the user's emotional state. Based on the recognized emotions, the content of the four-panel comics and videos is customized.

[0308] Examples:

[0309] The server analyzes the image from the user's facial recognition camera and recognizes that the user is "happy."

[0310] Based on this emotion, the emotion engine generates and displays four-panel comics with positive content to the user.

[0311] Similarly, the content and order of the four-frame videos are customized according to the user's emotions.

[0312] Example

[0313] When a user accesses a news site, their emotions are recognized through a camera or voice input device. Based on the recognized emotions, a four-panel comic strip or video that best suits the user is generated and displayed. Users can click on this content to read detailed news articles. They can also access the news site from a four-panel video they have viewed on a social media or video sharing platform.

[0314] example:

[0315] A user visits a news site and is perceived as "excited" based on five minutes of webcam footage.

[0316] Based on this, the server generates a four-panel comic strip of the active content and displays it on the news site's homepage.

[0317] Users click on the thumbnail of a four-panel comic that interests them and read a detailed news article.

[0318] The above is a specific embodiment of the present invention, and by combining emotion recognition, it is possible to provide personalized news content to users, which is more likely to attract users' attention and increase usage frequency than conventional news distribution methods.

[0319] The processing flow will be explained below.

[0320] Step 1:

[0321] The server retrieves the latest news articles using the news site's API or RSS feed. It periodically sends requests using the API key to collect news data. The retrieved news articles are then stored in a database.

[0322] Specific behavior:

[0323] The server accesses the "News API" and sends a request to retrieve the latest news articles.

[0324] The title, text, and URL of the news article received as a response are extracted and stored in a database.

[0325] Step 2:

[0326] The server invokes a natural language processing (NLP) engine to analyze the text of the stored news articles, thereby extracting important keywords and key points.

[0327] Specific behavior:

[0328] The server inputs the text of the news article into the NLP engine.

[0329] The NLP engine analyzes the text and extracts key keywords (e.g., "climate change," "extreme weather," "impact") and key points.

[0330] The extracted keywords and points are stored in a database.

[0331] Step 3:

[0332] The server uses generative AI to generate a four-panel comic scenario based on the extracted keywords and points. The generated scenario defines the situation and storyline of each panel.

[0333] Specific behavior:

[0334] The server inputs keywords and points into the generation AI and requests it to generate a four-panel scenario.

[0335] The generation AI creates a scenario and returns a description of each frame (e.g., "Frame 1: A character thinks about climate change," "Frame 2: Extreme weather occurs," etc.).

[0336] The generated scenario is saved in the database.

[0337] Step 4:

[0338] The server uses image generation AI to create a four-panel comic based on the generated scenario. The image generation AI generates illustrations based on the text description of each panel.

[0339] Specific behavior:

[0340] The server inputs the scenario text into the image generation AI and requests it to generate an illustration.

[0341] The image generation AI generates an illustration corresponding to each frame and returns it to the server.

[0342] The generated illustrations are arranged in the form of a four-panel comic.

[0343] Step 5:

[0344] The server generates animations based on each frame of the four-panel comic, creating a four-panel video. It displays still images in succession and adds movement to each scene.

[0345] Specific behavior:

[0346] The server provides the AI ​​with a script to animate the illustrations in each panel of the four-panel comic.

[0347] The generative AI generates the movement of each frame and creates a video file as an animation.

[0348] The generated video file is saved on the server.

[0349] Step 6:

[0350] The server places the generated four-panel comics and videos on the front page of the news site, and also posts them on video sharing platforms and social networking services to promote access.

[0351] Specific behavior:

[0352] The server generates HTML code to embed the four-panel comic on the top page of the news site.

[0353] The generated four-panel comic is displayed on the top page.

[0354] Upload your four-frame video to platforms such as YouTube and Twitter.

[0355] Link to the news article in the video description.

[0356] Step 7:

[0357] Users can click on thumbnails of four-panel comics or videos displayed on the homepage of news sites to read detailed news articles, or watch videos on social media or video-sharing platforms and access the original articles.

[0358] Specific behavior:

[0359] A user visits the homepage of a news site and clicks on a thumbnail of a four-panel comic strip.

[0360] Clicking on the thumbnail will take you to the news article details page.

[0361] Users watch four-frame videos on YouTube or Twitter and click on a link in the video description to access a news site.

[0362] Step 8:

[0363] The server uses an emotion engine to recognize the user's emotions. It analyzes the user's camera footage and voice input to recognize the user's emotional state.

[0364] Specific behavior:

[0365] When a user accesses a news site, the server collects camera footage and audio data.

[0366] The server calls the emotion engine and inputs the collected video and audio data.

[0367] The emotion engine analyzes the user's emotions (e.g., "happy" or "excited") and returns the results to the server.

[0368] Step 9:

[0369] The server customizes the content of the four-panel comic or video based on the recognized emotion.

[0370] Specific behavior:

[0371] The server adjusts the content of the four-panel comics and videos based on the emotional data obtained from the emotion engine.

[0372] For example, if the user is recognized as "happy," a four-panel comic strip with positive content is generated.

[0373] Similarly, the content and order of the four-frame video are changed according to the user's emotions.

[0374] The above are the specific processing steps of the present invention that incorporate emotion recognition, which enable providing users with personalized visual news content.

[0375] Example 2

[0376] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0377] In recent years, the distribution methods of news articles have become more diverse, making it increasingly difficult to keep readers interested. Furthermore, if the content of the article does not match the reader's emotions, it is difficult to maintain their interest. Therefore, there is a need for a method to provide news articles in a more user-friendly format and further customize the content based on the user's emotions.

[0378] The specific processing by the specific processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for acquiring news articles, means for analyzing the acquired news articles and extracting important data and key points, means for generating a four-panel cartoon scenario using a generative AI model based on the extracted data and key points, means for generating a four-panel cartoon using an image generation AI based on the generated scenario, means for generating animation based on the generated four-panel cartoon and creating a four-panel video, means for displaying the generated four-panel cartoon and four-panel video, means for posting the generated four-panel video to an information sharing platform or a social networking service to promote access to the original article, and means for recognizing user emotions and customizing content based on the user's emotions. This makes it possible to provide news articles in familiar four-panel cartoon or video formats and further provide personalized content tailored to the user's emotions.

[0379] A "news article" is a report of a fact or event provided online or in print.

[0380] A "server" is a computer or program that processes data on a network and provides services to other computers.

[0381] "Means of acquisition" are the methods and techniques used to collect specific data and store it in a usable format.

[0382] "Means of analysis" are methods and techniques for analyzing given data and understanding its meaning and structure.

[0383] "Important data" is information, such as news articles, that is deemed to be particularly relevant and valuable.

[0384] "Gist" refers to the most important or core part of information or data.

[0385] A "generative AI model" is an algorithm or program that uses artificial intelligence technology to generate new data and information.

[0386] "Scenario generation means" refers to methods and technologies for automatically creating stories and structures based on basic data.

[0387] "Image generation AI" refers to algorithms and programs that use artificial intelligence technology to generate image data based on input information.

[0388] "Four-panel comics" are a form of comic expression that consists of four images or illustrations arranged as a continuous story.

[0389] "Means for generating animation" refers to methods or techniques for expressing movement by displaying still images or illustrations in succession.

[0390] "4-frame video" is a video format that displays four scenes in succession and adds movement and effects.

[0391] The "display means" refers to a method or technology for providing the generated content in a format that can be visually confirmed by the user.

[0392] An "information sharing platform" is a website or application that allows users to share content or information with other users over the Internet.

[0393] A "social networking service (SNS)" is a service that allows users to connect with each other online and communicate and share information.

[0394] "Original article" refers to the original news article collected by the news article acquisition means and used for analysis and transformation.

[0395] "Means for recognizing emotions" refers to methods or technologies for identifying a user's emotional state by analyzing their facial expressions and voice.

[0396] "Customization means" are methods and techniques for modifying and tailoring content to a user's specific needs or circumstances.

[0397] overview

[0398] This invention is a system that converts news articles into four-panel comics and four-panel videos and customizes the content based on the user's emotions, making it possible to provide individual news content that matches the user's interests.

[0399] System Configuration

[0400] The system consists of the following main modules:

[0401] 1. News article acquisition module

[0402] 2. Text Analysis Module

[0403] 3. Four-panel comic generation module

[0404] 4. 4-frame video generation module

[0405] 5. Content Display and Sharing Module

[0406] 6. Emotion Recognition and Response Module

[0407] News article acquisition module

[0408] The server retrieves the latest news articles using the news site's API or RSS feed, which allows the latest news to be stored in the database at all times.

[0409] Examples:

[0410] The server accesses the "News API" to retrieve the latest news articles, then stores the article title, body text, and URL in a database.

[0411] Text Analysis Module

[0412] The server analyzes the text of the retrieved news articles and uses natural language processing (NLP) techniques, such as morphological analysis and topic modeling, to extract key data and key points.

[0413] Examples:

[0414] The server analyzes news articles and extracts the keywords "climate change," "abnormal weather," and "impact."

[0415] Four-panel comic generation module

[0416] The server uses a generative AI model to automatically generate a four-panel comic scenario based on the extracted data and key points. Based on this scenario, an image generation AI generates illustrations for each panel.

[0417] Examples:

[0418] Prompt: "A character who thinks about climate change"

[0419] 1. Frame 1: A scene where a character thinks about "climate change"

[0420] 2. Frame 2: A scene where abnormal weather is occurring

[0421] 3. Panel 3: People are surprised by the abnormal weather.

[0422] 4. Panel 4: A scene of characters thinking about climate change countermeasures

[0423] The server inputs a prompt sentence into the generative AI model, and passes the output scenario to the image generation AI to generate an illustration.

[0424] 4-frame video generation module

[0425] The server generates animations based on the generated four-panel comic illustrations, creating a four-panel video. Each still image is displayed in succession, and a short video is created by adding movement.

[0426] Examples:

[0427] The server uses animation production software (e.g., Adobe After Effects) to place the illustrations for each frame on a timeline, add transitions and effects, and generate the video.

[0428] Content Viewing and Sharing Module

[0429] The server places the generated four-panel comics and four-panel videos on news sites and also posts them on video sharing platforms and social networking services (SNS).

[0430] Examples:

[0431] The server uses HTML and JavaScript to display four-panel comics on the news site's homepage, and posts four-panel videos using YouTube and Twitter APIs.

[0432] Emotion Recognition and Response Module

[0433] The server uses an emotion engine to recognize the user's emotions. It analyzes the user's camera footage and voice input to recognize their emotional state. The content of the four-panel comics and videos is customized based on the recognized emotions.

[0434] Examples:

[0435] The server uses emotion recognition software (e.g., Microsoft Azure Cognitive Services) to analyze the user's camera footage, and if it recognizes the emotion as "happy," it generates and displays a four-panel comic strip with a positive message.

[0436] Example

[0437] When a user accesses a news site, their emotions are recognized through a camera or voice input device. Based on the recognized emotions, a four-panel comic strip or video that best suits the user is generated and displayed. Users can click on this content to read detailed news articles. They can also access the news site from a four-panel video they have viewed on a social media or video sharing platform.

[0438] Examples:

[0439] A user visits a news site and is identified as "excited" from five minutes of webcam footage. Based on this, the server generates a four-panel comic strip of the active content and displays it on the news site's homepage. The user clicks on the thumbnail of the four-panel comic strip that interests them and reads the detailed news article.

[0440] The above is a specific embodiment of the invention, and by combining it with emotion recognition, it is possible to provide personalized news content to users, which is more likely to attract users' attention and increase usage frequency compared to conventional news distribution methods.

[0441] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0442] Step 1: Get news articles

[0443] The server retrieves the latest news articles using the news site's API or RSS feed. The input is the news site's URL or API endpoint, and the output is the news article data. Specifically, the server sends an HTTP request, extracts the title, body, and URL of the news article from the received response, and stores them in a database.

[0444] Step 2: Analyzing the news article

[0445] The server retrieves news articles stored in a database and analyzes the text using natural language processing (NLP) technology. The input is the text data of the news article, and the output is extracted keywords and key points. Specifically, the server uses a morphological analysis tool to segment the text and performs topic modeling to extract important keywords.

[0446] Step 3: Creating a four-panel comic scenario

[0447] The server uses a generative AI model based on the extracted keywords and key points to automatically generate a four-panel comic scenario. The input is the extracted keywords, and the output is the four-panel comic scenario. Specifically, the server inputs a prompt statement (e.g., "A character thinking about climate change") into the generative AI model and retrieves the generated scenario.

[0448] Step 4: Generate a four-panel comic

[0449] The server uses image generation AI to generate illustrations for a four-panel comic based on the generated scenario. The input is the generated scenario, and the output is the illustrations for each panel. Specifically, the server passes the scenario for each panel to the image generation AI, which then generates the illustrations.

[0450] Step 5: Generate a 4-frame video

[0451] The server generates animation based on the generated four-panel cartoon illustrations, creating a four-panel video. The input is a four-panel cartoon illustration, and the output is a four-panel video. Specifically, the server uses animation production software to place the illustrations of each frame on a timeline, add effects and transitions, and create a short video.

[0452] Step 6: View and share content

[0453] The server places the generated four-panel comics and videos on news sites and also posts them on video sharing platforms and social networking services. The input is the generated four-panel comics and videos, and the output is public information on the news sites and sharing platforms. Specifically, the server updates the HTML and JavaScript on the news sites and posts the videos using the YouTube and Twitter APIs.

[0454] Step 7: User Emotion Recognition and Customization

[0455] The server uses an emotion engine to recognize the user's emotions. The input is the user's camera footage and audio input, and the output is the recognized emotional state. Specifically, the server uses emotion recognition software to analyze the camera footage and identify the user's emotions. Based on the emotional state, the content of the four-panel comic or video is customized.

[0456] By performing the above processing steps, news articles can be presented in the form of familiar four-panel comic strips or videos, and personalized content tailored to the user's emotions can be provided.

[0457] (Application example 2)

[0458] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0459] Modern news distribution services only provide information one-way to users, and have the problem of being unable to provide personalized content that reflects users' emotions and interests. Furthermore, they lack the visual and entertaining format to keep users engaged with the news. As a result, it is difficult to maintain users' interest, leading to a decline in frequency of use.

[0460] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0461] In this invention, the server includes: means for acquiring news articles; means for analyzing the acquired news articles and extracting important keywords and key points; means for generating a four-panel cartoon scenario using a generation AI based on the extracted keywords and key points; means for generating a four-panel cartoon using an image generation AI based on the generated scenario; means for generating an animation based on the generated four-panel cartoon and creating a four-panel video; means for displaying the generated four-panel cartoon and four-panel video; means for posting the generated four-panel video on a video sharing platform or a social networking service to promote access to the original article; means for recognizing a user's emotions using a camera or a voice input device; means for customizing the content of the generated four-panel cartoon and four-panel video based on the recognized emotion; and means for displaying the customized four-panel cartoon and four-panel video within the smartphone app. This allows news articles to be provided in a form adapted to the user's emotions, thereby increasing user interest and increasing usage frequency.

[0462] "Means for obtaining news articles" is a function for automatically collecting the latest news articles from news sites and APIs on the Internet.

[0463] "Means for analyzing news articles and extracting important keywords and key points" refers to technology for analyzing the content of acquired news articles and automatically extracting the main topics and important information of the articles.

[0464] "Generative AI" is a general term for artificial intelligence that uses natural language processing and machine learning techniques to generate stories and scenarios from input data.

[0465] "Image generation AI" is artificial intelligence that creates visual representations based on given scenarios or prompts.

[0466] The "means for generating four-panel comics" is a technology that automatically creates a comic consisting of four consecutive panels based on a generated scenario.

[0467] "Means for generating animations and creating four-frame videos" refers to a technology that converts the generated four-frame comics into moving animations and provides them in the form of short videos.

[0468] "Means for displaying the generated four-panel comics and four-panel videos" refers to technology for displaying the generated content in a manner that allows users to visually access it.

[0469] "A means of posting the generated four-frame video on video sharing platforms and social networking services to promote access to the original article" refers to a function that posts the generated video on YouTube or social media, directing viewers to the original news article.

[0470] "Means for recognizing a user's emotions using a camera or voice input device" refers to technology that analyzes camera footage and voice data to detect the user's emotional state.

[0471] "Means for customizing the content of generated four-panel comics and four-panel videos based on recognized emotions" refers to technology that adjusts the content and tone of generated content based on the user's emotional data.

[0472] "Means for displaying customized four-panel comics and four-panel videos within a smartphone app" is a function for providing users with emotionally tailored content through a smartphone app.

[0473] A specific embodiment for carrying out the present invention will be described.

[0474] Hardware and software used

[0475] The server uses the following major hardware and software:

[0476] Hardware

[0477] Internet connection required to retrieve news data

[0478] Camera and microphone for recognizing user emotions

[0479] Display devices such as smartphones and tablets

[0480] software

[0481] News Acquisition API

[0482] Natural language processing libraries (such as TextBlob)

[0483] Image generation AI (general image generation algorithms)

[0484] Video Generation Algorithm

[0485] Emotion recognition software (OpenCV and other facial recognition technologies)

[0486] Social media APIs (YouTube API, Twitter API, etc.)

[0487] Data processing and calculation flow

[0488] The server retrieves the latest news articles from the Internet, using a news API to pull news data into the server, and then analyzes the text data of the news articles using natural language processing technology to extract important keywords and key points.

[0489] Based on the extracted keywords and points, a scenario for a four-panel comic is automatically generated by the generative AI. This scenario is passed to the image generation AI used to generate a comic illustration consisting of four consecutive frames. The generated four-panel comic is then animated and presented as a short video.

[0490] The generated comic strips and videos are displayed on smartphone apps and other devices. When a user accesses the app, the app recognizes the user's emotions in real time via the camera and microphone, and the content of the comic strips and videos generated is customized based on those emotions. These contents are also posted on social networking services and video sharing platforms, and are used as a means to direct users to the original news article.

[0491] Specific examples

[0492] For example, if a news article mentions "climate change," the system might:

[0493] News acquisition

[0494] Use the News API to get the latest "climate change" articles.

[0495] Text analytics

[0496] Use a natural language processing library such as TextBlob to extract keywords such as "climate change," "extreme weather," and "impact" from articles.

[0497] Scenario generation and image generation

[0498] A scenario is generated based on the extracted keywords, and the image generation AI is used to draw the following four-panel comic scenario:

[0499] 1. A scene where a character thinks about climate change

[0500] 2. Scenes where abnormal weather is occurring

[0501] 3. Scenes of people being surprised by extreme weather

[0502] 4. A scene with characters thinking about climate change

[0503] Based on these scenarios, we generate prompts and pass them to the image generation AI:

[0504] "Generate a frame of a character talking at length about climate change."

[0505] "Draw a scene where extreme weather is occurring."

[0506] "Draw a frame showing people being surprised by the extreme weather."

[0507] Video Generation and Display

[0508] An animation is created based on the generated four-panel comic and converted into a short video format.

[0509] The video and four-panel comics are displayed within a smartphone app. When a user accesses the app, the camera and microphone are used to recognize their emotions, and content is customized accordingly.

[0510] Social Media Posts

[0511] The resulting four-frame video can then be posted on social media and video sharing platforms for a wider audience.

[0512] In this way, news articles can be provided in a form that is adapted to the user's emotions, which can increase the user's interest and increase the frequency of use.

[0513] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0514] Step 1:

[0515] The server uses a news API to retrieve the latest news articles from the Internet. The retrieved article data includes the article title, body text, URL, etc. The input is the response from the news API, and the output is text data for analysis.

[0516] Step 2:

[0517] The server analyzes the text data of the acquired news articles using natural language processing (NLP) techniques. Specifically, it uses libraries such as TextBlob to extract important keywords and key points. The input for this process is the text data, and the output is the extracted keywords and key points.

[0518] Step 3:

[0519] The server uses a generative AI model to generate a four-panel comic scenario based on the extracted keywords and points. Here, a group of keywords is passed as input, and the output is four prompt sentences containing the scenario for each panel. Specifically, the server converts the keywords into prompt sentences and passes them to the generative AI.

[0520] Step 4:

[0521] The server generates a four-panel comic using an image generation AI based on the generated scenario. At this stage, the input is a prompt sentence, and the output is four consecutive comic illustrations. Specifically, the prompt is passed to the image generation AI based on the scenario output from the generative AI model.

[0522] Step 5:

[0523] The server generates animation based on the generated four-panel comic to create a four-panel video. Here, still images of the four-panel comic are passed as input, and the output is a short video with continuous movements. Specifically, the video is generated by connecting still images and adding movement.

[0524] Step 6:

[0525] The server displays the generated four-panel comics and four-panel videos on a smartphone app or other device. The input is the generated four-panel comics and videos, and the output is visual data that is displayed to the user. Specifically, it includes the operation of displaying them in a specific section within the smartphone app.

[0526] Step 7:

[0527] When a user accesses the smartphone app, the app uses the device's camera and microphone to recognize emotions in real time. The input is the user's real-time video and audio data, and the output is analyzed emotional data. Specifically, the app uses emotion recognition software to analyze facial expressions and tone of voice.

[0528] Step 8:

[0529] The server customizes the content of the generated four-panel comics and four-panel videos based on the recognized emotions. The input is emotion data, and the output is customized four-panel comics and videos. Specific operations include reselecting scenarios and images according to the emotion.

[0530] Step 9:

[0531] The server posts customized four-panel comics and videos to social networking services and video sharing platforms to promote access to the original article. The input is the customized video and comic, and the output is a post on social media. Specifically, it automates the posting using the platform's API.

[0532] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0533] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0534] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0535] [Second embodiment]

[0536] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0537] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0538] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0539] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0540] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0541] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0542] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0543] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0544] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0545] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0546] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0547] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0548] overview

[0549] This invention is a system that converts news articles into visually intuitive four-panel comics and four-panel videos. This system utilizes generative AI and image generation AI technologies to extract the key points of news articles and present them in an easy-to-understand format for users.

[0550] System Configuration

[0551] The system of the present invention consists of the following main parts:

[0552] 1. News article acquisition module

[0553] 2. Text Analysis Module

[0554] 3. Four-panel comic generation module

[0555] 4. 4-frame video generation module

[0556] 5. Content Display and Sharing Module

[0557] News article acquisition module

[0558] The server retrieves the latest news articles through the news site's API or RSS feed. This retrieval operation can be set to occur periodically, and the latest articles are automatically imported into the system. The retrieved news articles are stored in a database.

[0559] Examples:

[0560] The server accesses the "News API" to retrieve the title, body text, and URL of the latest news article, then stores this information in a database.

[0561] Text Analysis Module

[0562] The server uses natural language processing techniques to analyze the text of the retrieved news articles and extract key keywords and points, a process that clarifies the core of the article.

[0563] Examples:

[0564] The server analyzes the news article "Impacts of Climate Change" and extracts the main keywords "climate change," "abnormal weather," and "impact."

[0565] Four-panel comic generation module

[0566] The server uses a generation AI to create a four-panel comic scenario based on the extracted keywords and points. This scenario includes the situation and storyline of each panel. The generated scenario is then passed to an image generation AI, which generates illustrations for each panel.

[0567] Examples:

[0568] The server creates the scenario as follows:

[0569] 1. Frame 1: A character is thinking about "climate change."

[0570] 2. Frame 2: A scene where abnormal weather is occurring

[0571] 3. Panel 3: People are surprised by the abnormal weather.

[0572] 4. Panel 4: A scene of characters thinking about climate change countermeasures

[0573] Based on this, the image generation AI generates four illustrations.

[0574] 4-frame video generation module

[0575] The server then creates an animation based on the four-panel comic, generating a four-panel video. During this process, the still images in each frame are animated and presented as a continuous video.

[0576] Examples:

[0577] The server breaks down a four-panel comic and animates each scene. For example, it animates the character's thoughts in the first panel, then moves to show the occurrence of extreme weather.

[0578] Content Viewing and Sharing Module

[0579] The server displays the generated four-panel comics and videos on the homepage of the news site, and also posts these videos on video sharing platforms and social networking services to promote access to the news articles.

[0580] Examples:

[0581] The server places the generated four-panel comic on the homepage of the news site, and then uploads the four-panel video to YouTube and Twitter, with a link to the article.

[0582] Example

[0583] When users visit a news site, they can click on a thumbnail of a four-panel comic that appears on the homepage to read a detailed news article. Users who are interested in a four-panel comic that has been shared on a video sharing platform or social networking service can click on the link to access the original article.

[0584] example:

[0585] A user clicks on a four-panel cartoon on the homepage of a news site to read a detailed article about the impacts of climate change. They also watch a four-panel video on YouTube and access the news site via a link in the video description.

[0586] As described above, the system of the present invention converts news articles into a visual and intuitive format, attracting the interest of the digital native generation and promoting access to news sites.

[0587] The processing flow will be explained below.

[0588] Step 1:

[0589] The server retrieves the latest news articles using the news site's API or RSS feed. It periodically sends requests using the API key to collect news data. The retrieved news articles are then stored in a database.

[0590] Specific behavior:

[0591] The server accesses the "News API" and sends a request to retrieve the latest news articles.

[0592] The title, text, and URL of the news article received as a response are extracted and stored in a database.

[0593] Step 2:

[0594] The server invokes a natural language processing (NLP) engine to analyze the text of the stored news articles, thereby extracting important keywords and key points.

[0595] Specific behavior:

[0596] The server inputs the text of the news article into the NLP engine.

[0597] The NLP engine analyzes the text and extracts key keywords (e.g., "climate change," "extreme weather," "impact") and key points.

[0598] The extracted keywords and points are stored in a database.

[0599] Step 3:

[0600] The server uses generative AI to generate a four-panel comic scenario based on the extracted keywords and points. The generated scenario defines the situation and storyline of each panel.

[0601] Specific behavior:

[0602] The server inputs keywords and points into the generation AI and requests it to generate a four-panel scenario.

[0603] The generation AI creates a scenario and returns a description of each frame (e.g., "Frame 1: A character thinks about climate change," "Frame 2: Extreme weather occurs," etc.).

[0604] The generated scenario is saved in the database.

[0605] Step 4:

[0606] The server uses image generation AI to create a four-panel comic based on the generated scenario. The image generation AI generates illustrations based on the text description of each panel.

[0607] Specific behavior:

[0608] The server inputs the scenario text into the image generation AI and requests it to generate an illustration.

[0609] The image generation AI generates an illustration corresponding to each frame and returns it to the server.

[0610] The generated illustrations are arranged in the form of a four-panel comic.

[0611] Step 5:

[0612] The server generates animations based on each frame of the four-panel comic, creating a four-panel video. It displays still images in succession and adds movement to each scene.

[0613] Specific behavior:

[0614] The server provides the AI ​​with a script to animate the illustrations in each panel of the four-panel comic.

[0615] The generative AI generates the movement of each frame and creates a video file as an animation.

[0616] The generated video file is saved on the server.

[0617] Step 6:

[0618] The server places the generated four-panel comics and videos on the front page of the news site, and also posts them on video sharing platforms and social networking services to promote access.

[0619] Specific behavior:

[0620] The server generates HTML code to embed the four-panel comic on the top page of the news site.

[0621] The generated four-panel comic is displayed on the top page.

[0622] Upload your four-frame video to platforms such as YouTube and Twitter.

[0623] Link to the news article in the video description.

[0624] Step 7:

[0625] Users can click on thumbnails of four-panel comics or videos displayed on the homepage of news sites to read detailed news articles, or watch videos on social media or video-sharing platforms and access the original articles.

[0626] Specific behavior:

[0627] A user visits the homepage of a news site and clicks on a thumbnail of a four-panel comic strip.

[0628] Clicking on the thumbnail will take you to the news article details page.

[0629] Users watch four-frame videos on YouTube or Twitter and click on a link in the video description to access a news site.

[0630] The above are the specific processing steps of the present invention. By executing these steps, it is possible to provide visual and intuitive news content to the digital native generation and promote access to news sites.

[0631] Example 1

[0632] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0633] Traditional news content is often expressed in text format, making it difficult to provide a visual and intuitive understanding, especially for the digital native generation. Furthermore, the process of visualizing and animating news articles is time-consuming, and automation is needed.

[0634] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0635] In this invention, the server includes a means for acquiring news content, a means for analyzing the acquired news content and extracting important keywords and information, and a means for generating a scenario for story-style content using a generative AI model based on the extracted keywords and information, thereby automating the task of visually and intuitively presenting news articles.

[0636] "News content" refers to information such as articles and reports obtained from news sites and other information platforms.

[0637] "Important keywords and information" refers to content and key points that are particularly noteworthy extracted from news content.

[0638] A "generative AI model" refers to an artificial intelligence model that understands natural language and generates new sentences and scenarios based on that information.

[0639] "Scenario for narrative content" refers to a storyline for constructing a continuous narrative, such as a four-panel comic strip, generated based on keywords and information extracted from news content.

[0640] An "image generation model" refers to an artificial intelligence model for generating images based on text information.

[0641] "Visual content" refers to visually presented information, such as four-panel comic strips, created using generative AI models and image generation models.

[0642] "Animation" refers to moving image content created by sequentially displaying visual content and adding movement.

[0643] "Video Content" refers to continuous, moving images generated from visual content.

[0644] A "video sharing platform" refers to a service that allows users to share videos over the Internet, such as YouTube.

[0645] "Social media" refers to online platforms such as Twitter and Facebook that allow users to communicate with each other and share content.

[0646] overview

[0647] This invention is a system that converts news content into visually intuitive four-panel comics and four-panel videos. This system utilizes generative AI and image generation model technologies to extract key points from news content and present them in an easy-to-understand format for users.

[0648] System Configuration

[0649] The system of the present invention consists of the following main modules:

[0650] 1. News article acquisition module

[0651] 2. Text Analysis Module

[0652] 3. Four-panel comic generation module

[0653] 4. 4-frame video generation module

[0654] 5. Content Display and Sharing Module

[0655] News article acquisition module

[0656] The server retrieves the latest news content using the news site's API or RSS feed. This retrieval operation can be performed periodically, and the latest news content is automatically imported into the system. The retrieved content is then stored in a database.

[0657] Examples:

[0658] The server accesses the "News API" to retrieve the title, body text, and URL of the latest news article, then stores this information in a database.

[0659] Text Analysis Module

[0660] The server uses natural language processing techniques to analyze the text of the retrieved news content and extract key keywords and information, thereby clarifying the core of the article.

[0661] Examples:

[0662] The server analyzes the news article "Impacts of Climate Change" and extracts the main keywords "climate change," "abnormal weather," and "impact."

[0663] Four-panel comic generation module

[0664] The server uses a generative AI model to create a four-panel comic scenario based on the extracted keywords and information. This scenario includes the situation and storyline of each panel. The generated scenario is then passed to an image generation model, which generates illustrations for each panel.

[0665] Examples:

[0666] The server creates the scenario as follows:

[0667] 1. Frame 1: A character is thinking about "climate change."

[0668] 2. Frame 2: A scene where abnormal weather is occurring

[0669] 3. Panel 3: People are surprised by the abnormal weather.

[0670] 4. Panel 4: A scene of characters thinking about climate change countermeasures

[0671] Based on this, an image generation model generates four illustrations.

[0672] 4-frame video generation module

[0673] The server then creates an animation based on the four-panel comic, generating a four-panel video. During this process, the still images in each frame are animated and presented as a continuous video.

[0674] Examples:

[0675] The server breaks down a four-panel comic and animates each scene. For example, it animates the character's thoughts in the first panel, then moves to show the occurrence of extreme weather.

[0676] Content Viewing and Sharing Module

[0677] The server displays the generated four-panel comics and videos on the homepage of the news site, and also posts these videos on video sharing platforms and social media to promote access to the news content.

[0678] Examples:

[0679] The server places the generated four-panel comic on the homepage of the news site, and also uploads the four-panel video to YouTube and Twitter, with a link to the article.

[0680] Examples of prompt statements

[0681] Below are some example prompts for using generative AI models:

[0682] Example prompt sentence:

[0683] 1. Prompts to extract necessary keywords from news articles:

[0684] Extract key keywords from the news article "Impacts of Climate Change."

[0685] 2. Prompts for generating a four-panel comic scenario:

[0686] Please create a four-panel comic scenario based on the extracted keywords "climate change," "extreme weather," and "impact." Please be specific and include a storyline for each panel.

[0687] 3. Prompt to pass to the image generation model:

[0688] Please generate an illustration for each frame based on the following scenario.

[0689] 1. Frame 1: A character is thinking about "climate change."

[0690] 2. Frame 2: A scene where abnormal weather is occurring

[0691] 3. Panel 3: People are surprised by the abnormal weather.

[0692] 4. Panel 4: A scene of characters thinking about climate change countermeasures

[0693] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0694] Step 1:

[0695] The server retrieves news content. It uses the news site's API or RSS feed as input. The output is news article data (title, body text, URL, etc.), which is stored in a database.

[0696] Specific behavior:

[0697] The server sends a request to the "News API" to receive the latest news article data. The received data (e.g., article data on the impact of climate change) is stored in the corresponding table in the database.

[0698] Step 2:

[0699] The server reads news content from the database and performs text analysis. The input is the text data of the news article obtained in step 1, and the output is a list of important keywords and information.

[0700] Specific behavior:

[0701] The server uses a natural language processing library (e.g., SpaCy or NLTK) to analyze the text of a news article (e.g., "impacts of climate change") and extracts the main keywords "climate change," "extreme weather," and "impact."

[0702] Step 3:

[0703] The server uses a generative AI model based on the extracted keywords and information to create a four-panel comic scenario. The input is the keywords and information extracted in step 2, and the output is the four-panel comic scenario.

[0704] Specific behavior:

[0705] The server sends a prompt to the generation AI model, for example, "Create a four-panel comic scenario based on the following keywords: Keywords: climate change, extreme weather, impact," and generates a scenario.

[0706] Example scenario:

[0707] 1. Frame 1: A character is thinking about "climate change."

[0708] 2. Frame 2: A scene where abnormal weather is occurring

[0709] 3. Panel 3: People are surprised by the abnormal weather.

[0710] 4. Panel 4: A scene of characters thinking about climate change countermeasures

[0711] Step 4:

[0712] Based on the generated scenario, the server uses an image generation model to generate illustrations for each frame of the four-panel comic. The input is the scenario created in step 3, and the output is a visual illustration for each frame.

[0713] Specific behavior:

[0714] The server sends prompts to the image generation model (e.g., DALL-E). For example, it uses the image generation prompt for the first frame, "A scene in which a character is thinking about climate change," to generate illustrations for each frame.

[0715] Step 5:

[0716] The server creates an animation based on the generated four-panel comic and generates a four-panel video. The input is the visual illustration generated in step 4, and the output is a four-panel video.

[0717] Specific behavior:

[0718] The server uses video editing software or libraries (e.g., OpenCV) to animate the images in each frame. For example, you could animate a character thinking in the first frame, then smoothly transition to an extreme weather scene.

[0719] Step 6:

[0720] The server displays the generated comic strips and videos on the homepage of the news site, and also posts these videos to video sharing platforms and social media. The input is the comic strips and videos generated in step 5, and the output is the web page where the information is displayed and the posted content.

[0721] Specific behavior:

[0722] The server places thumbnails of the four-panel comics on the homepage of the news site, and when users click on them, detailed news content is displayed. The server also uploads the four-panel comics to YouTube and Twitter, automatically linking them to the articles.

[0723] (Application example 1)

[0724] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0725] News articles are primarily text-based, making it difficult to understand their content quickly. Furthermore, digital natives, in particular, are demanding content in a visual, intuitive format. To meet these modern needs, a new way of presenting news articles in a visual, easy-to-understand format is needed.

[0726] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0727] In this invention, the server includes means for acquiring news articles, means for analyzing the acquired news articles and extracting important keywords and key points, means for generating a four-panel cartoon scenario using a generation AI based on the extracted keywords and key points, means for generating a four-panel cartoon using an image generation AI based on the generated scenario, means for generating animation based on the generated four-panel cartoon and creating a four-panel video, means for displaying the generated four-panel cartoon and four-panel video, means for posting the generated four-panel video on a video sharing platform or social networking service to promote access to the original article, and means for displaying and sharing the generated four-panel cartoon and four-panel video on a smart device, thereby enabling users to visually understand the content of the news article in a short amount of time.

[0728] A "means for acquiring news articles" is a system or module for automatically acquiring the latest news articles from news sites or RSS feeds on the Internet.

[0729] "Means for analyzing acquired news articles and extracting important keywords and key points" refers to a system or module that has the function of analyzing the text of news articles using natural language processing technology and extracting the subject matter and important information of the articles.

[0730] "Means for generating a four-panel comic scenario using generative AI" refers to a system or module that uses generative AI to construct a storyline for a four-panel comic and the content of each panel based on keywords and key points extracted from news articles.

[0731] "Means for generating four-panel comics using image generation AI" refers to a system or module that uses image generation AI to create illustrations for four-panel comics based on a generated scenario.

[0732] "Means for generating animation based on the generated four-panel comic and creating a four-panel video" refers to a system or module that animates the still images of each frame and presents them as a continuous four-panel video.

[0733] "Means for displaying the generated four-panel comics and four-panel videos" refers to a system or module for displaying the generated four-panel comics and four-panel videos on a user interface.

[0734] "Means for posting the generated four-panel videos on video sharing platforms or social networking services to promote access to the original article" refers to a system or module that uploads the generated four-panel videos to video sharing platforms or social networking services and provides users with a link that allows them to access the original news article.

[0735] "Means for displaying and sharing generated four-panel comics and four-panel videos on a smart device" refers to a system or module that allows users to display four-panel comics and four-panel videos generated on a smart device such as a smartphone or tablet, and share the content with other users via a link or share button.

[0736] "Natural language processing technology" is a computer technology for analyzing text data and extracting or summarizing important information.

[0737] A "content distribution service" is an online service that provides users with various types of digital content via the Internet.

[0738] Overall system overview

[0739] This system aims to provide news articles as visually intuitive four-panel comics and four-panel videos. Specifically, it supports processes including news article acquisition, analysis, scenario creation using generative AI, comic generation using image generation AI, animation, display, and sharing. The main components of this system are a news article acquisition module, text analysis module, four-panel comic generation module, four-panel video generation module, and content display and sharing module.

[0740] Hardware and Software

[0741] The hardware used includes servers and smart devices (e.g., smartphones). The main software includes:

[0742] Python 3.x: Programming Language

[0743] Requests: A library for making HTTP requests

[0744] BeautifulSoup: HTML parsing library

[0745] Spacy: A natural language processing library

[0746] Transformers: A library for using generative AI

[0747] Image generation API (hypothetical API)

[0748] Processing flow

[0749] The server first retrieves news articles from news sites and RSS feeds. Next, it analyzes the retrieved news articles and extracts important keywords and key points using natural language processing technology. Based on the extracted keywords and key points, it uses generative AI to generate a four-panel comic scenario. It then uses image generation AI to create four-panel comic illustrations based on the scenario. It then animates the generated four-panel comic to create a four-panel video. Finally, the generated four-panel comic and video can be displayed and shared on a smart device.

[0750] Specific examples

[0751] As a concrete example, we take the news article "The spread of the new coronavirus" and extract the key keywords "new coronavirus," "social distancing," and "vaccine." The following scenario is generated:

[0752] 1. A scene where a character is thinking about the new coronavirus

[0753] 2. Scenes where social distancing is being implemented

[0754] 3. Scenes of people maintaining social distancing

[0755] 4. A scene where a character thinks about getting vaccinated

[0756] An example of a prompt passed to the generative AI model is:

[0757] Prompt: Transform the contents of a news article into a four-panel comic book scenario.

[0758] Key Word 1: Novel Coronavirus

[0759] Key word 2: Social distancing

[0760] Key word 3: Vaccine

[0761] The generated four-panel comics and videos can be displayed on smart devices such as smartphones and can be easily shared via social media, allowing users to quickly and visually understand the content of news articles.

[0762] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0763] Step 1:

[0764] The server periodically retrieves news articles from news sites and RSS feeds on the Internet. Specifically, the server accesses a news API, retrieves the title, body, and URL of the latest news article, and stores them in a database. The URL of the news API is provided as input, and the JSON data of the news article is obtained as output.

[0765] Step 2:

[0766] The server analyzes the retrieved news articles and extracts important keywords and key points. In this process, it uses natural language processing technology (Spacy) to analyze the article text and extract key keywords and noun phrases. The body of the news article is provided as input, and a list of important keywords is obtained as output.

[0767] Step 3:

[0768] The server uses a generative AI model to generate a four-panel comic scenario based on the extracted keywords and key points. Specifically, a prompt is input to the generative AI, which outputs four scenario blocks. The extracted keywords and prompt are provided as input, and the scenario for each panel is generated as output.

[0769] Step 4:

[0770] The server uses image generation AI to generate a four-panel comic based on the generated scenario. The generated scenario is passed to an image generation API, which generates illustrations for each panel. The scenario is provided as input, and four illustrations are generated as output.

[0771] Step 5:

[0772] The server generates an animation based on the generated four-panel comic, creating a four-panel video. The illustrations for each frame are imported into animation software (e.g., Adobe After Effects) to create a continuous animation. Four illustrations are provided as input, and a four-panel video is generated as output.

[0773] Step 6:

[0774] The server displays the generated four-panel comic and four-panel video. In this step, the generated content is displayed on a user interface so that the user can view it. The four-panel comic and four-panel video are provided as input, and are displayed on the user interface as output.

[0775] Step 7:

[0776] The server posts the generated four-frame comics to video sharing platforms and social networking services to promote access to the original article. Specifically, it uploads the videos using APIs such as YouTube and Twitter and sets up links to the articles. The four-frame comics and the URLs of the original articles are provided as input, and the videos are published on the video platforms as output.

[0777] Step 8:

[0778] The server displays and shares the generated four-panel comics and four-panel videos on smart devices. Users view the generated content on their smartphones or tablets and click share links on social media to share the content with other users. Four-panel comics and four-panel videos are provided as input, and can be displayed and shared on smart devices as output.

[0779] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0780] overview

[0781] This invention combines an emotion engine with a system that converts news articles into four-panel comics and videos to customize the content based on the user's emotions, making it possible to provide individual news content that matches the user's interests and emotions.

[0782] System Configuration

[0783] The system of the present invention consists of the following main modules:

[0784] 1. News article acquisition module

[0785] 2. Text Analysis Module

[0786] 3. Four-panel comic generation module

[0787] 4. 4-frame video generation module

[0788] 5. Content Display and Sharing Module

[0789] 6. Emotion Recognition and Response Module

[0790] News article acquisition module

[0791] The server retrieves the latest news articles using the news site's API or RSS feed, and stores the retrieved news articles in a database.

[0792] Examples:

[0793] The server accesses the "News API" to retrieve the latest news articles, then stores the article title, text, and URL in a database.

[0794] Text Analysis Module

[0795] The server analyzes the text of the retrieved news articles and utilizes natural language processing (NLP) techniques to extract important keywords and key points.

[0796] Examples:

[0797] The server analyzes news articles and extracts the keywords "climate change," "abnormal weather," and "impact."

[0798] Four-panel comic generation module

[0799] The server uses a generation AI to automatically generate a four-panel comic scenario based on the extracted keywords and key points. This scenario includes the situation and storyline of each panel. The generated scenario is then passed to an image generation AI, which generates illustrations for each panel.

[0800] Examples:

[0801] The server creates the scenario as follows:

[0802] 1. Frame 1: A scene where a character thinks about "climate change"

[0803] 2. Frame 2: A scene where abnormal weather is occurring

[0804] 3. Panel 3: People are surprised by the abnormal weather.

[0805] 4. Panel 4: A scene of characters thinking about climate change countermeasures

[0806] Based on this, the image generation AI generates four illustrations.

[0807] 4-frame video generation module

[0808] The server generates an animation based on the generated four-panel comic, creating a four-panel video by displaying each still image in succession and adding movement.

[0809] Examples:

[0810] The server animates each frame of the four-panel comic to create a short video.

[0811] Content Viewing and Sharing Module

[0812] The server places the four-panel comics and videos on news sites and also posts them on video sharing platforms and social networking services (SNS).

[0813] Examples:

[0814] The server places the generated four-panel comic on the homepage of the news site, uploads the four-panel video to YouTube or Twitter, and sets up a link to the news article.

[0815] Emotion Recognition and Response Module

[0816] The server utilizes an emotion engine to recognize the user's emotions. The emotion engine analyzes the user's camera footage and voice input to recognize the user's emotional state. Based on the recognized emotions, the content of the four-panel comics and videos is customized.

[0817] Examples:

[0818] The server analyzes the image from the user's facial recognition camera and recognizes that the user is "happy."

[0819] Based on this emotion, the emotion engine generates and displays four-panel comics with positive content to the user.

[0820] Similarly, the content and order of the four-frame videos are customized according to the user's emotions.

[0821] Example

[0822] When a user accesses a news site, their emotions are recognized through a camera or voice input device. Based on the recognized emotions, a four-panel comic strip or video that best suits the user is generated and displayed. Users can click on this content to read detailed news articles. They can also access the news site from a four-panel video they have viewed on a social media or video sharing platform.

[0823] example:

[0824] A user visits a news site and is perceived as "excited" based on five minutes of webcam footage.

[0825] Based on this, the server generates a four-panel comic strip of the active content and displays it on the news site's homepage.

[0826] Users click on the thumbnail of a four-panel comic that interests them and read a detailed news article.

[0827] The above is a specific embodiment of the present invention, and by combining emotion recognition, it is possible to provide personalized news content to users, which is more likely to attract users' attention and increase usage frequency than conventional news distribution methods.

[0828] The processing flow will be explained below.

[0829] Step 1:

[0830] The server retrieves the latest news articles using the news site's API or RSS feed. It periodically sends requests using the API key to collect news data. The retrieved news articles are then stored in a database.

[0831] Specific behavior:

[0832] The server accesses the "News API" and sends a request to retrieve the latest news articles.

[0833] The title, text, and URL of the news article received as a response are extracted and stored in a database.

[0834] Step 2:

[0835] The server invokes a natural language processing (NLP) engine to analyze the text of the stored news articles, thereby extracting important keywords and key points.

[0836] Specific behavior:

[0837] The server inputs the text of the news article into the NLP engine.

[0838] The NLP engine analyzes the text and extracts key keywords (e.g., "climate change," "extreme weather," "impact") and key points.

[0839] The extracted keywords and points are stored in a database.

[0840] Step 3:

[0841] The server uses generative AI to generate a four-panel comic scenario based on the extracted keywords and points. The generated scenario defines the situation and storyline of each panel.

[0842] Specific behavior:

[0843] The server inputs keywords and points into the generation AI and requests it to generate a four-panel scenario.

[0844] The generation AI creates a scenario and returns a description of each frame (e.g., "Frame 1: A character thinks about climate change," "Frame 2: Extreme weather occurs," etc.).

[0845] The generated scenario is saved in the database.

[0846] Step 4:

[0847] The server uses image generation AI to create a four-panel comic based on the generated scenario. The image generation AI generates illustrations based on the text description of each panel.

[0848] Specific behavior:

[0849] The server inputs the scenario text into the image generation AI and requests it to generate an illustration.

[0850] The image generation AI generates an illustration corresponding to each frame and returns it to the server.

[0851] The generated illustrations are arranged in the form of a four-panel comic.

[0852] Step 5:

[0853] The server generates animations based on each frame of the four-panel comic, creating a four-panel video. It displays still images in succession and adds movement to each scene.

[0854] Specific behavior:

[0855] The server provides the AI ​​with a script to animate the illustrations in each panel of the four-panel comic.

[0856] The generative AI generates the movement of each frame and creates a video file as an animation.

[0857] The generated video file is saved on the server.

[0858] Step 6:

[0859] The server places the generated four-panel comics and videos on the front page of the news site, and also posts them on video sharing platforms and social networking services to promote access.

[0860] Specific behavior:

[0861] The server generates HTML code to embed the four-panel comic on the top page of the news site.

[0862] The generated four-panel comic is displayed on the top page.

[0863] Upload your four-frame video to platforms such as YouTube and Twitter.

[0864] Link to the news article in the video description.

[0865] Step 7:

[0866] Users can click on thumbnails of four-panel comics or videos displayed on the homepage of news sites to read detailed news articles, or watch videos on social media or video-sharing platforms and access the original articles.

[0867] Specific behavior:

[0868] A user visits the homepage of a news site and clicks on a thumbnail of a four-panel comic strip.

[0869] Clicking on the thumbnail will take you to the news article details page.

[0870] Users watch four-frame videos on YouTube or Twitter and click on a link in the video description to access a news site.

[0871] Step 8:

[0872] The server uses an emotion engine to recognize the user's emotions. It analyzes the user's camera footage and voice input to recognize the user's emotional state.

[0873] Specific behavior:

[0874] When a user accesses a news site, the server collects camera footage and audio data.

[0875] The server calls the emotion engine and inputs the collected video and audio data.

[0876] The emotion engine analyzes the user's emotions (e.g., "happy" or "excited") and returns the results to the server.

[0877] Step 9:

[0878] The server customizes the content of the four-panel comic or video based on the recognized emotion.

[0879] Specific behavior:

[0880] The server adjusts the content of the four-panel comics and videos based on the emotional data obtained from the emotion engine.

[0881] For example, if the user is recognized as "happy," a four-panel comic strip with positive content is generated.

[0882] Similarly, the content and order of the four-frame video are changed according to the user's emotions.

[0883] The above are the specific processing steps of the present invention that incorporate emotion recognition, which enable providing users with personalized visual news content.

[0884] Example 2

[0885] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0886] In recent years, the distribution methods of news articles have become more diverse, making it increasingly difficult to keep readers interested. Furthermore, if the content of the article does not match the reader's emotions, it is difficult to maintain their interest. Therefore, there is a need for a method to provide news articles in a more user-friendly format and further customize the content based on the user's emotions.

[0887] The specific processing by the specific processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for acquiring news articles, means for analyzing the acquired news articles and extracting important data and key points, means for generating a four-panel cartoon scenario using a generative AI model based on the extracted data and key points, means for generating a four-panel cartoon using an image generation AI based on the generated scenario, means for generating animation based on the generated four-panel cartoon and creating a four-panel video, means for displaying the generated four-panel cartoon and four-panel video, means for posting the generated four-panel video to an information sharing platform or a social networking service to promote access to the original article, and means for recognizing user emotions and customizing content based on the user's emotions. This makes it possible to provide news articles in familiar four-panel cartoon or video formats and further provide personalized content tailored to the user's emotions.

[0888] A "news article" is a report of a fact or event provided online or in print.

[0889] A "server" is a computer or program that processes data on a network and provides services to other computers.

[0890] "Means of acquisition" are the methods and techniques used to collect specific data and store it in a usable format.

[0891] "Means of analysis" are methods and techniques for analyzing given data and understanding its meaning and structure.

[0892] "Important data" is information, such as news articles, that is deemed to be particularly relevant and valuable.

[0893] "Gist" refers to the most important or core part of information or data.

[0894] A "generative AI model" is an algorithm or program that uses artificial intelligence technology to generate new data and information.

[0895] "Scenario generation means" refers to methods and technologies for automatically creating stories and structures based on basic data.

[0896] "Image generation AI" refers to algorithms or programs that use artificial intelligence technology to generate image data based on input information.

[0897] "Four-panel comics" are a form of comic expression that consists of four images or illustrations arranged as a continuous story.

[0898] "Means for generating animation" refers to methods or techniques for expressing movement by displaying still images or illustrations in succession.

[0899] "4-frame video" is a video format that displays four scenes in succession and adds movement and effects.

[0900] The "display means" refers to a method or technology for providing the generated content in a format that can be visually confirmed by the user.

[0901] An "information sharing platform" is a website or application that allows users to share content or information with other users over the Internet.

[0902] A "social networking service (SNS)" is a service that allows users to connect with each other online and communicate and share information.

[0903] "Original article" refers to the original news article collected by the news article acquisition means and used for analysis and transformation.

[0904] "Means for recognizing emotions" refers to methods or technologies for identifying a user's emotional state by analyzing their facial expressions and voice.

[0905] "Customization means" are methods and techniques for modifying and tailoring content to a user's specific needs or circumstances.

[0906] overview

[0907] This invention is a system that converts news articles into four-panel comics and four-panel videos and customizes the content based on the user's emotions, making it possible to provide individual news content that matches the user's interests.

[0908] System Configuration

[0909] The system consists of the following main modules:

[0910] 1. News article acquisition module

[0911] 2. Text Analysis Module

[0912] 3. Four-panel comic generation module

[0913] 4. 4-frame video generation module

[0914] 5. Content Display and Sharing Module

[0915] 6. Emotion Recognition and Response Module

[0916] News article acquisition module

[0917] The server retrieves the latest news articles using the news site's API or RSS feed, which allows the latest news to be stored in the database at all times.

[0918] Examples:

[0919] The server accesses the "News API" to retrieve the latest news articles, then stores the article title, body text, and URL in a database.

[0920] Text Analysis Module

[0921] The server analyzes the text of the retrieved news articles and uses natural language processing (NLP) techniques, such as morphological analysis and topic modeling, to extract key data and key points.

[0922] Examples:

[0923] The server analyzes news articles and extracts the keywords "climate change," "abnormal weather," and "impact."

[0924] Four-panel comic generation module

[0925] The server uses a generative AI model to automatically generate a four-panel comic scenario based on the extracted data and key points. Based on this scenario, an image generation AI generates illustrations for each panel.

[0926] Examples:

[0927] Prompt: "A character who thinks about climate change"

[0928] 1. Frame 1: A scene where a character thinks about "climate change"

[0929] 2. Frame 2: A scene where abnormal weather is occurring

[0930] 3. Panel 3: People are surprised by the abnormal weather.

[0931] 4. Panel 4: A scene of characters thinking about climate change countermeasures

[0932] The server inputs a prompt sentence into the generative AI model, and passes the output scenario to the image generation AI to generate an illustration.

[0933] 4-frame video generation module

[0934] The server generates animations based on the generated four-panel comic illustrations, creating a four-panel video. Each still image is displayed in succession, and a short video is created by adding movement.

[0935] Examples:

[0936] The server uses animation production software (e.g., Adobe After Effects) to place the illustrations for each frame on a timeline, add transitions and effects, and generate the video.

[0937] Content Viewing and Sharing Module

[0938] The server places the generated four-panel comics and four-panel videos on a news site and also posts them on video sharing platforms and social networking services (SNS).

[0939] Examples:

[0940] The server uses HTML and JavaScript to display four-panel comics on the news site's homepage, and posts four-panel videos using YouTube and Twitter APIs.

[0941] Emotion Recognition and Response Module

[0942] The server uses an emotion engine to recognize the user's emotions. It analyzes the user's camera footage and voice input to recognize their emotional state. The content of the four-panel comics and videos is customized based on the recognized emotions.

[0943] Examples:

[0944] The server uses emotion recognition software (e.g., Microsoft Azure Cognitive Services) to analyze the user's camera footage, and if it recognizes the emotion as "happy," it generates and displays a four-panel comic strip with a positive message.

[0945] Example

[0946] When a user accesses a news site, their emotions are recognized through a camera or voice input device. Based on the recognized emotions, a four-panel comic strip or video that best suits the user is generated and displayed. Users can click on this content to read detailed news articles. They can also access the news site from a four-panel video they have viewed on a social media or video sharing platform.

[0947] Examples:

[0948] A user visits a news site and is identified as "excited" from five minutes of webcam footage. Based on this, the server generates a four-panel comic strip of the active content and displays it on the news site's homepage. The user clicks on the thumbnail of the four-panel comic strip that interests them and reads the detailed news article.

[0949] The above is a specific embodiment of the invention, and by combining it with emotion recognition, it is possible to provide personalized news content to users, which is more likely to attract users' attention and increase usage frequency compared to conventional news distribution methods.

[0950] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0951] Step 1: Get news articles

[0952] The server retrieves the latest news articles using the news site's API or RSS feed. The input is the news site's URL or API endpoint, and the output is the news article data. Specifically, the server sends an HTTP request, extracts the title, body, and URL of the news article from the received response, and stores them in a database.

[0953] Step 2: Analyzing the news article

[0954] The server retrieves news articles stored in a database and analyzes the text using natural language processing (NLP) technology. The input is the text data of the news article, and the output is extracted keywords and key points. Specifically, the server uses a morphological analysis tool to segment the text and performs topic modeling to extract important keywords.

[0955] Step 3: Creating a four-panel comic scenario

[0956] The server uses a generative AI model based on the extracted keywords and key points to automatically generate a four-panel comic scenario. The input is the extracted keywords, and the output is the four-panel comic scenario. Specifically, the server inputs a prompt statement (e.g., "A character thinking about climate change") into the generative AI model and retrieves the generated scenario.

[0957] Step 4: Generate a four-panel comic

[0958] The server uses image generation AI to generate illustrations for a four-panel comic based on the generated scenario. The input is the generated scenario, and the output is the illustrations for each panel. Specifically, the server passes the scenario for each panel to the image generation AI, which then generates the illustrations.

[0959] Step 5: Generate a 4-frame video

[0960] The server generates animation based on the generated four-panel cartoon illustrations, creating a four-panel video. The input is a four-panel cartoon illustration, and the output is a four-panel video. Specifically, the server uses animation production software to place the illustrations of each frame on a timeline, add effects and transitions, and create a short video.

[0961] Step 6: View and share content

[0962] The server places the generated four-panel comics and videos on news sites and also posts them on video sharing platforms and social networking services. The input is the generated four-panel comics and videos, and the output is public information on the news sites and sharing platforms. Specifically, the server updates the HTML and JavaScript on the news sites and posts the videos using the YouTube and Twitter APIs.

[0963] Step 7: User Emotion Recognition and Customization

[0964] The server uses an emotion engine to recognize the user's emotions. The input is the user's camera footage and audio input, and the output is the recognized emotional state. Specifically, the server uses emotion recognition software to analyze the camera footage and identify the user's emotions. Based on the emotional state, the content of the four-panel comic or video is customized.

[0965] By performing the above processing steps, news articles can be presented in the form of familiar four-panel comic strips or videos, and personalized content tailored to the user's emotions can be provided.

[0966] (Application example 2)

[0967] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0968] Modern news distribution services only provide information one-way to users, and have the problem of being unable to provide personalized content that reflects users' emotions and interests. Furthermore, they lack the visual and entertaining format to keep users engaged with the news. As a result, it is difficult to maintain users' interest, leading to a decline in frequency of use.

[0969] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0970] In this invention, the server includes: means for acquiring news articles; means for analyzing the acquired news articles and extracting important keywords and key points; means for generating a four-panel cartoon scenario using a generation AI based on the extracted keywords and key points; means for generating a four-panel cartoon using an image generation AI based on the generated scenario; means for generating an animation based on the generated four-panel cartoon and creating a four-panel video; means for displaying the generated four-panel cartoon and four-panel video; means for posting the generated four-panel video on a video sharing platform or a social networking service to promote access to the original article; means for recognizing a user's emotions using a camera or a voice input device; means for customizing the content of the generated four-panel cartoon and four-panel video based on the recognized emotion; and means for displaying the customized four-panel cartoon and four-panel video within the smartphone app. This allows news articles to be provided in a form adapted to the user's emotions, thereby increasing user interest and increasing usage frequency.

[0971] "Means for obtaining news articles" is a function for automatically collecting the latest news articles from news sites and APIs on the Internet.

[0972] "Means for analyzing news articles and extracting important keywords and key points" refers to technology for analyzing the content of acquired news articles and automatically extracting the main topics and important information of the articles.

[0973] "Generative AI" is a general term for artificial intelligence that uses natural language processing and machine learning techniques to generate stories and scenarios from input data.

[0974] "Image generation AI" is artificial intelligence that creates visual representations based on given scenarios or prompts.

[0975] The "means for generating four-panel comics" is a technology that automatically creates a comic consisting of four consecutive panels based on a generated scenario.

[0976] "Means for generating animations and creating four-frame videos" refers to a technology that converts the generated four-frame comics into moving animations and provides them in the form of short videos.

[0977] "Means for displaying the generated four-panel comics and four-panel videos" refers to technology for displaying the generated content in a manner that allows users to visually access it.

[0978] "A means of posting the generated four-frame video on video sharing platforms and social networking services to promote access to the original article" refers to a function that posts the generated video on YouTube or social media, directing viewers to the original news article.

[0979] "Means for recognizing a user's emotions using a camera or voice input device" refers to technology that analyzes camera footage and voice data to detect the user's emotional state.

[0980] "Means for customizing the content of generated four-panel comics and four-panel videos based on recognized emotions" refers to technology that adjusts the content and tone of generated content based on the user's emotional data.

[0981] "Means for displaying customized four-panel comics and four-panel videos within a smartphone app" is a function for providing users with emotionally tailored content through a smartphone app.

[0982] A specific embodiment for carrying out the present invention will be described.

[0983] Hardware and software used

[0984] The server uses the following major hardware and software:

[0985] Hardware

[0986] Internet connection required to retrieve news data

[0987] Camera and microphone for recognizing user emotions

[0988] Display devices such as smartphones and tablets

[0989] software

[0990] News Acquisition API

[0991] Natural language processing libraries (such as TextBlob)

[0992] Image generation AI (general image generation algorithms)

[0993] Video Generation Algorithm

[0994] Emotion recognition software (OpenCV and other facial recognition technologies)

[0995] Social media APIs (YouTube API, Twitter API, etc.)

[0996] Data processing and calculation flow

[0997] The server retrieves the latest news articles from the Internet, using a news API to pull news data into the server, and then analyzes the text data of the news articles using natural language processing technology to extract important keywords and key points.

[0998] Based on the extracted keywords and points, a scenario for a four-panel comic is automatically generated by the generative AI. This scenario is passed to the image generation AI used to generate a comic illustration consisting of four consecutive frames. The generated four-panel comic is then animated and presented as a short video.

[0999] The generated comic strips and videos are displayed on smartphone apps and other devices. When a user accesses the app, the app recognizes the user's emotions in real time via the camera and microphone, and the content of the comic strips and videos generated is customized based on those emotions. These contents are also posted on social networking services and video sharing platforms, and are used as a means to direct users to the original news article.

[1000] Specific examples

[1001] For example, if a news article mentions "climate change," the system might:

[1002] News acquisition

[1003] Use the News API to get the latest "Climate Change" articles.

[1004] Text analytics

[1005] Use a natural language processing library such as TextBlob to extract keywords such as "climate change," "extreme weather," and "impact" from articles.

[1006] Scenario generation and image generation

[1007] A scenario is generated based on the extracted keywords, and the image generation AI is used to draw the following four-panel comic scenario:

[1008] 1. A scene where a character thinks about climate change

[1009] 2. Scenes where abnormal weather is occurring

[1010] 3. Scenes of people being surprised by extreme weather

[1011] 4. A scene with characters thinking about climate change

[1012] Based on these scenarios, we generate prompts and pass them to the image generation AI:

[1013] "Generate a frame of a character talking at length about climate change."

[1014] "Draw a scene where extreme weather is occurring."

[1015] "Draw a frame showing people being surprised by the extreme weather."

[1016] Video Generation and Display

[1017] An animation is created based on the generated four-panel comic and converted into a short video format.

[1018] The video and four-panel comics are displayed within a smartphone app. When a user accesses the app, the camera and microphone are used to recognize their emotions, and content is customized accordingly.

[1019] Social Media Posts

[1020] The resulting four-frame video can then be posted on social media and video sharing platforms for a wider audience.

[1021] In this way, news articles can be provided in a form that is adapted to the user's emotions, which can increase the user's interest and increase the frequency of use.

[1022] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1023] Step 1:

[1024] The server uses a news API to retrieve the latest news articles from the Internet. The retrieved article data includes the article title, body text, URL, etc. The input is the response from the news API, and the output is text data for analysis.

[1025] Step 2:

[1026] The server analyzes the text data of the acquired news articles using natural language processing (NLP) techniques. Specifically, it uses libraries such as TextBlob to extract important keywords and key points. The input for this process is the text data, and the output is the extracted keywords and key points.

[1027] Step 3:

[1028] The server uses a generative AI model to generate a four-panel comic scenario based on the extracted keywords and points. Here, a group of keywords is passed as input, and the output is four prompt sentences containing the scenario for each panel. Specifically, the server converts the keywords into prompt sentences and passes them to the generative AI.

[1029] Step 4:

[1030] The server generates a four-panel comic using an image generation AI based on the generated scenario. At this stage, the input is a prompt sentence, and the output is four consecutive comic illustrations. Specifically, the prompt is passed to the image generation AI based on the scenario output from the generative AI model.

[1031] Step 5:

[1032] The server generates animation based on the generated four-panel comic to create a four-panel video. Here, still images of the four-panel comic are passed as input, and the output is a short video with continuous movements. Specifically, the video is generated by connecting still images and adding movement.

[1033] Step 6:

[1034] The server displays the generated four-panel comics and four-panel videos on a smartphone app or other device. The input is the generated four-panel comics and videos, and the output is visual data that is displayed to the user. Specifically, it includes the operation of displaying them in a specific section within the smartphone app.

[1035] Step 7:

[1036] When a user accesses the smartphone app, the app uses the device's camera and microphone to recognize emotions in real time. The input is the user's real-time video and audio data, and the output is analyzed emotional data. Specifically, the app uses emotion recognition software to analyze facial expressions and tone of voice.

[1037] Step 8:

[1038] The server customizes the content of the generated four-panel comics and four-panel videos based on the recognized emotions. The input is emotion data, and the output is customized four-panel comics and videos. Specific operations include reselecting scenarios and images according to the emotion.

[1039] Step 9:

[1040] The server posts customized four-panel comics and videos to social networking services and video sharing platforms to promote access to the original article. The input is the customized video and comic, and the output is a post on social media. Specifically, it automates the posting using the platform's API.

[1041] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1042] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1043] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[1044] [Third embodiment]

[1045] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[1046] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[1047] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1048] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[1049] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1050] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1051] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1052] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1053] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1054] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1055] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1056] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[1057] overview

[1058] This invention is a system that converts news articles into visually intuitive four-panel comics and four-panel videos. This system utilizes generative AI and image generation AI technologies to extract the key points of news articles and present them in an easy-to-understand format for users.

[1059] System Configuration

[1060] The system of the present invention consists of the following main parts:

[1061] 1. News article acquisition module

[1062] 2. Text Analysis Module

[1063] 3. Four-panel comic generation module

[1064] 4. 4-frame video generation module

[1065] 5. Content Display and Sharing Module

[1066] News article acquisition module

[1067] The server retrieves the latest news articles through the news site's API or RSS feed. This retrieval operation can be set to occur periodically, and the latest articles are automatically imported into the system. The retrieved news articles are stored in a database.

[1068] Examples:

[1069] The server accesses the "News API" to retrieve the title, body text, and URL of the latest news article, then stores this information in a database.

[1070] Text Analysis Module

[1071] The server uses natural language processing techniques to analyze the text of the retrieved news articles and extract key keywords and points, a process that clarifies the core of the article.

[1072] Examples:

[1073] The server analyzes the news article "Impacts of Climate Change" and extracts the main keywords "climate change," "abnormal weather," and "impact."

[1074] Four-panel comic generation module

[1075] The server uses a generation AI to create a four-panel comic scenario based on the extracted keywords and points. This scenario includes the situation and storyline of each panel. The generated scenario is then passed to an image generation AI, which generates illustrations for each panel.

[1076] Examples:

[1077] The server creates the scenario as follows:

[1078] 1. Frame 1: A character is thinking about "climate change."

[1079] 2. Frame 2: A scene where abnormal weather is occurring

[1080] 3. Panel 3: People are surprised by the abnormal weather.

[1081] 4. Panel 4: A scene of characters thinking about climate change countermeasures

[1082] Based on this, the image generation AI generates four illustrations.

[1083] 4-frame video generation module

[1084] The server then creates an animation based on the four-panel comic, generating a four-panel video. During this process, the still images in each frame are animated and presented as a continuous video.

[1085] Examples:

[1086] The server breaks down a four-panel comic and animates each scene. For example, it animates the character's thoughts in the first panel, then moves to show the occurrence of extreme weather.

[1087] Content Viewing and Sharing Module

[1088] The server displays the generated four-panel comics and videos on the homepage of the news site, and also posts these videos on video sharing platforms and social networking services to promote access to the news articles.

[1089] Examples:

[1090] The server places the generated four-panel comic on the homepage of the news site, and then uploads the four-panel video to YouTube and Twitter, with a link to the article.

[1091] Example

[1092] When users visit a news site, they can click on a thumbnail of a four-panel comic that appears on the homepage to read a detailed news article. Users who are interested in a four-panel comic that has been shared on a video sharing platform or social networking service can click on the link to access the original article.

[1093] example:

[1094] A user clicks on a four-panel cartoon on the homepage of a news site to read a detailed article about the impacts of climate change. They also watch a four-panel video on YouTube and access the news site via a link in the video description.

[1095] As described above, the system of the present invention converts news articles into a visual and intuitive format, attracting the interest of the digital native generation and promoting access to news sites.

[1096] The processing flow will be explained below.

[1097] Step 1:

[1098] The server retrieves the latest news articles using the news site's API or RSS feed. It periodically sends requests using the API key to collect news data. The retrieved news articles are then stored in a database.

[1099] Specific behavior:

[1100] The server accesses the "News API" and sends a request to retrieve the latest news articles.

[1101] The title, text, and URL of the news article received as a response are extracted and stored in a database.

[1102] Step 2:

[1103] The server invokes a natural language processing (NLP) engine to analyze the text of the stored news articles, thereby extracting important keywords and key points.

[1104] Specific behavior:

[1105] The server inputs the text of the news article into the NLP engine.

[1106] The NLP engine analyzes the text and extracts key keywords (e.g., "climate change," "extreme weather," "impact") and key points.

[1107] The extracted keywords and points are stored in a database.

[1108] Step 3:

[1109] The server uses generative AI to generate a four-panel comic scenario based on the extracted keywords and points. The generated scenario defines the situation and storyline of each panel.

[1110] Specific behavior:

[1111] The server inputs keywords and points into the generation AI and requests it to generate a four-panel scenario.

[1112] The generation AI creates a scenario and returns a description of each frame (e.g., "Frame 1: A character thinks about climate change," "Frame 2: Extreme weather occurs," etc.).

[1113] The generated scenario is saved in the database.

[1114] Step 4:

[1115] The server uses image generation AI to create a four-panel comic based on the generated scenario. The image generation AI generates illustrations based on the text description of each panel.

[1116] Specific behavior:

[1117] The server inputs the scenario text into the image generation AI and requests it to generate an illustration.

[1118] The image generation AI generates an illustration corresponding to each frame and returns it to the server.

[1119] The generated illustrations are arranged in the form of a four-panel comic.

[1120] Step 5:

[1121] The server generates animations based on each frame of the four-panel comic, creating a four-panel video. It displays still images in succession and adds movement to each scene.

[1122] Specific behavior:

[1123] The server provides the AI ​​with a script to animate the illustrations in each panel of the four-panel comic.

[1124] The generative AI generates the movement of each frame and creates a video file as an animation.

[1125] The generated video file is saved on the server.

[1126] Step 6:

[1127] The server places the generated four-panel comics and videos on the front page of the news site, and also posts them on video sharing platforms and social networking services to promote access.

[1128] Specific behavior:

[1129] The server generates HTML code to embed the four-panel comic on the top page of the news site.

[1130] The generated four-panel comic is displayed on the top page.

[1131] Upload your four-frame video to platforms such as YouTube and Twitter.

[1132] Link to the news article in the video description.

[1133] Step 7:

[1134] Users can click on thumbnails of four-panel comics or videos displayed on the homepage of news sites to read detailed news articles, or watch videos on social media or video-sharing platforms and access the original articles.

[1135] Specific behavior:

[1136] A user visits the homepage of a news site and clicks on a thumbnail of a four-panel comic strip.

[1137] Clicking on the thumbnail will take you to the news article details page.

[1138] Users watch four-frame videos on YouTube or Twitter and click on a link in the video description to access a news site.

[1139] The above are the specific processing steps of the present invention. By executing these steps, it is possible to provide visual and intuitive news content to the digital native generation and promote access to news sites.

[1140] Example 1

[1141] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1142] Traditional news content is often expressed in text format, making it difficult to provide a visual and intuitive understanding, especially for the digital native generation. Furthermore, the process of visualizing and animating news articles is time-consuming, and automation is needed.

[1143] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1144] In this invention, the server includes a means for acquiring news content, a means for analyzing the acquired news content and extracting important keywords and information, and a means for generating a scenario for story-style content using a generative AI model based on the extracted keywords and information, thereby automating the task of visually and intuitively presenting news articles.

[1145] "News content" refers to information such as articles and reports obtained from news sites and other information platforms.

[1146] "Important keywords and information" refers to content and key points that are particularly noteworthy extracted from news content.

[1147] A "generative AI model" refers to an artificial intelligence model that understands natural language and generates new sentences and scenarios based on that information.

[1148] "Scenario for narrative content" refers to a storyline for constructing a continuous narrative, such as a four-panel comic strip, generated based on keywords and information extracted from news content.

[1149] An "image generation model" refers to an artificial intelligence model for generating images based on text information.

[1150] "Visual content" refers to visually presented information, such as four-panel comic strips, created using generative AI models and image generation models.

[1151] "Animation" refers to moving image content created by sequentially displaying visual content and adding movement.

[1152] "Video Content" refers to continuous, moving images generated from visual content.

[1153] A "video sharing platform" refers to a service that allows users to share videos over the Internet, such as YouTube.

[1154] "Social media" refers to online platforms such as Twitter and Facebook that allow users to communicate with each other and share content.

[1155] overview

[1156] This invention is a system that converts news content into visually intuitive four-panel comics and four-panel videos. This system utilizes generative AI and image generation model technologies to extract key points from news content and present them in an easy-to-understand format for users.

[1157] System Configuration

[1158] The system of the present invention consists of the following main modules:

[1159] 1. News article acquisition module

[1160] 2. Text Analysis Module

[1161] 3. Four-panel comic generation module

[1162] 4. 4-frame video generation module

[1163] 5. Content Display and Sharing Module

[1164] News article acquisition module

[1165] The server retrieves the latest news content using the news site's API or RSS feed. This retrieval operation can be performed periodically, and the latest news content is automatically imported into the system. The retrieved content is then stored in a database.

[1166] Examples:

[1167] The server accesses the "News API" to retrieve the title, body text, and URL of the latest news article, then stores this information in a database.

[1168] Text Analysis Module

[1169] The server uses natural language processing techniques to analyze the text of the retrieved news content and extract key keywords and information, thereby clarifying the core of the article.

[1170] Examples:

[1171] The server analyzes the news article "Impacts of Climate Change" and extracts the main keywords "climate change," "abnormal weather," and "impact."

[1172] Four-panel comic generation module

[1173] The server uses a generative AI model to create a four-panel comic scenario based on the extracted keywords and information. This scenario includes the situation and storyline of each panel. The generated scenario is then passed to an image generation model, which generates illustrations for each panel.

[1174] Examples:

[1175] The server creates the scenario as follows:

[1176] 1. Frame 1: A character is thinking about "climate change."

[1177] 2. Frame 2: A scene where abnormal weather is occurring

[1178] 3. Panel 3: People are surprised by the abnormal weather.

[1179] 4. Panel 4: A scene of characters thinking about climate change countermeasures

[1180] Based on this, an image generation model generates four illustrations.

[1181] 4-frame video generation module

[1182] The server then creates an animation based on the four-panel comic, generating a four-panel video. During this process, the still images in each frame are animated and presented as a continuous video.

[1183] Examples:

[1184] The server breaks down a four-panel comic and animates each scene. For example, it animates the character's thoughts in the first panel, then moves to show the occurrence of extreme weather.

[1185] Content Viewing and Sharing Module

[1186] The server displays the generated four-panel comics and videos on the homepage of the news site, and also posts these videos on video sharing platforms and social media to promote access to the news content.

[1187] Examples:

[1188] The server places the generated four-panel comic on the homepage of the news site, and also uploads the four-panel video to YouTube and Twitter, with a link to the article.

[1189] Examples of prompt statements

[1190] Below are some example prompts for using generative AI models:

[1191] Example prompt sentence:

[1192] 1. Prompts to extract necessary keywords from news articles:

[1193] Extract key keywords from the news article "Impacts of Climate Change."

[1194] 2. Prompts for generating a four-panel comic scenario:

[1195] Please create a four-panel comic scenario based on the extracted keywords "climate change," "extreme weather," and "impact." Please be specific and include a storyline for each panel.

[1196] 3. Prompt to pass to the image generation model:

[1197] Please generate an illustration for each frame based on the following scenario.

[1198] 1. Frame 1: A character is thinking about "climate change."

[1199] 2. Frame 2: A scene where abnormal weather is occurring

[1200] 3. Panel 3: People are surprised by the abnormal weather.

[1201] 4. Panel 4: A scene of characters thinking about climate change countermeasures

[1202] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1203] Step 1:

[1204] The server retrieves news content. It uses the news site's API or RSS feed as input. The output is news article data (title, body text, URL, etc.), which is stored in a database.

[1205] Specific behavior:

[1206] The server sends a request to the "News API" to receive the latest news article data. The received data (e.g., article data on the impact of climate change) is stored in the corresponding table in the database.

[1207] Step 2:

[1208] The server reads news content from the database and performs text analysis. The input is the text data of the news article obtained in step 1, and the output is a list of important keywords and information.

[1209] Specific behavior:

[1210] The server uses a natural language processing library (e.g., SpaCy or NLTK) to analyze the text of a news article (e.g., "impacts of climate change") and extracts the main keywords "climate change," "extreme weather," and "impact."

[1211] Step 3:

[1212] The server uses a generative AI model based on the extracted keywords and information to create a four-panel comic scenario. The input is the keywords and information extracted in step 2, and the output is the four-panel comic scenario.

[1213] Specific behavior:

[1214] The server sends a prompt to the generation AI model, for example, "Create a four-panel comic scenario based on the following keywords: Keywords: climate change, extreme weather, impact," and generates a scenario.

[1215] Example scenario:

[1216] 1. Frame 1: A character is thinking about "climate change."

[1217] 2. Frame 2: A scene where abnormal weather is occurring

[1218] 3. Panel 3: People are surprised by the abnormal weather.

[1219] 4. Panel 4: A scene of characters thinking about climate change countermeasures

[1220] Step 4:

[1221] Based on the generated scenario, the server uses an image generation model to generate illustrations for each frame of the four-panel comic. The input is the scenario created in step 3, and the output is a visual illustration for each frame.

[1222] Specific behavior:

[1223] The server sends prompts to the image generation model (e.g., DALL-E). For example, it uses the image generation prompt for the first frame, "A scene in which a character is thinking about climate change," to generate illustrations for each frame.

[1224] Step 5:

[1225] The server creates an animation based on the generated four-panel comic and generates a four-panel video. The input is the visual illustration generated in step 4, and the output is a four-panel video.

[1226] Specific behavior:

[1227] The server uses video editing software or libraries (e.g., OpenCV) to animate the images in each frame. For example, you could animate a character thinking in the first frame, then smoothly transition to an extreme weather scene.

[1228] Step 6:

[1229] The server displays the generated comic strips and videos on the homepage of the news site, and also posts these videos to video sharing platforms and social media. The input is the comic strips and videos generated in step 5, and the output is the web page where the information is displayed and the posted content.

[1230] Specific behavior:

[1231] The server places thumbnails of the four-panel comics on the homepage of the news site, and when users click on them, detailed news content is displayed. The server also uploads the four-panel comics to YouTube and Twitter, automatically linking them to the articles.

[1232] (Application example 1)

[1233] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1234] News articles are primarily text-based, making it difficult to understand their content quickly. Furthermore, digital natives, in particular, are demanding content in a visual, intuitive format. To meet these modern needs, a new way of presenting news articles in a visual, easy-to-understand format is needed.

[1235] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1236] In this invention, the server includes means for acquiring news articles, means for analyzing the acquired news articles and extracting important keywords and key points, means for generating a four-panel cartoon scenario using a generation AI based on the extracted keywords and key points, means for generating a four-panel cartoon using an image generation AI based on the generated scenario, means for generating animation based on the generated four-panel cartoon and creating a four-panel video, means for displaying the generated four-panel cartoon and four-panel video, means for posting the generated four-panel video on a video sharing platform or social networking service to promote access to the original article, and means for displaying and sharing the generated four-panel cartoon and four-panel video on a smart device, thereby enabling users to visually understand the content of the news article in a short amount of time.

[1237] A "means for acquiring news articles" is a system or module for automatically acquiring the latest news articles from news sites or RSS feeds on the Internet.

[1238] "Means for analyzing acquired news articles and extracting important keywords and key points" refers to a system or module that has the function of analyzing the text of news articles using natural language processing technology and extracting the subject matter and important information of the articles.

[1239] "Means for generating a four-panel comic scenario using generative AI" refers to a system or module that uses generative AI to construct a storyline for a four-panel comic and the content of each panel based on keywords and key points extracted from news articles.

[1240] "Means for generating four-panel comics using image generation AI" refers to a system or module that uses image generation AI to create illustrations for four-panel comics based on a generated scenario.

[1241] "Means for generating animation based on the generated four-panel comic and creating a four-panel video" refers to a system or module that animates the still images of each frame and presents them as a continuous four-panel video.

[1242] "Means for displaying the generated four-panel comics and four-panel videos" refers to a system or module for displaying the generated four-panel comics and four-panel videos on a user interface.

[1243] "Means for posting the generated four-panel videos on video sharing platforms or social networking services to promote access to the original article" refers to a system or module that uploads the generated four-panel videos to video sharing platforms or social networking services and provides users with a link that allows them to access the original news article.

[1244] "Means for displaying and sharing generated four-panel comics and four-panel videos on a smart device" refers to a system or module that allows users to display four-panel comics and four-panel videos generated on a smart device such as a smartphone or tablet, and share the content with other users via a link or share button.

[1245] "Natural language processing technology" is a computer technology for analyzing text data and extracting or summarizing important information.

[1246] A "content distribution service" is an online service that provides users with various types of digital content via the Internet.

[1247] Overall system overview

[1248] This system aims to provide news articles as visually intuitive four-panel comics and four-panel videos. Specifically, it supports processes including news article acquisition, analysis, scenario creation using generative AI, comic generation using image generation AI, animation, display, and sharing. The main components of this system are a news article acquisition module, text analysis module, four-panel comic generation module, four-panel video generation module, and content display and sharing module.

[1249] Hardware and Software

[1250] The hardware used includes servers and smart devices (e.g., smartphones). The main software includes:

[1251] Python 3.x: Programming Language

[1252] Requests: A library for making HTTP requests

[1253] BeautifulSoup: HTML parsing library

[1254] Spacy: A natural language processing library

[1255] Transformers: A library for using generative AI

[1256] Image generation API (hypothetical API)

[1257] Processing flow

[1258] The server first retrieves news articles from news sites and RSS feeds. Next, it analyzes the retrieved news articles and extracts important keywords and key points using natural language processing technology. Based on the extracted keywords and key points, it uses generative AI to generate a four-panel comic scenario. It then uses image generation AI to create four-panel comic illustrations based on the scenario. It then animates the generated four-panel comic to create a four-panel video. Finally, the generated four-panel comic and video can be displayed and shared on a smart device.

[1259] Specific examples

[1260] As a concrete example, we take the news article "The spread of the new coronavirus" and extract the key keywords "new coronavirus," "social distancing," and "vaccine." The following scenario is generated:

[1261] 1. A scene where a character is thinking about the new coronavirus

[1262] 2. Scenes where social distancing is being implemented

[1263] 3. Scenes of people maintaining social distancing

[1264] 4. A scene where a character thinks about getting vaccinated

[1265] An example of a prompt passed to the generative AI model is:

[1266] Prompt: Transform the contents of a news article into a four-panel comic book scenario.

[1267] Key Word 1: Novel Coronavirus

[1268] Key word 2: Social distancing

[1269] Key word 3: Vaccine

[1270] The generated four-panel comics and videos can be displayed on smart devices such as smartphones and can be easily shared via social media, allowing users to quickly and visually understand the content of news articles.

[1271] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1272] Step 1:

[1273] The server periodically retrieves news articles from news sites and RSS feeds on the Internet. Specifically, the server accesses a news API, retrieves the title, body, and URL of the latest news article, and stores them in a database. The URL of the news API is provided as input, and the JSON data of the news article is obtained as output.

[1274] Step 2:

[1275] The server analyzes the retrieved news articles and extracts important keywords and key points. In this process, it uses natural language processing technology (Spacy) to analyze the article text and extract key keywords and noun phrases. The body of the news article is provided as input, and a list of important keywords is obtained as output.

[1276] Step 3:

[1277] The server uses a generative AI model to generate a four-panel comic scenario based on the extracted keywords and key points. Specifically, a prompt is input to the generative AI, which outputs four scenario blocks. The extracted keywords and prompt are provided as input, and the scenario for each panel is generated as output.

[1278] Step 4:

[1279] The server uses image generation AI to generate a four-panel comic based on the generated scenario. The generated scenario is passed to an image generation API, which generates illustrations for each panel. The scenario is provided as input, and four illustrations are generated as output.

[1280] Step 5:

[1281] The server generates an animation based on the generated four-panel comic, creating a four-panel video. The illustrations for each frame are imported into animation software (e.g., Adobe After Effects) to create a continuous animation. Four illustrations are provided as input, and a four-panel video is generated as output.

[1282] Step 6:

[1283] The server displays the generated four-panel comic and four-panel video. In this step, the generated content is displayed on a user interface so that the user can view it. The four-panel comic and four-panel video are provided as input, and are displayed on the user interface as output.

[1284] Step 7:

[1285] The server posts the generated four-frame comics to video sharing platforms and social networking services to promote access to the original article. Specifically, it uploads the videos using APIs such as YouTube and Twitter and sets up links to the articles. The four-frame comics and the URLs of the original articles are provided as input, and the videos are published on the video platforms as output.

[1286] Step 8:

[1287] The server displays and shares the generated four-panel comics and four-panel videos on smart devices. Users view the generated content on their smartphones or tablets and click share links on social media to share the content with other users. Four-panel comics and four-panel videos are provided as input, and can be displayed and shared on smart devices as output.

[1288] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1289] overview

[1290] This invention combines an emotion engine with a system that converts news articles into four-panel comics and videos to customize the content based on the user's emotions, making it possible to provide individual news content that matches the user's interests and emotions.

[1291] System Configuration

[1292] The system of the present invention consists of the following main modules:

[1293] 1. News article acquisition module

[1294] 2. Text Analysis Module

[1295] 3. Four-panel comic generation module

[1296] 4. 4-frame video generation module

[1297] 5. Content Display and Sharing Module

[1298] 6. Emotion Recognition and Response Module

[1299] News article acquisition module

[1300] The server retrieves the latest news articles using the news site's API or RSS feed, and stores the retrieved news articles in a database.

[1301] Examples:

[1302] The server accesses the "News API" to retrieve the latest news articles, then stores the article title, text, and URL in a database.

[1303] Text Analysis Module

[1304] The server analyzes the text of the retrieved news articles and utilizes natural language processing (NLP) techniques to extract important keywords and key points.

[1305] Examples:

[1306] The server analyzes news articles and extracts the keywords "climate change," "abnormal weather," and "impact."

[1307] Four-panel comic generation module

[1308] The server uses a generation AI to automatically generate a four-panel comic scenario based on the extracted keywords and key points. This scenario includes the situation and storyline of each panel. The generated scenario is then passed to an image generation AI, which generates illustrations for each panel.

[1309] Examples:

[1310] The server creates the scenario as follows:

[1311] 1. Frame 1: A scene where a character thinks about "climate change"

[1312] 2. Frame 2: A scene where abnormal weather is occurring

[1313] 3. Panel 3: People are surprised by the abnormal weather.

[1314] 4. Panel 4: A scene of characters thinking about climate change countermeasures

[1315] Based on this, the image generation AI generates four illustrations.

[1316] 4-frame video generation module

[1317] The server generates an animation based on the generated four-panel comic, creating a four-panel video by displaying each still image in succession and adding movement.

[1318] Examples:

[1319] The server animates each frame of the four-panel comic to create a short video.

[1320] Content Viewing and Sharing Module

[1321] The server places the four-panel comics and videos on news sites and also posts them on video sharing platforms and social networking services (SNS).

[1322] Examples:

[1323] The server places the generated four-panel comic on the homepage of the news site, uploads the four-panel video to YouTube or Twitter, and sets up a link to the news article.

[1324] Emotion Recognition and Response Module

[1325] The server utilizes an emotion engine to recognize the user's emotions. The emotion engine analyzes the user's camera footage and voice input to recognize the user's emotional state. Based on the recognized emotions, the content of the four-panel comics and videos is customized.

[1326] Examples:

[1327] The server analyzes the image from the user's facial recognition camera and recognizes that the user is "happy."

[1328] Based on this emotion, the emotion engine generates and displays four-panel comics with positive content to the user.

[1329] Similarly, the content and order of the four-frame videos are customized according to the user's emotions.

[1330] Example

[1331] When a user accesses a news site, their emotions are recognized through a camera or voice input device. Based on the recognized emotions, a four-panel comic strip or video that best suits the user is generated and displayed. Users can click on this content to read detailed news articles. They can also access the news site from a four-panel video they have viewed on a social media or video sharing platform.

[1332] example:

[1333] A user visits a news site and is perceived as "excited" based on five minutes of webcam footage.

[1334] Based on this, the server generates a four-panel comic strip of the active content and displays it on the news site's homepage.

[1335] Users click on the thumbnail of a four-panel comic that interests them and read a detailed news article.

[1336] The above is a specific embodiment of the present invention, and by combining emotion recognition, it is possible to provide personalized news content to users, which is more likely to attract users' attention and increase usage frequency than conventional news distribution methods.

[1337] The processing flow will be explained below.

[1338] Step 1:

[1339] The server retrieves the latest news articles using the news site's API or RSS feed. It periodically sends requests using the API key to collect news data. The retrieved news articles are then stored in a database.

[1340] Specific behavior:

[1341] The server accesses the "News API" and sends a request to retrieve the latest news articles.

[1342] The title, text, and URL of the news article received as a response are extracted and stored in a database.

[1343] Step 2:

[1344] The server invokes a natural language processing (NLP) engine to analyze the text of the stored news articles, thereby extracting important keywords and key points.

[1345] Specific behavior:

[1346] The server inputs the text of the news article into the NLP engine.

[1347] The NLP engine analyzes the text and extracts key keywords (e.g., "climate change," "extreme weather," "impact") and key points.

[1348] The extracted keywords and points are stored in a database.

[1349] Step 3:

[1350] The server uses generative AI to generate a four-panel comic scenario based on the extracted keywords and points. The generated scenario defines the situation and storyline of each panel.

[1351] Specific behavior:

[1352] The server inputs keywords and points into the generation AI and requests it to generate a four-panel scenario.

[1353] The generation AI creates a scenario and returns a description of each frame (e.g., "Frame 1: A character thinks about climate change," "Frame 2: Extreme weather occurs," etc.).

[1354] The generated scenario is saved in the database.

[1355] Step 4:

[1356] The server uses image generation AI to create a four-panel comic based on the generated scenario. The image generation AI generates illustrations based on the text description of each panel.

[1357] Specific behavior:

[1358] The server inputs the scenario text into the image generation AI and requests it to generate an illustration.

[1359] The image generation AI generates an illustration corresponding to each frame and returns it to the server.

[1360] The generated illustrations are arranged in the form of a four-panel comic.

[1361] Step 5:

[1362] The server generates animations based on each frame of the four-panel comic, creating a four-panel video. It displays still images in succession and adds movement to each scene.

[1363] Specific behavior:

[1364] The server provides the AI ​​with a script to animate the illustrations in each panel of the four-panel comic.

[1365] The generative AI generates the movement of each frame and creates a video file as an animation.

[1366] The generated video file is saved on the server.

[1367] Step 6:

[1368] The server places the generated four-panel comics and videos on the front page of the news site, and also posts them on video sharing platforms and social networking services to promote access.

[1369] Specific behavior:

[1370] The server generates HTML code to embed the four-panel comic on the top page of the news site.

[1371] The generated four-panel comic is displayed on the top page.

[1372] Upload your four-frame video to platforms such as YouTube and Twitter.

[1373] Link to the news article in the video description.

[1374] Step 7:

[1375] Users can click on thumbnails of four-panel comics or videos displayed on the homepage of news sites to read detailed news articles, or watch videos on social media or video-sharing platforms and access the original articles.

[1376] Specific behavior:

[1377] A user visits the homepage of a news site and clicks on a thumbnail of a four-panel comic strip.

[1378] Clicking on the thumbnail will take you to the news article details page.

[1379] Users watch four-frame videos on YouTube or Twitter and click on a link in the video description to access a news site.

[1380] Step 8:

[1381] The server uses an emotion engine to recognize the user's emotions. It analyzes the user's camera footage and voice input to recognize the user's emotional state.

[1382] Specific behavior:

[1383] When a user accesses a news site, the server collects camera footage and audio data.

[1384] The server calls the emotion engine and inputs the collected video and audio data.

[1385] The emotion engine analyzes the user's emotions (e.g., "happy" or "excited") and returns the results to the server.

[1386] Step 9:

[1387] The server customizes the content of the four-panel comic or video based on the recognized emotion.

[1388] Specific behavior:

[1389] The server adjusts the content of the four-panel comics and videos based on the emotional data obtained from the emotion engine.

[1390] For example, if the user is recognized as "happy," a four-panel comic strip with positive content is generated.

[1391] Similarly, the content and order of the four-frame video are changed according to the user's emotions.

[1392] The above are the specific processing steps of the present invention that incorporate emotion recognition, which enable providing users with personalized visual news content.

[1393] Example 2

[1394] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1395] In recent years, the distribution methods of news articles have become more diverse, making it increasingly difficult to keep readers interested. Furthermore, if the content of the article does not match the reader's emotions, it is difficult to maintain their interest. Therefore, there is a need for a method to provide news articles in a more user-friendly format and further customize the content based on the user's emotions.

[1396] The specific processing by the specific processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for acquiring news articles, means for analyzing the acquired news articles and extracting important data and key points, means for generating a four-panel cartoon scenario using a generative AI model based on the extracted data and key points, means for generating a four-panel cartoon using an image generation AI based on the generated scenario, means for generating animation based on the generated four-panel cartoon and creating a four-panel video, means for displaying the generated four-panel cartoon and four-panel video, means for posting the generated four-panel video to an information sharing platform or a social networking service to promote access to the original article, and means for recognizing user emotions and customizing content based on the user's emotions. This makes it possible to provide news articles in familiar four-panel cartoon or video formats and further provide personalized content tailored to the user's emotions.

[1397] A "news article" is a report of a fact or event provided online or in print.

[1398] A "server" is a computer or program that processes data on a network and provides services to other computers.

[1399] "Means of acquisition" are the methods and techniques used to collect specific data and store it in a usable format.

[1400] "Means of analysis" are methods and techniques for analyzing given data and understanding its meaning and structure.

[1401] "Important data" is information, such as news articles, that is deemed to be particularly relevant and valuable.

[1402] "Gist" refers to the most important or core part of information or data.

[1403] A "generative AI model" is an algorithm or program that uses artificial intelligence technology to generate new data and information.

[1404] "Scenario generation means" refers to methods and technologies for automatically creating stories and structures based on basic data.

[1405] "Image generation AI" refers to algorithms and programs that use artificial intelligence technology to generate image data based on input information.

[1406] "Four-panel comics" are a form of comic expression that consists of four images or illustrations arranged as a continuous story.

[1407] "Means for generating animation" refers to methods or techniques for expressing movement by displaying still images or illustrations in succession.

[1408] "4-frame video" is a video format that displays four scenes in succession and adds movement and effects.

[1409] The "display means" refers to a method or technology for providing the generated content in a format that can be visually confirmed by the user.

[1410] An "information sharing platform" is a website or application that allows users to share content or information with other users over the Internet.

[1411] A "social networking service (SNS)" is a service that allows users to connect with each other online and communicate and share information.

[1412] "Original article" refers to the original news article collected by the news article acquisition means and used for analysis and transformation.

[1413] "Means for recognizing emotions" refers to methods or technologies for identifying a user's emotional state by analyzing their facial expressions and voice.

[1414] "Customization means" are methods and techniques for modifying and tailoring content to a user's specific needs or circumstances.

[1415] overview

[1416] This invention is a system that converts news articles into four-panel comics and four-panel videos and customizes the content based on the user's emotions, making it possible to provide individual news content that matches the user's interests.

[1417] System Configuration

[1418] The system consists of the following main modules:

[1419] 1. News article acquisition module

[1420] 2. Text Analysis Module

[1421] 3. Four-panel comic generation module

[1422] 4. 4-frame video generation module

[1423] 5. Content Display and Sharing Module

[1424] 6. Emotion Recognition and Response Module

[1425] News article acquisition module

[1426] The server retrieves the latest news articles using the news site's API or RSS feed, which allows the latest news to be stored in the database at all times.

[1427] Examples:

[1428] The server accesses the "News API" to retrieve the latest news articles, then stores the article title, body text, and URL in a database.

[1429] Text Analysis Module

[1430] The server analyzes the text of the retrieved news articles and uses natural language processing (NLP) techniques, such as morphological analysis and topic modeling, to extract key data and key points.

[1431] Examples:

[1432] The server analyzes news articles and extracts the keywords "climate change," "abnormal weather," and "impact."

[1433] Four-panel comic generation module

[1434] The server uses a generative AI model to automatically generate a four-panel comic scenario based on the extracted data and key points. Based on this scenario, an image generation AI generates illustrations for each panel.

[1435] Examples:

[1436] Prompt: "A character who thinks about climate change"

[1437] 1. Frame 1: A scene where a character thinks about "climate change"

[1438] 2. Frame 2: A scene where abnormal weather is occurring

[1439] 3. Panel 3: People are surprised by the abnormal weather.

[1440] 4. Panel 4: A scene of characters thinking about climate change countermeasures

[1441] The server inputs a prompt sentence into the generative AI model, and passes the output scenario to the image generation AI to generate an illustration.

[1442] 4-frame video generation module

[1443] The server generates animations based on the generated four-panel comic illustrations, creating a four-panel video. Each still image is displayed in succession, and a short video is created by adding movement.

[1444] Examples:

[1445] The server uses animation production software (e.g., Adobe After Effects) to place the illustrations for each frame on a timeline, add transitions and effects, and generate the video.

[1446] Content Viewing and Sharing Module

[1447] The server places the generated four-panel comics and four-panel videos on a news site and also posts them on video sharing platforms and social networking services (SNS).

[1448] Examples:

[1449] The server uses HTML and JavaScript to display four-panel comics on the news site's homepage, and posts four-panel videos using YouTube and Twitter APIs.

[1450] Emotion Recognition and Response Module

[1451] The server uses an emotion engine to recognize the user's emotions. It analyzes the user's camera footage and voice input to recognize their emotional state. The content of the four-panel comics and videos is customized based on the recognized emotions.

[1452] Examples:

[1453] The server uses emotion recognition software (e.g., Microsoft Azure Cognitive Services) to analyze the user's camera footage, and if it recognizes the emotion as "happy," it generates and displays a four-panel comic strip with a positive message.

[1454] Example

[1455] When a user accesses a news site, their emotions are recognized through a camera or voice input device. Based on the recognized emotions, a four-panel comic strip or video that best suits the user is generated and displayed. Users can click on this content to read detailed news articles. They can also access the news site from a four-panel video they have viewed on a social media or video sharing platform.

[1456] Examples:

[1457] A user visits a news site and is identified as "excited" from five minutes of webcam footage. Based on this, the server generates a four-panel comic strip of the active content and displays it on the news site's homepage. The user clicks on the thumbnail of the four-panel comic strip that interests them and reads the detailed news article.

[1458] The above is a specific embodiment of the invention, and by combining it with emotion recognition, it is possible to provide personalized news content to users, which is more likely to attract users' attention and increase usage frequency compared to conventional news distribution methods.

[1459] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1460] Step 1: Get news articles

[1461] The server retrieves the latest news articles using the news site's API or RSS feed. The input is the news site's URL or API endpoint, and the output is the news article data. Specifically, the server sends an HTTP request, extracts the title, body, and URL of the news article from the received response, and stores them in a database.

[1462] Step 2: Analyzing the news article

[1463] The server retrieves news articles stored in a database and analyzes the text using natural language processing (NLP) technology. The input is the text data of the news article, and the output is extracted keywords and key points. Specifically, the server uses a morphological analysis tool to segment the text and performs topic modeling to extract important keywords.

[1464] Step 3: Creating a four-panel comic scenario

[1465] The server uses a generative AI model based on the extracted keywords and key points to automatically generate a four-panel comic scenario. The input is the extracted keywords, and the output is the four-panel comic scenario. Specifically, the server inputs a prompt statement (e.g., "A character thinking about climate change") into the generative AI model and retrieves the generated scenario.

[1466] Step 4: Generate a four-panel comic

[1467] The server uses image generation AI to generate illustrations for a four-panel comic based on the generated scenario. The input is the generated scenario, and the output is the illustrations for each panel. Specifically, the server passes the scenario for each panel to the image generation AI, which then generates the illustrations.

[1468] Step 5: Generate a 4-frame video

[1469] The server generates animation based on the generated four-panel cartoon illustrations, creating a four-panel video. The input is a four-panel cartoon illustration, and the output is a four-panel video. Specifically, the server uses animation production software to place the illustrations of each frame on a timeline, add effects and transitions, and create a short video.

[1470] Step 6: View and share content

[1471] The server places the generated four-panel comics and videos on news sites and also posts them on video sharing platforms and social networking services. The input is the generated four-panel comics and videos, and the output is public information on the news sites and sharing platforms. Specifically, the server updates the HTML and JavaScript on the news sites and posts the videos using the YouTube and Twitter APIs.

[1472] Step 7: User Emotion Recognition and Customization

[1473] The server uses an emotion engine to recognize the user's emotions. The input is the user's camera footage and audio input, and the output is the recognized emotional state. Specifically, the server uses emotion recognition software to analyze the camera footage and identify the user's emotions. Based on the emotional state, the content of the four-panel comic or video is customized.

[1474] By performing the above processing steps, news articles can be presented in the form of familiar four-panel comic strips or videos, and personalized content tailored to the user's emotions can be provided.

[1475] (Application example 2)

[1476] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1477] Modern news distribution services only provide information one-way to users, and have the problem of being unable to provide personalized content that reflects users' emotions and interests. Furthermore, they lack the visual and entertaining format to keep users engaged with the news. As a result, it is difficult to maintain users' interest, leading to a decline in frequency of use.

[1478] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1479] In this invention, the server includes: means for acquiring news articles; means for analyzing the acquired news articles and extracting important keywords and key points; means for generating a four-panel cartoon scenario using a generation AI based on the extracted keywords and key points; means for generating a four-panel cartoon using an image generation AI based on the generated scenario; means for generating an animation based on the generated four-panel cartoon and creating a four-panel video; means for displaying the generated four-panel cartoon and four-panel video; means for posting the generated four-panel video on a video sharing platform or a social networking service to promote access to the original article; means for recognizing a user's emotions using a camera or a voice input device; means for customizing the content of the generated four-panel cartoon and four-panel video based on the recognized emotion; and means for displaying the customized four-panel cartoon and four-panel video within the smartphone app. This allows news articles to be provided in a form adapted to the user's emotions, thereby increasing user interest and increasing usage frequency.

[1480] "Means for obtaining news articles" is a function for automatically collecting the latest news articles from news sites and APIs on the Internet.

[1481] "Means for analyzing news articles and extracting important keywords and key points" refers to technology for analyzing the content of acquired news articles and automatically extracting the main topics and important information of the articles.

[1482] "Generative AI" is a general term for artificial intelligence that uses natural language processing and machine learning techniques to generate stories and scenarios from input data.

[1483] "Image generation AI" is artificial intelligence that creates visual representations based on given scenarios or prompts.

[1484] The "means for generating four-panel comics" is a technology that automatically creates a comic consisting of four consecutive panels based on a generated scenario.

[1485] "Means for generating animations and creating four-frame videos" refers to a technology that converts the generated four-frame comics into moving animations and provides them in the form of short videos.

[1486] "Means for displaying the generated four-panel comics and four-panel videos" refers to technology for displaying the generated content in a manner that allows users to visually access it.

[1487] "A means of posting the generated four-frame video on video sharing platforms and social networking services to promote access to the original article" refers to a function that posts the generated video on YouTube or social media, directing viewers to the original news article.

[1488] "Means for recognizing a user's emotions using a camera or voice input device" refers to technology that analyzes camera footage and voice data to detect the user's emotional state.

[1489] "Means for customizing the content of generated four-panel comics and four-panel videos based on recognized emotions" refers to technology that adjusts the content and tone of generated content based on the user's emotional data.

[1490] "Means for displaying customized four-panel comics and four-panel videos within a smartphone app" is a function for providing users with emotionally tailored content through a smartphone app.

[1491] A specific embodiment for carrying out the present invention will be described.

[1492] Hardware and software used

[1493] The server uses the following major hardware and software:

[1494] Hardware

[1495] Internet connection required to retrieve news data

[1496] Camera and microphone for recognizing user emotions

[1497] Display devices such as smartphones and tablets

[1498] software

[1499] News Acquisition API

[1500] Natural language processing libraries (such as TextBlob)

[1501] Image generation AI (general image generation algorithms)

[1502] Video Generation Algorithm

[1503] Emotion recognition software (OpenCV and other facial recognition technologies)

[1504] Social media APIs (YouTube API, Twitter API, etc.)

[1505] Data processing and calculation flow

[1506] The server retrieves the latest news articles from the Internet, using a news API to pull news data into the server, and then analyzes the text data of the news articles using natural language processing technology to extract important keywords and key points.

[1507] Based on the extracted keywords and points, a scenario for a four-panel comic is automatically generated by the generative AI. This scenario is passed to the image generation AI used to generate a comic illustration consisting of four consecutive frames. The generated four-panel comic is then animated and presented as a short video.

[1508] The generated comic strips and videos are displayed on smartphone apps and other devices. When a user accesses the app, the app recognizes the user's emotions in real time via the camera and microphone, and the content of the comic strips and videos generated is customized based on those emotions. These contents are also posted on social networking services and video sharing platforms, and are used as a means to direct users to the original news article.

[1509] Specific examples

[1510] For example, if a news article mentions "climate change," the system might:

[1511] News acquisition

[1512] Use the News API to get the latest "Climate Change" articles.

[1513] Text analytics

[1514] Use a natural language processing library such as TextBlob to extract keywords such as "climate change," "extreme weather," and "impact" from articles.

[1515] Scenario generation and image generation

[1516] A scenario is generated based on the extracted keywords, and the image generation AI is used to draw the following four-panel comic scenario:

[1517] 1. A scene where a character thinks about climate change

[1518] 2. Scenes where abnormal weather is occurring

[1519] 3. Scenes of people being surprised by extreme weather

[1520] 4. A scene with characters thinking about climate change

[1521] Based on these scenarios, we generate prompts and pass them to the image generation AI:

[1522] "Generate a frame of a character talking at length about climate change."

[1523] "Draw a scene where extreme weather is occurring."

[1524] "Draw a frame showing people being surprised by the extreme weather."

[1525] Video Generation and Display

[1526] An animation is created based on the generated four-panel comic and converted into a short video format.

[1527] The video and four-panel comics are displayed within a smartphone app. When a user accesses the app, the camera and microphone are used to recognize their emotions, and content is customized accordingly.

[1528] Social Media Posts

[1529] The resulting four-frame video can then be posted on social media and video sharing platforms for a wider audience.

[1530] In this way, news articles can be provided in a form that is adapted to the user's emotions, which can increase the user's interest and increase the frequency of use.

[1531] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1532] Step 1:

[1533] The server uses a news API to retrieve the latest news articles from the Internet. The retrieved article data includes the article title, body text, URL, etc. The input is the response from the news API, and the output is text data for analysis.

[1534] Step 2:

[1535] The server analyzes the text data of the acquired news articles using natural language processing (NLP) techniques. Specifically, it uses libraries such as TextBlob to extract important keywords and key points. The input for this process is the text data, and the output is the extracted keywords and key points.

[1536] Step 3:

[1537] The server uses a generative AI model to generate a four-panel comic scenario based on the extracted keywords and points. Here, a group of keywords is passed as input, and the output is four prompt sentences containing the scenario for each panel. Specifically, the server converts the keywords into prompt sentences and passes them to the generative AI.

[1538] Step 4:

[1539] The server generates a four-panel comic using an image generation AI based on the generated scenario. At this stage, the input is a prompt sentence, and the output is four consecutive comic illustrations. Specifically, the prompt is passed to the image generation AI based on the scenario output from the generative AI model.

[1540] Step 5:

[1541] The server generates animation based on the generated four-panel comic to create a four-panel video. Here, still images of the four-panel comic are passed as input, and the output is a short video with continuous movements. Specifically, the video is generated by connecting still images and adding movement.

[1542] Step 6:

[1543] The server displays the generated four-panel comics and four-panel videos on a smartphone app or other device. The input is the generated four-panel comics and videos, and the output is visual data that is displayed to the user. Specifically, it includes the operation of displaying them in a specific section within the smartphone app.

[1544] Step 7:

[1545] When a user accesses the smartphone app, the app uses the device's camera and microphone to recognize emotions in real time. The input is the user's real-time video and audio data, and the output is analyzed emotional data. Specifically, the app uses emotion recognition software to analyze facial expressions and tone of voice.

[1546] Step 8:

[1547] The server customizes the content of the generated four-panel comics and four-panel videos based on the recognized emotions. The input is emotion data, and the output is customized four-panel comics and videos. Specific operations include reselecting scenarios and images according to the emotion.

[1548] Step 9:

[1549] The server posts customized four-panel comics and videos to social networking services and video sharing platforms to promote access to the original article. The input is the customized video and comic, and the output is a post on social media. Specifically, it automates the posting using the platform's API.

[1550] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1551] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1552] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1553] [Fourth embodiment]

[1554] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1555] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1556] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1557] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1558] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1559] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1560] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1561] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1562] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1563] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1564] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1565] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1566] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1567] overview

[1568] This invention is a system that converts news articles into visually intuitive four-panel comics and four-panel videos. This system utilizes generative AI and image generation AI technologies to extract the key points of news articles and present them in an easy-to-understand format for users.

[1569] System Configuration

[1570] The system of the present invention consists of the following main parts:

[1571] 1. News article acquisition module

[1572] 2. Text Analysis Module

[1573] 3. Four-panel comic generation module

[1574] 4. 4-frame video generation module

[1575] 5. Content Display and Sharing Module

[1576] News article acquisition module

[1577] The server retrieves the latest news articles through the news site's API or RSS feed. This retrieval operation can be set to occur periodically, and the latest articles are automatically imported into the system. The retrieved news articles are stored in a database.

[1578] Examples:

[1579] The server accesses the "News API" to retrieve the title, body text, and URL of the latest news article, then stores this information in a database.

[1580] Text Analysis Module

[1581] The server uses natural language processing techniques to analyze the text of the retrieved news articles and extract key keywords and points, a process that clarifies the core of the article.

[1582] Examples:

[1583] The server analyzes the news article "Impacts of Climate Change" and extracts the main keywords "climate change," "abnormal weather," and "impact."

[1584] Four-panel comic generation module

[1585] The server uses a generation AI to create a four-panel comic scenario based on the extracted keywords and points. This scenario includes the situation and storyline of each panel. The generated scenario is then passed to an image generation AI, which generates illustrations for each panel.

[1586] Examples:

[1587] The server creates the scenario as follows:

[1588] 1. Frame 1: A character is thinking about "climate change."

[1589] 2. Frame 2: A scene where abnormal weather is occurring

[1590] 3. Panel 3: People are surprised by the abnormal weather.

[1591] 4. Panel 4: A scene of characters thinking about climate change countermeasures

[1592] Based on this, the image generation AI generates four illustrations.

[1593] 4-frame video generation module

[1594] The server then creates an animation based on the four-panel comic, generating a four-panel video. During this process, the still images in each frame are animated and presented as a continuous video.

[1595] Examples:

[1596] The server breaks down a four-panel comic and animates each scene. For example, it animates the character's thoughts in the first panel, then moves to show the occurrence of extreme weather.

[1597] Content Viewing and Sharing Module

[1598] The server displays the generated four-panel comics and videos on the homepage of the news site, and also posts these videos on video sharing platforms and social networking services to promote access to the news articles.

[1599] Examples:

[1600] The server places the generated four-panel comic on the homepage of the news site, and then uploads the four-panel video to YouTube and Twitter, with a link to the article.

[1601] Example

[1602] When users visit a news site, they can click on a thumbnail of a four-panel comic that appears on the homepage to read a detailed news article. Users who are interested in a four-panel comic that has been shared on a video sharing platform or social networking service can click on the link to access the original article.

[1603] example:

[1604] A user clicks on a four-panel cartoon on the homepage of a news site to read a detailed article about the impacts of climate change. They also watch a four-panel video on YouTube and access the news site via a link in the video description.

[1605] As described above, the system of the present invention converts news articles into a visual and intuitive format, attracting the interest of the digital native generation and promoting access to news sites.

[1606] The processing flow will be explained below.

[1607] Step 1:

[1608] The server retrieves the latest news articles using the news site's API or RSS feed. It periodically sends requests using the API key to collect news data. The retrieved news articles are then stored in a database.

[1609] Specific behavior:

[1610] The server accesses the "News API" and sends a request to retrieve the latest news articles.

[1611] The title, text, and URL of the news article received as a response are extracted and stored in a database.

[1612] Step 2:

[1613] The server invokes a natural language processing (NLP) engine to analyze the text of the stored news articles, thereby extracting important keywords and key points.

[1614] Specific behavior:

[1615] The server inputs the text of the news article into the NLP engine.

[1616] The NLP engine analyzes the text and extracts key keywords (e.g., "climate change," "extreme weather," "impact") and key points.

[1617] The extracted keywords and points are stored in a database.

[1618] Step 3:

[1619] The server uses generative AI to generate a four-panel comic scenario based on the extracted keywords and points. The generated scenario defines the situation and storyline of each panel.

[1620] Specific behavior:

[1621] The server inputs keywords and points into the generation AI and requests it to generate a four-panel scenario.

[1622] The generation AI creates a scenario and returns a description of each frame (e.g., "Frame 1: A character thinks about climate change," "Frame 2: Extreme weather occurs," etc.).

[1623] The generated scenario is saved in the database.

[1624] Step 4:

[1625] The server uses image generation AI to create a four-panel comic based on the generated scenario. The image generation AI generates illustrations based on the text description of each panel.

[1626] Specific behavior:

[1627] The server inputs the scenario text into the image generation AI and requests it to generate an illustration.

[1628] The image generation AI generates an illustration corresponding to each frame and returns it to the server.

[1629] The generated illustrations are arranged in the form of a four-panel comic.

[1630] Step 5:

[1631] The server generates animations based on each frame of the four-panel comic, creating a four-panel video. It displays still images in succession and adds movement to each scene.

[1632] Specific behavior:

[1633] The server provides the AI ​​with a script to animate the illustrations in each panel of the four-panel comic.

[1634] The generative AI generates the movement of each frame and creates a video file as an animation.

[1635] The generated video file is saved on the server.

[1636] Step 6:

[1637] The server places the generated four-panel comics and videos on the front page of the news site, and also posts them on video sharing platforms and social networking services to promote access.

[1638] Specific behavior:

[1639] The server generates HTML code to embed the four-panel comic on the top page of the news site.

[1640] The generated four-panel comic is displayed on the top page.

[1641] Upload your four-frame video to platforms such as YouTube and Twitter.

[1642] Link to the news article in the video description.

[1643] Step 7:

[1644] Users can click on thumbnails of four-panel comics or videos displayed on the homepage of news sites to read detailed news articles, or watch videos on social media or video-sharing platforms and access the original articles.

[1645] Specific behavior:

[1646] A user visits the homepage of a news site and clicks on a thumbnail of a four-panel comic strip.

[1647] Clicking on the thumbnail will take you to the news article details page.

[1648] Users watch four-frame videos on YouTube or Twitter and click on a link in the video description to access a news site.

[1649] The above are the specific processing steps of the present invention. By executing these steps, it is possible to provide visual and intuitive news content to the digital native generation and promote access to news sites.

[1650] Example 1

[1651] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1652] Traditional news content is often expressed in text format, making it difficult to provide a visual and intuitive understanding, especially for the digital native generation. Furthermore, the process of visualizing and animating news articles is time-consuming, and automation is needed.

[1653] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1654] In this invention, the server includes a means for acquiring news content, a means for analyzing the acquired news content and extracting important keywords and information, and a means for generating a scenario for story-style content using a generative AI model based on the extracted keywords and information, thereby automating the task of visually and intuitively presenting news articles.

[1655] "News content" refers to information such as articles and reports obtained from news sites and other information platforms.

[1656] "Important keywords and information" refers to content and key points that are particularly noteworthy extracted from news content.

[1657] A "generative AI model" refers to an artificial intelligence model that understands natural language and generates new sentences and scenarios based on that information.

[1658] "Scenario for narrative content" refers to a storyline for constructing a continuous narrative, such as a four-panel comic strip, generated based on keywords and information extracted from news content.

[1659] An "image generation model" refers to an artificial intelligence model for generating images based on text information.

[1660] "Visual content" refers to visually presented information, such as four-panel comic strips, created using generative AI models and image generation models.

[1661] "Animation" refers to moving image content created by sequentially displaying visual content and adding movement.

[1662] "Video Content" refers to continuous, moving images generated from visual content.

[1663] A "video sharing platform" refers to a service that allows users to share videos over the Internet, such as YouTube.

[1664] "Social media" refers to online platforms such as Twitter and Facebook that allow users to communicate with each other and share content.

[1665] overview

[1666] This invention is a system that converts news content into visually intuitive four-panel comics and four-panel videos. This system utilizes generative AI and image generation model technologies to extract key points from news content and present them in an easy-to-understand format for users.

[1667] System Configuration

[1668] The system of the present invention consists of the following main modules:

[1669] 1. News article acquisition module

[1670] 2. Text Analysis Module

[1671] 3. Four-panel comic generation module

[1672] 4. 4-frame video generation module

[1673] 5. Content Display and Sharing Module

[1674] News article acquisition module

[1675] The server retrieves the latest news content using the news site's API or RSS feed. This retrieval operation can be performed periodically, and the latest news content is automatically imported into the system. The retrieved content is then stored in a database.

[1676] Examples:

[1677] The server accesses the "News API" to retrieve the title, body text, and URL of the latest news article, then stores this information in a database.

[1678] Text Analysis Module

[1679] The server uses natural language processing techniques to analyze the text of the retrieved news content and extract key keywords and information, thereby clarifying the core of the article.

[1680] Examples:

[1681] The server analyzes the news article "Impacts of Climate Change" and extracts the main keywords "climate change," "abnormal weather," and "impact."

[1682] Four-panel comic generation module

[1683] The server uses a generative AI model to create a four-panel comic scenario based on the extracted keywords and information. This scenario includes the situation and storyline of each panel. The generated scenario is then passed to an image generation model, which generates illustrations for each panel.

[1684] Examples:

[1685] The server creates the scenario as follows:

[1686] 1. Frame 1: A character is thinking about "climate change."

[1687] 2. Frame 2: A scene where abnormal weather is occurring

[1688] 3. Panel 3: People are surprised by the abnormal weather.

[1689] 4. Panel 4: A scene of characters thinking about climate change countermeasures

[1690] Based on this, an image generation model generates four illustrations.

[1691] 4-frame video generation module

[1692] The server then creates an animation based on the four-panel comic, generating a four-panel video. During this process, the still images in each frame are animated and presented as a continuous video.

[1693] Examples:

[1694] The server breaks down a four-panel comic and animates each scene. For example, it animates the character's thoughts in the first panel, then moves to show the occurrence of extreme weather.

[1695] Content Viewing and Sharing Module

[1696] The server displays the generated four-panel comics and videos on the homepage of the news site, and also posts these videos on video sharing platforms and social media to promote access to the news content.

[1697] Examples:

[1698] The server places the generated four-panel comic on the homepage of the news site, and also uploads the four-panel video to YouTube and Twitter, with a link to the article.

[1699] Examples of prompt statements

[1700] Below are some example prompts for using generative AI models:

[1701] Example prompt sentence:

[1702] 1. Prompts to extract necessary keywords from news articles:

[1703] Extract key keywords from the news article "Impacts of Climate Change."

[1704] 2. Prompts for generating a four-panel comic scenario:

[1705] Please create a four-panel comic scenario based on the extracted keywords "climate change," "extreme weather," and "impact." Please be specific and include a storyline for each panel.

[1706] 3. Prompt to pass to the image generation model:

[1707] Please generate an illustration for each frame based on the following scenario.

[1708] 1. Frame 1: A character is thinking about "climate change."

[1709] 2. Frame 2: A scene where abnormal weather is occurring

[1710] 3. Panel 3: People are surprised by the abnormal weather.

[1711] 4. Panel 4: A scene of characters thinking about climate change countermeasures

[1712] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1713] Step 1:

[1714] The server retrieves news content. It uses the news site's API or RSS feed as input. The output is news article data (title, body text, URL, etc.), which is stored in a database.

[1715] Specific behavior:

[1716] The server sends a request to the "News API" to receive the latest news article data. The received data (e.g., article data on the impact of climate change) is stored in the corresponding table in the database.

[1717] Step 2:

[1718] The server reads news content from the database and performs text analysis. The input is the text data of the news article obtained in step 1, and the output is a list of important keywords and information.

[1719] Specific behavior:

[1720] The server uses a natural language processing library (e.g., SpaCy or NLTK) to analyze the text of a news article (e.g., "impacts of climate change") and extracts the main keywords "climate change," "extreme weather," and "impact."

[1721] Step 3:

[1722] The server uses a generative AI model based on the extracted keywords and information to create a four-panel comic scenario. The input is the keywords and information extracted in step 2, and the output is the four-panel comic scenario.

[1723] Specific behavior:

[1724] The server sends a prompt to the generation AI model, for example, "Create a four-panel comic scenario based on the following keywords: Keywords: climate change, extreme weather, impact," and generates a scenario.

[1725] Example scenario:

[1726] 1. Frame 1: A character is thinking about "climate change."

[1727] 2. Frame 2: A scene where abnormal weather is occurring

[1728] 3. Panel 3: People are surprised by the abnormal weather.

[1729] 4. Panel 4: A scene of characters thinking about climate change countermeasures

[1730] Step 4:

[1731] Based on the generated scenario, the server uses an image generation model to generate illustrations for each frame of the four-panel comic. The input is the scenario created in step 3, and the output is a visual illustration for each frame.

[1732] Specific behavior:

[1733] The server sends prompts to the image generation model (e.g., DALL-E). For example, it uses the image generation prompt for the first frame, "A scene in which a character is thinking about climate change," to generate illustrations for each frame.

[1734] Step 5:

[1735] The server creates an animation based on the generated four-panel comic and generates a four-panel video. The input is the visual illustration generated in step 4, and the output is a four-panel video.

[1736] Specific behavior:

[1737] The server uses video editing software or libraries (e.g., OpenCV) to animate the images in each frame. For example, you could animate a character thinking in the first frame, then smoothly transition to an extreme weather scene.

[1738] Step 6:

[1739] The server displays the generated comic strips and videos on the homepage of the news site, and also posts these videos to video sharing platforms and social media. The input is the comic strips and videos generated in step 5, and the output is the web page where the information is displayed and the posted content.

[1740] Specific behavior:

[1741] The server places thumbnails of the four-panel comics on the homepage of the news site, and when users click on them, detailed news content is displayed. The server also uploads the four-panel comics to YouTube and Twitter, automatically linking them to the articles.

[1742] (Application example 1)

[1743] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1744] News articles are primarily text-based, making it difficult to understand their content quickly. Furthermore, digital natives, in particular, are demanding content in a visual, intuitive format. To meet these modern needs, a new way of presenting news articles in a visual, easy-to-understand format is needed.

[1745] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1746] In this invention, the server includes means for acquiring news articles, means for analyzing the acquired news articles and extracting important keywords and key points, means for generating a four-panel cartoon scenario using a generation AI based on the extracted keywords and key points, means for generating a four-panel cartoon using an image generation AI based on the generated scenario, means for generating animation based on the generated four-panel cartoon and creating a four-panel video, means for displaying the generated four-panel cartoon and four-panel video, means for posting the generated four-panel video on a video sharing platform or social networking service to promote access to the original article, and means for displaying and sharing the generated four-panel cartoon and four-panel video on a smart device, thereby enabling users to visually understand the content of the news article in a short amount of time.

[1747] A "means for acquiring news articles" is a system or module for automatically acquiring the latest news articles from news sites or RSS feeds on the Internet.

[1748] "Means for analyzing acquired news articles and extracting important keywords and key points" refers to a system or module that has the function of analyzing the text of news articles using natural language processing technology and extracting the subject matter and important information of the articles.

[1749] "Means for generating a four-panel comic scenario using generative AI" refers to a system or module that uses generative AI to construct a storyline for a four-panel comic and the content of each panel based on keywords and key points extracted from news articles.

[1750] "Means for generating four-panel comics using image generation AI" refers to a system or module that uses image generation AI to create illustrations for four-panel comics based on a generated scenario.

[1751] "Means for generating animation based on the generated four-panel comic and creating a four-panel video" refers to a system or module that animates the still images of each frame and presents them as a continuous four-panel video.

[1752] "Means for displaying the generated four-panel comics and four-panel videos" refers to a system or module for displaying the generated four-panel comics and four-panel videos on a user interface.

[1753] "Means for posting the generated four-panel videos on video sharing platforms or social networking services to promote access to the original article" refers to a system or module that uploads the generated four-panel videos to video sharing platforms or social networking services and provides users with a link that allows them to access the original news article.

[1754] "Means for displaying and sharing generated four-panel comics and four-panel videos on a smart device" refers to a system or module that allows users to display four-panel comics and four-panel videos generated on a smart device such as a smartphone or tablet, and share the content with other users via a link or share button.

[1755] "Natural language processing technology" is a computer technology for analyzing text data and extracting or summarizing important information.

[1756] A "content distribution service" is an online service that provides users with various types of digital content via the Internet.

[1757] Overall system overview

[1758] This system aims to provide news articles as visually intuitive four-panel comics and four-panel videos. Specifically, it supports processes including news article acquisition, analysis, scenario creation using generative AI, comic generation using image generation AI, animation, display, and sharing. The main components of this system are a news article acquisition module, text analysis module, four-panel comic generation module, four-panel video generation module, and content display and sharing module.

[1759] Hardware and Software

[1760] The hardware used includes servers and smart devices (e.g., smartphones). The main software includes:

[1761] Python 3.x: Programming Language

[1762] Requests: A library for making HTTP requests

[1763] BeautifulSoup: HTML parsing library

[1764] Spacy: A natural language processing library

[1765] Transformers: A library for using generative AI

[1766] Image generation API (hypothetical API)

[1767] Processing flow

[1768] The server first retrieves news articles from news sites and RSS feeds. Next, it analyzes the retrieved news articles and extracts important keywords and key points using natural language processing technology. Based on the extracted keywords and key points, it uses generative AI to generate a four-panel comic scenario. It then uses image generation AI to create four-panel comic illustrations based on the scenario. It then animates the generated four-panel comic to create a four-panel video. Finally, the generated four-panel comic and video can be displayed and shared on a smart device.

[1769] Specific examples

[1770] As a concrete example, we take the news article "The spread of the new coronavirus" and extract the key keywords "new coronavirus," "social distancing," and "vaccine." The following scenario is generated:

[1771] 1. A scene where a character is thinking about the new coronavirus

[1772] 2. Scenes where social distancing is being implemented

[1773] 3. Scenes of people maintaining social distancing

[1774] 4. A scene where a character thinks about getting vaccinated

[1775] An example of a prompt passed to the generative AI model is:

[1776] Prompt: Transform the contents of a news article into a four-panel comic book scenario.

[1777] Key Word 1: Novel Coronavirus

[1778] Key word 2: Social distancing

[1779] Key word 3: Vaccine

[1780] The generated four-panel comics and videos can be displayed on smart devices such as smartphones and can be easily shared via social media, allowing users to quickly and visually understand the content of news articles.

[1781] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1782] Step 1:

[1783] The server periodically retrieves news articles from news sites and RSS feeds on the Internet. Specifically, the server accesses a news API, retrieves the title, body, and URL of the latest news article, and stores them in a database. The URL of the news API is provided as input, and the JSON data of the news article is obtained as output.

[1784] Step 2:

[1785] The server analyzes the retrieved news articles and extracts important keywords and key points. In this process, it uses natural language processing technology (Spacy) to analyze the article text and extract key keywords and noun phrases. The body of the news article is provided as input, and a list of important keywords is obtained as output.

[1786] Step 3:

[1787] The server uses a generative AI model to generate a four-panel comic scenario based on the extracted keywords and key points. Specifically, a prompt is input to the generative AI, which outputs four scenario blocks. The extracted keywords and prompt are provided as input, and the scenario for each panel is generated as output.

[1788] Step 4:

[1789] The server uses image generation AI to generate a four-panel comic based on the generated scenario. The generated scenario is passed to an image generation API, which generates illustrations for each panel. The scenario is provided as input, and four illustrations are generated as output.

[1790] Step 5:

[1791] The server generates an animation based on the generated four-panel comic, creating a four-panel video. The illustrations for each frame are imported into animation software (e.g., Adobe After Effects) to create a continuous animation. Four illustrations are provided as input, and a four-panel video is generated as output.

[1792] Step 6:

[1793] The server displays the generated four-panel comic and four-panel video. In this step, the generated content is displayed on a user interface so that the user can view it. The four-panel comic and four-panel video are provided as input, and are displayed on the user interface as output.

[1794] Step 7:

[1795] The server posts the generated four-frame comics to video sharing platforms and social networking services to promote access to the original article. Specifically, it uploads the videos using APIs such as YouTube and Twitter and sets up links to the articles. The four-frame comics and the URLs of the original articles are provided as input, and the videos are published on the video platforms as output.

[1796] Step 8:

[1797] The server displays and shares the generated four-panel comics and four-panel videos on smart devices. Users view the generated content on their smartphones or tablets and click share links on social media to share the content with other users. Four-panel comics and four-panel videos are provided as input, and can be displayed and shared on smart devices as output.

[1798] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1799] overview

[1800] This invention combines an emotion engine with a system that converts news articles into four-panel comics and videos to customize the content based on the user's emotions, making it possible to provide individual news content that matches the user's interests and emotions.

[1801] System Configuration

[1802] The system of the present invention consists of the following main modules:

[1803] 1. News article acquisition module

[1804] 2. Text Analysis Module

[1805] 3. Four-panel comic generation module

[1806] 4. 4-frame video generation module

[1807] 5. Content Display and Sharing Module

[1808] 6. Emotion Recognition and Response Module

[1809] News article acquisition module

[1810] The server retrieves the latest news articles using the news site's API or RSS feed, and stores the retrieved news articles in a database.

[1811] Examples:

[1812] The server accesses the "News API" to retrieve the latest news articles, then stores the article title, text, and URL in a database.

[1813] Text Analysis Module

[1814] The server analyzes the text of the retrieved news articles and utilizes natural language processing (NLP) techniques to extract important keywords and key points.

[1815] Examples:

[1816] The server analyzes news articles and extracts the keywords "climate change," "abnormal weather," and "impact."

[1817] Four-panel comic generation module

[1818] The server uses a generation AI to automatically generate a four-panel comic scenario based on the extracted keywords and key points. This scenario includes the situation and storyline of each panel. The generated scenario is then passed to an image generation AI, which generates illustrations for each panel.

[1819] Examples:

[1820] The server creates the scenario as follows:

[1821] 1. Frame 1: A scene where a character thinks about "climate change"

[1822] 2. Frame 2: A scene where abnormal weather is occurring

[1823] 3. Panel 3: People are surprised by the abnormal weather.

[1824] 4. Panel 4: A scene of characters thinking about climate change countermeasures

[1825] Based on this, the image generation AI generates four illustrations.

[1826] 4-frame video generation module

[1827] The server generates an animation based on the generated four-panel comic, creating a four-panel video by displaying each still image in succession and adding movement.

[1828] Examples:

[1829] The server animates each frame of the four-panel comic to create a short video.

[1830] Content Viewing and Sharing Module

[1831] The server places the four-panel comics and videos on news sites and also posts them on video sharing platforms and social networking services (SNS).

[1832] Examples:

[1833] The server places the generated four-panel comic on the homepage of the news site, uploads the four-panel video to YouTube or Twitter, and sets up a link to the news article.

[1834] Emotion Recognition and Response Module

[1835] The server utilizes an emotion engine to recognize the user's emotions. The emotion engine analyzes the user's camera footage and voice input to recognize the user's emotional state. Based on the recognized emotions, the content of the four-panel comics and videos is customized.

[1836] Examples:

[1837] The server analyzes the image from the user's facial recognition camera and recognizes that the user is "happy."

[1838] Based on this emotion, the emotion engine generates and displays four-panel comics with positive content to the user.

[1839] Similarly, the content and order of the four-frame videos are customized according to the user's emotions.

[1840] Example

[1841] When a user accesses a news site, their emotions are recognized through a camera or voice input device. Based on the recognized emotions, a four-panel comic strip or video that best suits the user is generated and displayed. Users can click on this content to read detailed news articles. They can also access the news site from a four-panel video they have viewed on a social media or video sharing platform.

[1842] example:

[1843] A user visits a news site and is perceived as "excited" based on five minutes of webcam footage.

[1844] Based on this, the server generates a four-panel comic strip of the active content and displays it on the news site's homepage.

[1845] Users click on the thumbnail of a four-panel comic that interests them and read a detailed news article.

[1846] The above is a specific embodiment of the present invention, and by combining emotion recognition, it is possible to provide personalized news content to users, which is more likely to attract users' attention and increase usage frequency than conventional news distribution methods.

[1847] The processing flow will be explained below.

[1848] Step 1:

[1849] The server retrieves the latest news articles using the news site's API or RSS feed. It periodically sends requests using the API key to collect news data. The retrieved news articles are then stored in a database.

[1850] Specific behavior:

[1851] The server accesses the "News API" and sends a request to retrieve the latest news articles.

[1852] The title, text, and URL of the news article received as a response are extracted and stored in a database.

[1853] Step 2:

[1854] The server invokes a natural language processing (NLP) engine to analyze the text of the stored news articles, thereby extracting important keywords and key points.

[1855] Specific behavior:

[1856] The server inputs the text of the news article into the NLP engine.

[1857] The NLP engine analyzes the text and extracts key keywords (e.g., "climate change," "extreme weather," "impact") and key points.

[1858] The extracted keywords and points are stored in a database.

[1859] Step 3:

[1860] The server uses generative AI to generate a four-panel comic scenario based on the extracted keywords and points. The generated scenario defines the situation and storyline of each panel.

[1861] Specific behavior:

[1862] The server inputs keywords and points into the generation AI and requests it to generate a four-panel scenario.

[1863] The generation AI creates a scenario and returns a description of each frame (e.g., "Frame 1: A character thinks about climate change," "Frame 2: Extreme weather occurs," etc.).

[1864] The generated scenario is saved in the database.

[1865] Step 4:

[1866] The server uses image generation AI to create a four-panel comic based on the generated scenario. The image generation AI generates illustrations based on the text description of each panel.

[1867] Specific behavior:

[1868] The server inputs the scenario text into the image generation AI and requests it to generate an illustration.

[1869] The image generation AI generates an illustration corresponding to each frame and returns it to the server.

[1870] The generated illustrations are arranged in the form of a four-panel comic.

[1871] Step 5:

[1872] The server generates animations based on each frame of the four-panel comic, creating a four-panel video. It displays still images in succession and adds movement to each scene.

[1873] Specific behavior:

[1874] The server provides the AI ​​with a script to animate the illustrations in each panel of the four-panel comic.

[1875] The generative AI generates the movement of each frame and creates a video file as an animation.

[1876] The generated video file is saved on the server.

[1877] Step 6:

[1878] The server places the generated four-panel comics and videos on the front page of the news site, and also posts them on video sharing platforms and social networking services to promote access.

[1879] Specific behavior:

[1880] The server generates HTML code to embed the four-panel comic on the top page of the news site.

[1881] The generated four-panel comic is displayed on the top page.

[1882] Upload your four-frame video to platforms such as YouTube and Twitter.

[1883] Link to the news article in the video description.

[1884] Step 7:

[1885] Users can click on thumbnails of four-panel comics or videos displayed on the homepage of news sites to read detailed news articles, or watch videos on social media or video-sharing platforms and access the original articles.

[1886] Specific behavior:

[1887] A user visits the homepage of a news site and clicks on a thumbnail of a four-panel comic strip.

[1888] Clicking on the thumbnail will take you to the news article details page.

[1889] Users watch four-frame videos on YouTube or Twitter and click on a link in the video description to access a news site.

[1890] Step 8:

[1891] The server uses an emotion engine to recognize the user's emotions. It analyzes the user's camera footage and voice input to recognize the user's emotional state.

[1892] Specific behavior:

[1893] When a user accesses a news site, the server collects camera footage and audio data.

[1894] The server calls the emotion engine and inputs the collected video and audio data.

[1895] The emotion engine analyzes the user's emotions (e.g., "happy" or "excited") and returns the results to the server.

[1896] Step 9:

[1897] The server customizes the content of the four-panel comic or video based on the recognized emotion.

[1898] Specific behavior:

[1899] The server adjusts the content of the four-panel comics and videos based on the emotional data obtained from the emotion engine.

[1900] For example, if the user is recognized as "happy," a four-panel comic strip with positive content is generated.

[1901] Similarly, the content and order of the four-frame video are changed according to the user's emotions.

[1902] The above are the specific processing steps of the present invention that incorporate emotion recognition, which enable providing users with personalized visual news content.

[1903] Example 2

[1904] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1905] In recent years, the distribution methods of news articles have become more diverse, making it increasingly difficult to keep readers interested. Furthermore, if the content of the article does not match the reader's emotions, it is difficult to maintain their interest. Therefore, there is a need for a method to provide news articles in a more user-friendly format and further customize the content based on the user's emotions.

[1906] The specific processing by the specific processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for acquiring news articles, means for analyzing the acquired news articles and extracting important data and key points, means for generating a four-panel cartoon scenario using a generative AI model based on the extracted data and key points, means for generating a four-panel cartoon using an image generation AI based on the generated scenario, means for generating animation based on the generated four-panel cartoon and creating a four-panel video, means for displaying the generated four-panel cartoon and four-panel video, means for posting the generated four-panel video to an information sharing platform or a social networking service to promote access to the original article, and means for recognizing user emotions and customizing content based on the user's emotions. This makes it possible to provide news articles in familiar four-panel cartoon or video formats and further provide personalized content tailored to the user's emotions.

[1907] A "news article" is a report of a fact or event provided online or in print.

[1908] A "server" is a computer or program that processes data on a network and provides services to other computers.

[1909] "Means of acquisition" are the methods and techniques used to collect specific data and store it in a usable format.

[1910] "Means of analysis" are methods and techniques for analyzing given data and understanding its meaning and structure.

[1911] "Important data" is information, such as news articles, that is deemed to be particularly relevant and valuable.

[1912] "Gist" refers to the most important or core part of information or data.

[1913] A "generative AI model" is an algorithm or program that uses artificial intelligence technology to generate new data and information.

[1914] "Scenario generation means" refers to methods and technologies for automatically creating stories and structures based on basic data.

[1915] "Image generation AI" refers to algorithms and programs that use artificial intelligence technology to generate image data based on input information.

[1916] "Four-panel comics" are a form of comic expression that consists of four images or illustrations arranged as a continuous story.

[1917] "Means for generating animation" refers to methods or techniques for expressing movement by displaying still images or illustrations in succession.

[1918] "4-frame video" is a video format that displays four scenes in succession and adds movement and effects.

[1919] The "display means" refers to a method or technology for providing the generated content in a format that can be visually confirmed by the user.

[1920] An "information sharing platform" is a website or application that allows users to share content or information with other users over the Internet.

[1921] A "social networking service (SNS)" is a service that allows users to connect with each other online and communicate and share information.

[1922] "Original article" refers to the original news article collected by the news article acquisition means and used for analysis and transformation.

[1923] "Means for recognizing emotions" refers to methods or technologies for identifying a user's emotional state by analyzing their facial expressions and voice.

[1924] "Customization means" are methods and techniques for modifying and tailoring content to a user's specific needs or circumstances.

[1925] overview

[1926] This invention is a system that converts news articles into four-panel comics and four-panel videos and customizes the content based on the user's emotions, making it possible to provide individual news content that matches the user's interests.

[1927] System Configuration

[1928] The system consists of the following main modules:

[1929] 1. News article acquisition module

[1930] 2. Text Analysis Module

[1931] 3. Four-panel comic generation module

[1932] 4. 4-frame video generation module

[1933] 5. Content Display and Sharing Module

[1934] 6. Emotion Recognition and Response Module

[1935] News article acquisition module

[1936] The server retrieves the latest news articles using the news site's API or RSS feed, which allows the latest news to be stored in the database at all times.

[1937] Examples:

[1938] The server accesses the "News API" to retrieve the latest news articles, then stores the article title, body text, and URL in a database.

[1939] Text Analysis Module

[1940] The server analyzes the text of the retrieved news articles and uses natural language processing (NLP) techniques, such as morphological analysis and topic modeling, to extract key data and key points.

[1941] Examples:

[1942] The server analyzes news articles and extracts the keywords "climate change," "abnormal weather," and "impact."

[1943] Four-panel comic generation module

[1944] The server uses a generative AI model to automatically generate a four-panel comic scenario based on the extracted data and key points. Based on this scenario, an image generation AI generates illustrations for each panel.

[1945] Examples:

[1946] Prompt: "A character who thinks about climate change"

[1947] 1. Frame 1: A scene where a character thinks about "climate change"

[1948] 2. Frame 2: A scene where abnormal weather is occurring

[1949] 3. Panel 3: People are surprised by the abnormal weather.

[1950] 4. Panel 4: A scene of characters thinking about climate change countermeasures

[1951] The server inputs a prompt sentence into the generative AI model, and passes the output scenario to the image generation AI to generate an illustration.

[1952] 4-frame video generation module

[1953] The server generates animations based on the generated four-panel comic illustrations, creating a four-panel video. Each still image is displayed in succession, and a short video is created by adding movement.

[1954] Examples:

[1955] The server uses animation production software (e.g., Adobe After Effects) to place the illustrations for each frame on a timeline, add transitions and effects, and generate the video.

[1956] Content Viewing and Sharing Module

[1957] The server places the generated four-panel comics and four-panel videos on a news site and also posts them on video sharing platforms and social networking services (SNS).

[1958] Examples:

[1959] The server uses HTML and JavaScript to display four-panel comics on the news site's homepage, and posts four-panel videos using YouTube and Twitter APIs.

[1960] Emotion Recognition and Response Module

[1961] The server uses an emotion engine to recognize the user's emotions. It analyzes the user's camera footage and voice input to recognize their emotional state. The content of the four-panel comics and videos is customized based on the recognized emotions.

[1962] Examples:

[1963] The server uses emotion recognition software (e.g., Microsoft Azure Cognitive Services) to analyze the user's camera footage, and if it recognizes the emotion as "happy," it generates and displays a four-panel comic strip with a positive message.

[1964] Example

[1965] When a user accesses a news site, their emotions are recognized through a camera or voice input device. Based on the recognized emotions, a four-panel comic strip or video that best suits the user is generated and displayed. Users can click on this content to read detailed news articles. They can also access the news site from a four-panel video they have viewed on a social media or video sharing platform.

[1966] Examples:

[1967] A user visits a news site and is identified as "excited" from five minutes of webcam footage. Based on this, the server generates a four-panel comic strip of the active content and displays it on the news site's homepage. The user clicks on the thumbnail of the four-panel comic strip that interests them and reads the detailed news article.

[1968] The above is a specific embodiment of the invention, and by combining it with emotion recognition, it is possible to provide personalized news content to users, which is more likely to attract users' attention and increase usage frequency compared to conventional news distribution methods.

[1969] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1970] Step 1: Get news articles

[1971] The server retrieves the latest news articles using the news site's API or RSS feed. The input is the news site's URL or API endpoint, and the output is the news article data. Specifically, the server sends an HTTP request, extracts the title, body, and URL of the news article from the received response, and stores them in a database.

[1972] Step 2: Analyzing the news article

[1973] The server retrieves news articles stored in a database and analyzes the text using natural language processing (NLP) technology. The input is the text data of the news article, and the output is extracted keywords and key points. Specifically, the server uses a morphological analysis tool to segment the text and performs topic modeling to extract important keywords.

[1974] Step 3: Creating a four-panel comic scenario

[1975] The server uses a generative AI model based on the extracted keywords and key points to automatically generate a four-panel comic scenario. The input is the extracted keywords, and the output is the four-panel comic scenario. Specifically, the server inputs a prompt statement (e.g., "A character thinking about climate change") into the generative AI model and retrieves the generated scenario.

[1976] Step 4: Generate a four-panel comic

[1977] The server uses image generation AI to generate illustrations for a four-panel comic based on the generated scenario. The input is the generated scenario, and the output is the illustrations for each panel. Specifically, the server passes the scenario for each panel to the image generation AI, which then generates the illustrations.

[1978] Step 5: Generate a 4-frame video

[1979] The server generates animation based on the generated four-panel cartoon illustrations, creating a four-panel video. The input is a four-panel cartoon illustration, and the output is a four-panel video. Specifically, the server uses animation production software to place the illustrations of each frame on a timeline, add effects and transitions, and create a short video.

[1980] Step 6: View and share content

[1981] The server places the generated four-panel comics and videos on news sites and also posts them on video sharing platforms and social networking services. The input is the generated four-panel comics and videos, and the output is public information on the news sites and sharing platforms. Specifically, the server updates the HTML and JavaScript on the news sites and posts the videos using the YouTube and Twitter APIs.

[1982] Step 7: User Emotion Recognition and Customization

[1983] The server uses an emotion engine to recognize the user's emotions. The input is the user's camera footage and audio input, and the output is the recognized emotional state. Specifically, the server uses emotion recognition software to analyze the camera footage and identify the user's emotions. Based on the emotional state, the content of the four-panel comic or video is customized.

[1984] By performing the above processing steps, news articles can be presented in the form of familiar four-panel comic strips or videos, and personalized content tailored to the user's emotions can be provided.

[1985] (Application example 2)

[1986] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1987] Modern news distribution services only provide information one-way to users, and have the problem of being unable to provide personalized content that reflects users' emotions and interests. Furthermore, they lack the visual and entertaining format to keep users engaged with the news. As a result, it is difficult to maintain users' interest, leading to a decline in frequency of use.

[1988] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1989] In this invention, the server includes: means for acquiring news articles; means for analyzing the acquired news articles and extracting important keywords and key points; means for generating a four-panel cartoon scenario using a generation AI based on the extracted keywords and key points; means for generating a four-panel cartoon using an image generation AI based on the generated scenario; means for generating an animation based on the generated four-panel cartoon and creating a four-panel video; means for displaying the generated four-panel cartoon and four-panel video; means for posting the generated four-panel video on a video sharing platform or a social networking service to promote access to the original article; means for recognizing a user's emotions using a camera or a voice input device; means for customizing the content of the generated four-panel cartoon and four-panel video based on the recognized emotion; and means for displaying the customized four-panel cartoon and four-panel video within the smartphone app. This allows news articles to be provided in a form adapted to the user's emotions, thereby increasing user interest and increasing usage frequency.

[1990] "Means for obtaining news articles" is a function for automatically collecting the latest news articles from news sites and APIs on the Internet.

[1991] "Means for analyzing news articles and extracting important keywords and key points" refers to technology for analyzing the content of acquired news articles and automatically extracting the main topics and important information of the articles.

[1992] "Generative AI" is a general term for artificial intelligence that uses natural language processing and machine learning techniques to generate stories and scenarios from input data.

[1993] "Image generation AI" is artificial intelligence that creates visual representations based on given scenarios or prompts.

[1994] The "means for generating four-panel comics" is a technology that automatically creates a comic consisting of four consecutive panels based on a generated scenario.

[1995] "Means for generating animations and creating four-frame videos" refers to a technology that converts the generated four-frame comics into moving animations and provides them in the form of short videos.

[1996] "Means for displaying the generated four-panel comics and four-panel videos" refers to technology for displaying the generated content in a manner that allows users to visually access it.

[1997] "A means of posting the generated four-frame video on video sharing platforms and social networking services to promote access to the original article" refers to a function that posts the generated video on YouTube or social media, directing viewers to the original news article.

[1998] "Means for recognizing a user's emotions using a camera or voice input device" refers to technology that analyzes camera footage and voice data to detect the user's emotional state.

[1999] "Means for customizing the content of generated four-panel comics and four-panel videos based on recognized emotions" refers to technology that adjusts the content and tone of generated content based on the user's emotional data.

[2000] "Means for displaying customized four-panel comics and four-panel videos within a smartphone app" is a function for providing users with emotionally tailored content through a smartphone app.

[2001] A specific embodiment for carrying out the present invention will be described.

[2002] Hardware and software used

[2003] The server uses the following major hardware and software:

[2004] Hardware

[2005] Internet connection required to retrieve news data

[2006] Camera and microphone for recognizing user emotions

[2007] Display devices such as smartphones and tablets

[2008] software

[2009] News Acquisition API

[2010] Natural language processing libraries (such as TextBlob)

[2011] Image generation AI (general image generation algorithms)

[2012] Video Generation Algorithm

[2013] Emotion recognition software (OpenCV and other facial recognition technologies)

[2014] Social media APIs (YouTube API, Twitter API, etc.)

[2015] Data processing and calculation flow

[2016] The server retrieves the latest news articles from the Internet, using a news API to pull news data into the server, and then analyzes the text data of the news articles using natural language processing technology to extract important keywords and key points.

[2017] Based on the extracted keywords and points, a scenario for a four-panel comic is automatically generated by the generative AI. This scenario is passed to the image generation AI used to generate a comic illustration consisting of four consecutive frames. The generated four-panel comic is then animated and presented as a short video.

[2018] The generated comic strips and videos are displayed on smartphone apps and other devices. When a user accesses the app, the app recognizes the user's emotions in real time via the camera and microphone, and the content of the comic strips and videos generated is customized based on those emotions. These contents are also posted on social networking services and video sharing platforms, and are used as a means to direct users to the original news article.

[2019] Specific examples

[2020] For example, if a news article mentions "climate change," the system might:

[2021] News acquisition

[2022] Use the News API to get the latest "Climate Change" articles.

[2023] Text analytics

[2024] Use a natural language processing library such as TextBlob to extract keywords such as "climate change," "extreme weather," and "impact" from articles.

[2025] Scenario generation and image generation

[2026] A scenario is generated based on the extracted keywords, and the image generation AI is used to draw the following four-panel comic scenario:

[2027] 1. A scene where a character thinks about climate change

[2028] 2. Scenes where abnormal weather is occurring

[2029] 3. Scenes of people being surprised by extreme weather

[2030] 4. A scene with characters thinking about climate change

[2031] Based on these scenarios, we generate prompts and pass them to the image generation AI:

[2032] "Generate a frame of a character talking at length about climate change."

[2033] "Draw a scene where extreme weather is occurring."

[2034] "Draw a frame showing people being surprised by the extreme weather."

[2035] Video Generation and Display

[2036] An animation is created based on the generated four-panel comic and converted into a short video format.

[2037] The video and four-panel comics are displayed within a smartphone app. When a user accesses the app, the camera and microphone are used to recognize their emotions, and content is customized accordingly.

[2038] Social Media Posts

[2039] The resulting four-frame video can then be posted on social media and video sharing platforms for a wider audience.

[2040] In this way, news articles can be provided in a form that is adapted to the user's emotions, which can increase the user's interest and increase the frequency of use.

[2041] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[2042] Step 1:

[2043] The server uses a news API to retrieve the latest news articles from the Internet. The retrieved article data includes the article title, body text, URL, etc. The input is the response from the news API, and the output is text data for analysis.

[2044] Step 2:

[2045] The server analyzes the text data of the acquired news articles using natural language processing (NLP) techniques. Specifically, it uses libraries such as TextBlob to extract important keywords and key points. The input for this process is the text data, and the output is the extracted keywords and key points.

[2046] Step 3:

[2047] The server uses a generative AI model to generate a four-panel comic scenario based on the extracted keywords and points. Here, a group of keywords is passed as input, and the output is four prompt sentences containing the scenario for each panel. Specifically, the server converts the keywords into prompt sentences and passes them to the generative AI.

[2048] Step 4:

[2049] The server generates a four-panel comic using an image generation AI based on the generated scenario. At this stage, the input is a prompt sentence, and the output is four consecutive comic illustrations. Specifically, the prompt is passed to the image generation AI based on the scenario output from the generative AI model.

[2050] Step 5:

[2051] The server generates animation based on the generated four-panel comic to create a four-panel video. Here, still images of the four-panel comic are passed as input, and the output is a short video with continuous movements. Specifically, the video is generated by connecting still images and adding movement.

[2052] Step 6:

[2053] The server displays the generated four-panel comics and four-panel videos on a smartphone app or other device. The input is the generated four-panel comics and videos, and the output is visual data that is displayed to the user. Specifically, it includes the operation of displaying them in a specific section within the smartphone app.

[2054] Step 7:

[2055] When a user accesses the smartphone app, the app uses the device's camera and microphone to recognize emotions in real time. The input is the user's real-time video and audio data, and the output is analyzed emotional data. Specifically, the app uses emotion recognition software to analyze facial expressions and tone of voice.

[2056] Step 8:

[2057] The server customizes the content of the generated four-panel comics and four-panel videos based on the recognized emotions. The input is emotion data, and the output is customized four-panel comics and videos. Specific operations include reselecting scenarios and images according to the emotion.

[2058] Step 9:

[2059] The server posts customized four-panel comics and videos to social networking services and video sharing platforms to promote access to the original article. The input is the customized video and comic, and the output is a post on social media. Specifically, it automates the posting using the platform's API.

[2060] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[2061] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[2062] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[2063] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[2064] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[2065] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[2066] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[2067] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[2068] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[2069] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[2070] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[2071] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[2072] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[2073] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[2074] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[2075] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[2076] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[2077] As an example of a system configu...

Claims

1. a means for obtaining news articles; A means to analyze the acquired news articles and extract important keywords and points, A method to generate a four-panel comic scenario using generative AI based on the extracted keywords and points, A method for generating a four-panel comic using image generation AI based on the generated scenario, and A method to generate animation based on the generated four-frame comic and create a four-frame video. A means for displaying the generated four-frame comic and four-frame video; The generated four-frame videos will be posted on video sharing platforms and social networking services to promote access to the original articles. A system including:

2. 10. The system of claim 1, wherein natural language processing techniques are used to analyze news articles.

3. The system according to claim 1, further comprising means for placing the generated four-panel cartoon on the top page of a news site.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A

Cited By

  • Information processing methods, information processing systems, and programs

    JP7863699B1