system

The system addresses the challenge of accessing foreign literature by generating concise summaries and videos of best-selling books, allowing users to efficiently understand book content through visually engaging summaries.

JP2026018086APending Publication Date: 2026-02-05SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024119147
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-24
Publication Date
2026-02-05

AI Technical Summary

Technical Problem

There is a growing demand for quick and concise access to information about foreign literature and best-selling books, as many books have not been translated, making it tedious to read them in other languages, and users have limited means to understand the contents efficiently.

Method used

A system that includes acquiring book information from a database or API, generating a summary using a generative AI model, creating a video based on the summary, uploading the video to a distribution service, and notifying the user, with the summary limited to 200 characters and video length to 30 seconds, and including related images.

Benefits of technology

Enables users to quickly obtain the essence of a book without language barriers, providing efficient and easy-to-understand information through visually appealing content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026018086000001_ABST
    Figure 2026018086000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system comprising: means for acquiring book information from a book database or an API; generative AI model means for generating a summary using the acquired book information; means for generating a video based on the generated summary; means for uploading the generated video to a video distribution service; and means for notifying a user.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] In modern society, there is a growing demand for quick and concise access to information about foreign literature and best-selling books. However, many books have not been translated, making it tedious to read books in other languages. Furthermore, users have limited means of quickly understanding the contents of books. There is a need for a solution to these problems and a way to easily obtain new information from overseas. [Means for solving the problem]

[0005] The present invention provides a system that includes a means for acquiring book information from a book database or API, a generative AI model means for generating a summary using the acquired book information, a means for generating a video based on the generated summary, a means for uploading the generated video to a video distribution service, and a means for notifying the user. This system allows users to quickly obtain the essence of a book of interest without experiencing language barriers. In addition, by limiting the generated summary to 200 characters or less and the video length to 30 seconds or less, efficient and easy-to-understand information provision is achieved. Furthermore, by including a means for automatically generating related images based on the summary, visually appealing content can be provided.

[0006] A "book database" is a collection of data for storing and managing book information.

[0007] "API" stands for Application Program Interface, an interface for communication between software programs.

[0008] "Book information" is data about a book, such as the book title, author, synopsis, and publication date.

[0009] A summary is a short sentence that succinctly summarizes the main content of a book.

[0010] A "generative AI model" is an algorithm or method that uses artificial intelligence to generate text or content based on specific inputs.

[0011] "Video" refers to moving video data that combines images and audio.

[0012] A "video distribution service" is a platform that distributes videos to users via the Internet in the form of streaming or downloads.

[0013] A "notification" is a message that informs the user of a specific piece of information or action.

[0014] "User" refers to a person or organization that uses the system.

[0015] "Image generation" is the process of generating images based on text or other input data.

[0016] A "push notification" is a notification message sent from a server to a device in real time. [Brief explanation of the drawings]

[0017] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11]FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0018] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0019] First, the terms used in the following description will be explained.

[0020] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0021] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0022] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0023] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0024] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0025] [First embodiment]

[0026] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0027] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0028] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0029] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0030] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0031] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0032] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0033] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0034] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0035] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0036] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0037] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0038] A natural language description of the program's operation

[0039] This system provides an automated process to summarize international best-selling books in 200 characters or less and distribute them as 30-second videos. The following explains the roles and processes of the server, terminal, and user in detail.

[0040] Get book data

[0041] The server retrieves the book data.

[0042] First, the server connects to a specific book database or API (Application Program Interface) to retrieve new book information. The server runs a scheduled task at a regular time, for example, retrieving the latest best-selling book list at midnight every day. The retrieved data includes the book title, author, synopsis, publication date, etc.

[0043] Summary generation

[0044] The server generates a summary

[0045] Next, the server generates a summary using a generative AI model (such as GPT-3) based on the acquired book information. The server inputs the book title, author, and synopsis into the generative AI model in a prompt format, and the AI ​​model outputs a summary of up to 200 characters. The generated summary is then verified by the server to ensure it does not contain any inappropriate content.

[0046] Video generation

[0047] The server generates a video from the summary text

[0048] Based on the verified summary, the server automatically generates or collects related images. The server extracts keywords from the summary and, depending on the keywords, collects related images from the Internet or selects them from an internal image database. The server then combines these images with the summary, overlaying text on the images to generate a slideshow-style video. Video generation software such as FFmpeg is used to generate video files of up to 30 seconds.

[0049] Video distribution

[0050] The server uploads the video to the video streaming service.

[0051] The server automatically uploads the generated video file to a video distribution service (such as YouTube or Vimeo). The server uses the video distribution service's API to set the information required when uploading the video (title, description, tags, etc.). Once the upload is complete, the video's URL is obtained from the service and saved in an internal database.

[0052] User Notifications

[0053] The device notifies the user of the video link

[0054] The server sends a message to the user's device to notify them that a new video has been uploaded. The device (smartphone or computer) then displays a push notification to the user with a message such as "A new summary video has been released." After confirming the notification, the user clicks on the provided video link to access the video streaming service and watch the 30-second summary video.

[0055] Specific examples

[0056] 1. Acquiring book data

[0057] The server retrieves the bestselling book data from the API at midnight every day. For example, it retrieves the data for the book "The Silent Patient."

[0058] 2. Summary Generation

[0059] The server inputs the data from "The Silent Patient" into a generative AI model and generates the summary: "Alicia is an artist who goes silent after a mysterious incident, and the story is about a therapist who searches for the truth behind it."

[0060] 3. Video Generation

[0061] The server collects free images based on keywords such as "Alicia" and "therapist," and combines the summary text and images to generate a 30-second video.

[0062] 4. Video distribution

[0063] The server uploads the generated video to YouTube, retrieves the video URL and stores it in the database.

[0064] 5. User Notices

[0065] The server sends a push notification to the user's smartphone, and the user clicks on the notification to watch the video.

[0066] This system allows users to easily view summaries of international best-selling books in a short amount of time, enabling them to acquire information efficiently.

[0067] The processing flow will be explained below.

[0068] Step 1:

[0069] The server connects to the book database or API at midnight every day to retrieve the latest bestselling book data, including the book title, author, synopsis, and publication date.

[0070] Step 2:

[0071] The server stores the acquired book information in an internal database and adds it to a processing queue, so that subsequent processes can retrieve and process the book information sequentially.

[0072] Step 3:

[0073] The server retrieves a book from the processing queue and inputs the book's title, author, and synopsis into the generative AI model, providing appropriate input to the AI ​​model in the form of prompts.

[0074] Step 4:

[0075] The generative AI model outputs a summary of up to 200 characters based on the input information. The server receives the summary and prepares it for the next process.

[0076] Step 5:

[0077] The server then validates the generated summary to check for inappropriate content, for example, by checking for inappropriate language or unclear sentences.

[0078] Step 6:

[0079] The server extracts keywords from the verified abstract and passes the list of keywords, including people's names, places, and events, to the image generation module.

[0080] Step 7:

[0081] The image generation module collects free stock images corresponding to each keyword from the Internet or selects them from an internal image database, and the server downloads or retrieves the images.

[0082] Step 8:

[0083] The server combines the abstract with the acquired images to generate a slideshow-style video by overlaying the text on the images. Video generation software (e.g., FFmpeg) is used to generate a video file of up to 30 seconds.

[0084] Step 9:

[0085] The server uploads the generated video file to a video distribution service (e.g., YouTube or Vimeo), and also sets the video's metadata (title, description, tags, etc.) at the same time.

[0086] Step 10:

[0087] Once the upload is complete, the server retrieves the generated video URL from the video streaming service, which is then stored in an internal database.

[0088] Step 11:

[0089] The server sends a message to the user's device informing them that a new video has been uploaded, which includes sending a push notification.

[0090] Step 12:

[0091] The device displays a notification to the user saying, "A new summary video has been released." The user checks the notification and clicks on the provided video link.

[0092] Step 13:

[0093] Users can easily obtain summaries of international best-selling books by watching 30-second summary videos on their device's browser or YouTube app.

[0094] This detailed processing flow automates and efficiently executes the entire process of obtaining book information, generating a summary and video, and finally notifying the user.

[0095] Example 1

[0096] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0097] Current book information systems do not allow users to quickly obtain summaries of international best-selling books, making it difficult for users to grasp the content in a short time. Furthermore, when summaries are provided only in text, the lack of visual elements makes it difficult for users to understand or engage with the information. Therefore, there is a need for a method to provide book information efficiently and in a visually understandable way.

[0098] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0099] In this invention, the server includes means for acquiring book information from a book database or API, means for creating a prompt based on the acquired book information and sending it to the generative AI model, means for verifying the summary returned from the generative AI model, means for extracting keywords from the summary and collecting related images, means for generating a slideshow video based on the collected images and summary, means for uploading the generated video to a video distribution service, and means for sending a notification message including a link to the video to the user terminal. This allows users to quickly and visually understand summaries of international best-selling books.

[0100] A "book database" is a system that systematically stores information about books and has a data structure that allows for searching and extraction.

[0101] An "API" is an interface that allows various pieces of software to communicate with each other and share functions, and is often provided in the form of a web service.

[0102] A "prompt sentence" is an input sentence provided to a generative AI model, a series of text that is generated in anticipation of specific conditions or responses.

[0103] A "generative AI model" is an artificial intelligence algorithm that automatically generates text and information from large amounts of data.

[0104] "Validation" is the process of verifying whether the generated data is appropriate and checking that it does not contain any inappropriate content.

[0105] "Keywords" are important words extracted from summaries and texts, and are used for information retrieval and classification.

[0106] "Image gathering" is the process of acquiring relevant image data based on specific criteria or keywords.

[0107] "Slideshow format" is a visual presentation format that displays multiple images and text sequentially.

[0108] A "video distribution service" is an online platform for uploading videos over the Internet and making them available to viewers.

[0109] A "notification message" is a message sent by the system to convey specific information to the user.

[0110] A "user terminal" is a device such as a smartphone or computer that allows a user to access the Internet and applications.

[0111] To implement this invention, a system is required that automates the process of acquiring book information from a book database or API, generating a summary using a generative AI model based on that information, creating a video using that summary, and finally notifying the user. The following describes in detail the hardware and software used in each processing step, as well as the details of data processing and data calculation.

[0112] Get book data

[0113] Server Roles

[0114] The server connects to a specific book database or API to retrieve new book information. For example, it can use the Goodreads or Google Books API. The server runs a periodic task every day at midnight, sending an HTTP request to retrieve the latest best-selling book list. The retrieved data is provided in JSON format and includes the book title, author, synopsis, publication date, etc.

[0115] Summary generation

[0116] Server Roles

[0117] The server creates a prompt based on the acquired book information and sends it to the generative AI model. Specifically, it combines the book title, author, and synopsis to create a prompt to be input to the generative AI model (e.g., GPT-3). The following is an example of a prompt:

[0118] Example prompt sentence:

[0119] "Book title: 'XXX', author: XXX, summary: XXX. Please write a summary in 200 characters or less."

[0120] The server receives the summary returned by the generative AI model and uses an internal validation algorithm to check for inappropriate content. If there are no problems, the summary is stored in the database.

[0121] Video generation

[0122] Server Roles

[0123] The server extracts keywords from the summary and collects related images. It uses a natural language processing algorithm to extract key keywords from the summary and collects images via HTTP requests from free image services (such as Unsplash and Pexels). It then combines the collected images with the summary to generate a slideshow-style video using video generation software such as FFmpeg. The video must be no longer than 30 seconds.

[0124] Video distribution

[0125] Server Roles

[0126] The server uploads the generated video to a video distribution service (such as YouTube or Vimeo). The server uses the API of the video distribution service to set the necessary information such as title, description, and tags when uploading the video. If the upload is successful, the server obtains the URL of the video returned by the video service and saves it in the database.

[0127] User Notifications

[0128] Device Role

[0129] The server generates a notification message containing the URL of the new video and sends a push notification to the user's device. The device (smartphone or computer) then displays a notification to the user saying, "A new summary video has been released." The user can click the notification to access the video streaming service and watch a 30-second summary video.

[0130] Specific examples

[0131] For example, for the book "Silent Patient," the server sends an HTTP request to the Goodreads API at midnight every day to retrieve information about the book. Next, the following prompt is sent to the generative AI model: "Book Title: 'Silent Patient,' Author: Alex Michaelides, Synopsis: Alicia is an artist who falls silent after a mysterious incident, and the story of a therapist who searches for the truth behind this. Please provide a summary of 200 characters or less." Based on the summary returned by the generative AI model, "Alicia is an artist who falls silent after a mysterious incident, and the story of a therapist who searches for the truth behind this," the server extracts keywords such as "Alicia" and "therapist" and collects related free images using the Unsplash API. These images and the summary are then processed into a slideshow video using FFmpeg. The server then uploads this video through the YouTube API and gives it a title such as "Silent Patient Summary Video." The user's device receives a notification stating, "A new summary video of 'Silent Patient' has been uploaded," and they can click the notification to watch the video.

[0132] This system allows users to quickly obtain summaries of international best-selling books in a visually easy-to-understand manner.

[0133] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0134] Step 1:

[0135] Get book data

[0136] The server connects to the book database or API to retrieve new book information. The server sends an HTTP request at midnight every day to retrieve the latest bestselling book list. The retrieved data is provided in JSON format and includes the book title, author, synopsis, publication date, etc. The input is the book information retrieved from the book database or API, and the output is the parsed book information. The server parses the retrieved data and saves it in its internal database.

[0137] Step 2:

[0138] Creating and sending prompts

[0139] The server creates a prompt based on the acquired book information and sends it to the generative AI model. Specifically, it combines the book title, author, and synopsis to create a prompt to be input to the generative AI model (e.g., GPT-3). The input is the analyzed book information, and the output is the generated prompt. Below is an example of a prompt:

[0140] Example prompt sentence:

[0141] "Book title: 'XXX', author: XXX, summary: XXX. Please write a summary in 200 characters or less."

[0142] Step 3:

[0143] Summary generation and verification

[0144] The server receives the summary returned by the generative AI model. The server checks the generated summary with an internal verification algorithm to ensure that it does not contain inappropriate content. The input is the summary returned by the generative AI model, and the output is the summary after verification. If it is confirmed that it does not contain inappropriate content, the summary is saved in the database.

[0145] Step 4:

[0146] Keyword extraction and image collection

[0147] The server extracts keywords from the abstract and collects related images. It uses a natural language processing algorithm to extract key keywords from the abstract and collects images via HTTP requests from services that provide free images (e.g., Unsplash and Pexels). The input is the verified abstract, and the output is the extracted keywords and collected images.

[0148] Step 5:

[0149] Video generation

[0150] The server generates a slideshow-style video based on the collected images and summaries. The collected images are concatenated and the summaries are overlaid on the images as text. Video generation software such as FFmpeg is used to generate videos of up to 30 seconds. The input is the collected images and summaries, and the output is the generated video file.

[0151] Step 6:

[0152] Uploading videos

[0153] The server uploads the generated video to a video distribution service. Using the API of the video distribution service (for example, YouTube or Vimeo), necessary information such as title, description, and tags is set when uploading the video. The input is the generated video file, and the output is the URL of the uploaded video. If the upload is successful, the server saves the video URL returned by the video service in a database.

[0154] Step 7:

[0155] User Notifications

[0156] The server generates a notification message containing the URL of the new video and sends a push notification to the user's device. The device (smartphone or computer) receives the notification and displays a message to the user saying, "A new summary video has been released." The input is the URL of the uploaded video, and the output is the notification message displayed on the user's device. The user can click the notification to access the video streaming service and watch a 30-second summary video.

[0157] (Application example 1)

[0158] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0159] In today's world, busy people want to obtain information efficiently within a limited amount of time, but it is often difficult to find the time to read long books. Another issue is that simply reading a book summary can weaken comprehension of the content and the impact of the information. Furthermore, there are limited ways for users to easily view summarized book information, and this needs to be addressed.

[0160] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0161] In this invention, the server includes a means for acquiring book information from a book database or an API, a generation AI model means for generating a summary using the acquired book information, and a means for generating a video based on the generated summary, thereby enabling users to efficiently view summarized book information in video format.

[0162] "Book database" refers to one or more data collections that store information about books.

[0163] "API" refers to an interface that allows application programs to communicate with each other.

[0164] A "generative AI model" refers to artificial intelligence that automatically generates text, images, videos, etc. based on specified input data.

[0165] A "summary" is a short summary of the book's contents.

[0166] "Video distribution service" refers to a service that provides video content to users via the Internet.

[0167] "User" refers to an entity that uses the system to receive information or services.

[0168] "Smart devices" refer to mobile information devices with internet connectivity, such as smartphones and tablets.

[0169] MODE FOR CARRYING OUT THE INVENTION

[0170] System Program

[0171] This system retrieves book information from a book database or API, and generates a summary based on that information using a generative AI model. It consists of a series of processes: generating a video based on the generated summary, uploading the video to a video distribution service, and sending a notification to the user. Users can also watch the summary video on their smart devices. The program design for realizing this system is as follows:

[0172] Hardware and software used

[0173] Server hardware: Amazon EC2

[0174] Book data acquisition: Google Books API

[0175] Summary sentence generation: OpenAI GPT-3

[0176] Video generation: FFmpeg

[0177] Video streaming: YouTube Data API

[0178] Push notifications: Firebase Cloud Messaging (FCM)

[0179] Smart devices: smartphones, tablets, etc.

[0180] Data processing and calculation

[0181] The details of how this system works are as follows:

[0182] 1. Acquiring book data

[0183] The server runs a scheduled task at regular intervals, connecting to a book database or API (e.g., Google Books API) to retrieve the latest best-selling book list, including book title, author, synopsis, publication date, etc.

[0184] 2. Summary Generation

[0185] Based on the acquired book information, the server uses a generative AI model (e.g., OpenAI GPT-3) to generate a summary. Specifically, the book title, author, and synopsis are input into the generative AI model in the form of a prompt, and the AI ​​model outputs a summary of up to 200 characters. This summary is then verified by the server to ensure it does not contain any inappropriate content.

[0186] 3. Video Generation

[0187] Based on the verified summary, the server automatically generates or collects related images. The server extracts keywords from the summary and, based on those keywords, collects related images from the Internet or selects them from an internal image database. It then combines these images with the summary, overlaying text on the images to generate a slideshow-style video. It uses video generation software such as FFmpeg to generate video files of up to 30 seconds.

[0188] 4. Video distribution

[0189] The server automatically uploads the generated video file using the API of a video distribution service (such as YouTube or Vimeo). The necessary information (title, description, tags, etc.) is also set when uploading. Once the upload is complete, the video URL is obtained from the service and saved in an internal database.

[0190] 5. User Notices

[0191] The server sends a message to the user's smart device to notify them that a new video has been uploaded. A push notification displays a message to the user, such as "A new summary video has been uploaded." The user checks the notification and clicks on the provided video link to access the video streaming service and watch the 30-second summary video.

[0192] Specific examples

[0193] 1. Acquiring book data

[0194] Book: "The Silent Patient"

[0195] Author: "Alex Michaelides"

[0196] Synopsis: "Alicia Belson is a celebrated artist..."

[0197] 2. Summary Generation

[0198] Prompt statement:

[0199] Title: The Silent Patient

[0200] Author: Alex Michaelides

[0201] Summary: Alicia Belson is a renowned artist whose work is celebrated around the world. One day, she shoots and kills her husband in their home and never speaks again. This story is told from the perspective of a therapist who tries to unravel this mystery.

[0202] Please summarize in 200 characters or less.

[0203] Generated summary:

[0204] Alicia is an artist who goes silent after a mysterious incident, and the story revolves around a therapist who searches for the truth.

[0205] 3. Video Generation

[0206] The generated summary is combined with images related to keywords such as "Alicia" and "therapist," and a video of less than 30 seconds is generated using FFmpeg.

[0207] 4. Video distribution

[0208] Upload a video to YouTube, get the URL and save it in the database.

[0209] 5. User Notices

[0210] Use Firebase Cloud Messaging to send notifications of new video streams to users' smart devices.

[0211] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0212] Step 1:

[0213] The server periodically connects to a book database or API to retrieve the latest best-selling book information, including book title, author, synopsis, and publication date. This allows you to always have access to the latest book information.

[0214] Input: Scheduled task, book database or API

[0215] Output: Retrieved book information (title, author, synopsis, publication date)

[0216] Step 2:

[0217] The server generates a summary using a generative AI model (e.g., OpenAI GPT-3) based on the acquired book information. The book title, author, and synopsis are input into the generative AI model in the form of a prompt, and the AI ​​outputs a summary of up to 200 characters. At this time, the summary is verified to ensure that it does not contain any inappropriate content.

[0218] Input: Book information (title, author, synopsis), generative AI model

[0219] Output: Summary of up to 200 characters

[0220] Step 3:

[0221] The server extracts keywords from the verified summary and uses them to gather relevant images from the internet or select them from an internal image database. It then combines the summary with these images and generates a slideshow-style video by overlaying text on the images. It then uses video generation software such as FFmpeg to generate a video file of up to 30 seconds.

[0222] Input: Abstract, keywords, related images

[0223] Output: Video file up to 30 seconds long

[0224] Step 4:

[0225] The server uploads the generated video via the API of a video distribution service (e.g., YouTube). Information such as the title, description, and tags are also set. Once the upload is complete, the video URL is obtained from the service and saved in an internal database.

[0226] Input: Video file, video streaming service API

[0227] Output: Video URL

[0228] Step 5:

[0229] The server sends a push notification message via Firebase Cloud Messaging to the user's smart device to notify them that a new video has been uploaded, and the user receives the notification and clicks the video link to watch the video.

[0230] Input: Video URL, Firebase Cloud Messaging

[0231] Output: User notification

[0232] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0233] A natural language description of the program's operation

[0234] This system acquires book information from a book database or API, generates a summary based on the acquired information, and generates a video based on the summary. Furthermore, this system incorporates an emotion engine that recognizes the user's emotions, and provides the user with an optimal book summary video based on the acquired emotion information.

[0235] Get book data

[0236] The server retrieves the book data.

[0237] The server connects to a specific book database or API to retrieve data on new bestselling books, including the book's title, author, synopsis, and publication date. The server runs a scheduled task at a regular time to retrieve the data and store it in an internal database.

[0238] Summary generation

[0239] The server generates a summary

[0240] The server uses a generative AI model to generate a summary of up to 200 characters based on the acquired book information. By inputting the book title, author, and synopsis into the generative AI model, the AI ​​model outputs a summary. This summary is then verified by the server to ensure it does not contain any inappropriate content.

[0241] Video generation

[0242] The server generates a video from the summary text

[0243] Based on the verified summary, the server automatically generates or collects relevant images. Depending on the extracted keywords, appropriate images are collected from the Internet or selected from an internal image database. The server then combines these images with the summary, overlaying text on the images to generate a slideshow-style video. Video generation software (e.g., FFmpeg) is used to generate a video file of up to 30 seconds.

[0244] Video distribution

[0245] The server uploads the video to the video streaming service.

[0246] The server automatically uploads the generated video file to a video distribution service (for example, YouTube or Vimeo). The server uses the API of the video distribution service to set the information required when uploading the video (title, description, tags, etc.). Once the upload is complete, the server obtains the URL of the generated video from the video distribution service and saves it in its internal database.

[0247] User Notifications

[0248] The device notifies the user of the video link

[0249] The server sends a message to the user's device to notify them that a new video has been uploaded. The device (smartphone or computer) then displays a push notification to the user, saying something like, "A new summary video has been released." After checking the notification, the user clicks on the provided video link to access the video streaming service and watch the 30-second summary video.

[0250] Sentiment Analysis and Content Optimization

[0251] The server analyzes the user's emotions

[0252] The emotion engine analyzes user emotions based on their viewing history and reaction data. For example, it learns the genres and content that users have particularly rated among the videos they have watched in the past.

[0253] The server delivers the best content

[0254] Based on the emotional data analyzed by the emotion engine, the server selects and prioritizes the book summaries and videos that best fit the user's emotions. Specifically, it prioritizes books in the user's favorite genres and themes, and generates and distributes summaries and videos based on those.

[0255] Specific examples

[0256] 1. Acquiring book data

[0257] The server retrieves data on bestselling books from the API at midnight every day. For example, it retrieves data on the book "Introduction to Algorithms for Engineers."

[0258] 2. Summary Generation

[0259] The server inputs data from "Introduction to Algorithms for Engineers" into a generative AI model and generates a summary statement that reads, "This book introduces the basic concepts and practical applications of algorithms."

[0260] 3. Video Generation

[0261] The server collects related images based on keywords such as "algorithm" and "application," and combines the summary text and images to generate a 30-second video.

[0262] 4. Video distribution

[0263] The server uploads the generated video to YouTube, retrieves the video URL and stores it in the database.

[0264] 5. User Notices

[0265] The server sends a push notification to the user's smartphone, and the user clicks on the notification to watch the video.

[0266] 6. Sentiment Analysis and Content Optimization

[0267] The emotion engine analyzes the user's viewing history, and if, for example, the user often watches technical books, from the next time onwards, it will prioritize summarizing technical books and generating and delivering videos.

[0268] This detailed process allows users to efficiently watch video summaries of books that suit their preferences and emotions.

[0269] The processing flow will be explained below.

[0270] Step 1:

[0271] The server connects to the book database or API at midnight every day to retrieve the latest bestselling book data, including the book title, author, synopsis, and publication date.

[0272] Step 2:

[0273] The server stores the acquired book information in an internal database and adds it to a processing queue, so that subsequent processes can retrieve and process the book information sequentially.

[0274] Step 3:

[0275] The server retrieves a book from the processing queue and inputs the book's title, author, and synopsis into the generative AI model, providing appropriate input to the AI ​​model in the form of prompts.

[0276] Step 4:

[0277] The generative AI model outputs a summary of up to 200 characters based on the input information. The server receives the summary and prepares it for the next process.

[0278] Step 5:

[0279] The server then validates the generated summary to check for inappropriate content, for example, by checking for inappropriate language or unclear sentences.

[0280] Step 6:

[0281] The server extracts keywords from the verified abstract and passes the list of keywords, including people's names, places, and events, to the image generation module.

[0282] Step 7:

[0283] The image generation module collects free stock images corresponding to each keyword from the Internet or selects them from an internal image database, and the server downloads or retrieves the images.

[0284] Step 8:

[0285] The server combines the abstract with the acquired images to generate a slideshow-style video by overlaying the text on the images. Video generation software (e.g., FFmpeg) is used to generate a video file of up to 30 seconds.

[0286] Step 9:

[0287] The server uploads the generated video file to a video distribution service (e.g., YouTube or Vimeo), and also sets the video's metadata (title, description, tags, etc.) at the same time.

[0288] Step 10:

[0289] Once the upload is complete, the server retrieves the generated video URL from the video streaming service, which is then stored in an internal database.

[0290] Step 11:

[0291] The server sends a message to the user's device informing them that a new video has been uploaded, which includes sending a push notification.

[0292] Step 12:

[0293] The device displays a notification to the user saying, "A new summary video has been released." The user checks the notification and clicks on the provided video link.

[0294] Step 13:

[0295] Users can easily obtain summaries of international best-selling books by watching 30-second summary videos on their device's browser or YouTube app.

[0296] Sentiment Analysis and Content Optimization

[0297] Step 14:

[0298] The server collects the user's viewing history and reaction data and sends it to the emotion engine, which analyzes the user's viewing history and reaction data to estimate the user's emotions.

[0299] Step 15:

[0300] Based on the emotional data analyzed by the emotion engine, the server selects the most suitable book summary and video for the user. Specifically, it prioritizes books in genres and themes that the user has given high ratings to, and generates summaries and videos based on those.

[0301] Step 16:

[0302] The server uploads the generated optimal book summary video to a video distribution service and obtains a distribution link.

[0303] Step 17:

[0304] The server sends push notifications to the user's device at the appropriate time, allowing the user to receive a summary video that matches their emotions.

[0305] Step 18:

[0306] The device receives a notification and displays a message to the user saying, "A summary video has been delivered that is recommended for you." The user clicks the link to watch the video.

[0307] This allows the emotional engine to be used to provide summarized videos based on the user's emotions, resulting in even higher satisfaction. To give a concrete example, for example, a user who frequently reads technical books could be given preferential access to videos summarizing new best-selling books in the same genre. In this way, more personalized content is provided to users through sentiment analysis and content optimization.

[0308] Example 2

[0309] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0310] Conventional book recommendation systems have the problem that the process of generating a book summary and distributing it as a video incorporating visuals is complicated and difficult to automate. It is also difficult to provide personalized content that takes into account the user's preferences and emotions. This makes it difficult to attract the user's interest, resulting in a poor content consumption experience.

[0311] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for acquiring book information from a book database or API, a generation AI model means for generating a summary using the acquired book information, means for verifying the generated summary, means for automatically generating or collecting related images based on the verified summary, means for generating a video by combining the generated images and summary, means for uploading the generated video to a video distribution service, means for acquiring and saving the URL of the uploaded video, means for notifying the user, means for analyzing the user's emotions, and means for providing an optimal book summary video based on the analysis results. This automates the process from generating book summaries to distributing videos and providing content optimized for users, enabling the provision of an efficient and high-quality user experience.

[0312] "Book information" refers to basic data such as the book's title, author, synopsis, and publication date.

[0313] A "generative AI model" refers to an algorithm or software that automatically generates a summary from input data using natural language processing technology.

[0314] "Video Streaming Service" means a platform for hosting generated video content online and for users to stream or download it.

[0315] A "means" refers to a specific method, device, or technique for performing a particular function.

[0316] "Means for analyzing emotions" refers to algorithms and software for estimating and analyzing a user's emotional state based on their viewing history and reaction data.

[0317] A "summary" is a text that compresses the original book information into 200 characters or less and succinctly presents the main content of the book.

[0318] "Verification measures" refers to processes or algorithms that check and evaluate the accuracy and appropriateness of generated summaries and other data.

[0319] "Means of notification" refers to the technology and process used to communicate new video distribution information, etc. to users' devices.

[0320] "Optimal book summary video" refers to the book summary video that is determined to be most relevant based on the user's interests and emotional state.

[0321] "Means for automatically generating or collecting relevant images" refers to algorithms that generate relevant images from text, or technologies that collect appropriate images from existing databases or the internet.

[0322] This invention relates to a system that acquires book information from a book database or API, generates a summary using that information, and then generates a video based on the summary. The system includes an emotion engine that recognizes a user's emotions and provides optimal content based on those emotions. The following describes in detail the embodiments of the invention.

[0323] Get book data

[0324] The server retrieves the book data.

[0325] The server connects to a book database or API (for example, a common book database API) and periodically retrieves data on new best-selling books. This data includes the book's title, author, synopsis, and publication date. The server runs a scheduled task and stores the retrieved data in an internal database. Specifically, the server retrieves the data from the API at midnight every day and stores it in an internal database (for example, MySQL or MongoDB).

[0326] Summary generation

[0327] The server generates a summary

[0328] The server uses a generative AI model (e.g., GPT-4) to generate a summary of up to 200 characters based on the acquired book information. The input includes the book title, author, and synopsis. For example, data from the book "Introduction to Algorithms for Engineers" is input into the generative AI model, and the resulting summary is, "This book introduces the basic concepts and practical applications of algorithms."

[0329] Video generation

[0330] The server generates a video from the summary text

[0331] The server automatically generates or collects related images based on the verified summary. Depending on the extracted keywords, appropriate images are collected from the Internet or selected from an internal image database. For example, related images are collected based on keywords such as "algorithm" and "application." The server then combines the images with the summary, overlaying text on the images to generate a slideshow-style video. Specifically, video generation software such as FFmpeg is used to generate video files of up to 30 seconds.

[0332] Video distribution

[0333] The server uploads the video to the video streaming service.

[0334] The server automatically uploads the generated video file using the API of a video distribution service (for example, YouTube or Vimeo). When uploading, necessary information (title, description, tags, etc.) is set, and once the upload is complete, the server obtains the URL of the generated video from the video distribution service and saves it in an internal database. For example, the video title could be "Introduction to Algorithms for Engineers - Summary Video" and the description "Introducing the basic concepts of algorithms in 30 seconds."

[0335] User Notifications

[0336] The device notifies the user of the video link

[0337] The server sends a message to the user's device to notify them that a new video has been uploaded. The device (smartphone or computer) then displays a push notification to the user, saying something like, "A new summary video has been released." After checking the notification, the user clicks on the provided video link to access the video streaming service and watch the 30-second summary video.

[0338] Sentiment Analysis and Content Optimization

[0339] The server analyzes the user's emotions

[0340] The emotion engine analyzes user emotions based on their viewing history and reaction data, for example, learning the genres and content of videos they have particularly rated.

[0341] The server delivers the best content

[0342] Based on the analysis results, the server selects and delivers the book summaries and videos that best suit the user's emotions. Specifically, it prioritizes books in the user's favorite genres and themes, and generates and delivers summaries and videos based on that. For example, if a user frequently reads technical books, the server will continue to generate summaries and videos for technical books from the next time.

[0343] Specific examples

[0344] Specific examples of book data acquisition

[0345] The server retrieves data for "Introduction to Algorithms for Engineers" from the book database API at midnight every day. This data includes the book title, author, synopsis, and publication date.

[0346] A concrete example of summary generation

[0347] The server inputs data from "Introduction to Algorithms for Engineers" into a generative AI model and generates a summary statement that reads, "This book introduces the basic concepts and practical applications of algorithms."

[0348] Example of video generation

[0349] The server collects relevant images based on keywords such as "algorithm" and "application," and uses FFmpeg to generate a 30-second video that combines the summary text and images.

[0350] Specific examples of video distribution

[0351] The server uploads the generated video to YouTube, obtains the video URL "https: / / www.youtube.com / watch?v=example", and stores it in the database.

[0352] Example of user notification

[0353] The server sends a push notification to the user's smartphone saying, "A new summary video has been released." The user clicks the notification to watch the video.

[0354] Examples of sentiment analysis and content optimization

[0355] The emotion engine analyzes the user's viewing history, and if they frequently watch technical books, it will prioritize similar genres for summarization and generate videos from them next time.

[0356] Prompt Sentence Examples

[0357] "Book Title: Introduction to Algorithms for Engineers Author: Taro Yamada Summary: This book introduces the basic concepts and practical applications of algorithms."

[0358] This invention allows users to efficiently watch video summaries of books that match their emotions and interests.

[0359] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0360] Step 1:

[0361] Get book data

[0362] The server connects to a book database or API and periodically retrieves data on new bestselling books. First, the server sends a request to a specific API (e.g., a general book database API) at midnight every day. In response to this request, the API returns book information (title, author, synopsis, publication date). The server stores the retrieved data in an internal database (e.g., MySQL or MongoDB).

[0363] Input: API request / response

[0364] Output: Book information stored in the internal database

[0365] Step 2:

[0366] Summary generation

[0367] The server generates a summary using book information retrieved from an internal database. Specifically, it inputs the book title, author, and synopsis into a generative AI model (e.g., GPT-4). Example prompt:

[0368] "Book Title: Introduction to Algorithms for Engineers Author: Taro Yamada Summary: This book introduces the basic concepts and practical applications of algorithms."

[0369] The generative AI model outputs a summary of up to 200 characters, which is then verified by the server to ensure it does not contain inappropriate content.

[0370] Input: Book information, prompt for the generative AI model

[0371] Output: Summary of up to 200 characters

[0372] Step 3:

[0373] Abstract verification

[0374] The server verifies the generated summary. It uses an automatic verification algorithm to check whether the summary contains any prohibited words. It also checks for grammatical errors and inappropriate content. Once the verification is complete, the summary proceeds to the next process.

[0375] Input: Generated summary

[0376] Output: Verified summary

[0377] Step 4:

[0378] Collection of related images

[0379] The server collects related images based on keywords extracted from the abstract. For example, it searches for and retrieves appropriate images from image databases on the Internet or internal image databases based on keywords such as "algorithm" or "application."

[0380] Input: Keywords extracted from the abstract

[0381] Output: Associated image data

[0382] Step 5:

[0383] Combining images and abstracts

[0384] The server combines the collected images with the verified summaries to generate a slideshow-style video. Specifically, it uses video generation software such as FFmpeg to overlay the summaries on the images and create a video file of up to 30 seconds in length.

[0385] Input: Related images, verified summary

[0386] Output: Generated video file

[0387] Step 6:

[0388] Uploading videos

[0389] The server uploads the generated video file to a video distribution service (e.g., YouTube or Vimeo). At this time, information such as the video title, description, and tags are set. Once the upload is complete, the server obtains the video URL from the video distribution service and saves it in an internal database.

[0390] Input: Generated video file, video information (title, description, tags)

[0391] Output: Video URL, saved to internal database

[0392] Step 7:

[0393] User Notifications

[0394] The server sends a message to the user's device to notify them that a new video has been uploaded. The device (smartphone or computer) then displays the message to the user via a push notification. For example, a notification could be sent to the user saying, "A new summary video has been uploaded."

[0395] Input: Video URL, notification message

[0396] Output: Notification displayed on the user's device

[0397] Step 8:

[0398] sentiment analysis

[0399] The server collects user viewing history and reaction data and uses an emotion engine to analyze the user's emotions. For example, it infers the user's emotional state based on data such as ratings, viewing time, and comments on videos viewed in the past.

[0400] Input: Viewing history, reaction data

[0401] Output: User emotion data

[0402] Step 9:

[0403] Content Optimization

[0404] The server selects the most suitable book summary video for the user based on the analysis results of the emotion engine. It prioritizes books in the user's favorite genres and themes, and generates the summary text and video again based on that. This makes it possible to provide personalized content to the user.

[0405] Input: User emotion data, book information

[0406] Output: Optimized book summary video

[0407] (Application example 2)

[0408] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0409] When viewing book summaries and videos, it is difficult to provide optimal content that matches the user's interests. Another issue is that the content presented to the user does not necessarily match the user's preferences. Furthermore, there is a need for a system that can efficiently generate summaries and videos from a vast amount of book information and optimize them based on the user's emotions.

[0410] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0411] In this invention, the server includes means for acquiring book information from a book database or API, a generation AI model means for generating summaries using the acquired book information, means for generating videos based on the generated summaries, means for uploading the generated videos to a video distribution service, means for notifying the user, an emotion engine means for analyzing the user's emotions, and means for optimizing and distributing book summary videos to the user based on the emotion engine means. This makes it possible to provide optimal book summary videos based on the user's emotions and improve the viewing experience.

[0412] A "book database" is a database that accumulates and organizes information about books.

[0413] "API" stands for Application Programming Interface, an interface for exchanging functions and data between different software programs.

[0414] A "generative AI model" is a model that uses artificial intelligence to generate output data from specific input data. In this invention, it refers to a model that generates a summary from book information.

[0415] The "emotion engine" is an engine that analyzes users' emotions and reactions and provides optimal content based on that data.

[0416] A "video distribution service" is a service that distributes video content over the Internet.

[0417] A "summary" is a sentence that summarizes and shortens a long sentence, and in the present invention, it is intended to express the main content of a book concisely.

[0418] "Optimization" means adjusting and improving to best suit a specific purpose or condition. In this invention, it refers to optimizing the book summary video based on user sentiment.

[0419] "Notification" is the act of informing a user of a certain fact or information, and in the present invention, it is intended to notify the user of the distribution of a new digest video.

[0420] This system acquires book information from a book database or API, generates a summary based on the acquired information, and generates a video based on the summary. Furthermore, this system incorporates an emotion engine that recognizes the user's emotions, and provides the user with an optimal book summary video based on the acquired emotion information.

[0421] Get book data

[0422] The server connects to a specific book database or API to retrieve data on new bestselling books, including the book's title, author, synopsis, and publication date. The server runs a scheduled task at a regular time to retrieve the data and store it in an internal database.

[0423] Summary generation

[0424] The server uses a generative AI model to generate a summary of up to 200 characters based on the acquired book information. By inputting the book title, author, and synopsis into the generative AI model, the AI ​​model outputs a summary. This summary is then verified by the server to ensure it does not contain any inappropriate content.

[0425] Video generation

[0426] The server automatically generates or collects relevant images based on the verified summary. Depending on the extracted keywords, appropriate images are collected from the Internet or selected from an internal image database. The server then combines these images with the summary, overlaying text on the images to generate a slideshow-style video. Video generation software (e.g., FFmpeg) is used to generate a video file of up to 30 seconds.

[0427] Video distribution

[0428] The server automatically uploads the generated video file to a video distribution service (for example, YouTube or Vimeo). The server uses the API of the video distribution service to set the information required when uploading the video (title, description, tags, etc.). Once the upload is complete, the server obtains the URL of the generated video from the video distribution service and saves it in its internal database.

[0429] User Notifications

[0430] The server sends a message to the user's device to notify them that a new video has been uploaded. The device (smartphone or computer) then displays a push notification to the user, saying something like, "A new summary video has been released." After checking the notification, the user clicks on the provided video link to access the video streaming service and watch the 30-second summary video.

[0431] Sentiment Analysis and Content Optimization

[0432] The emotion engine analyzes a user's emotions based on their viewing history and reaction data. For example, the emotion engine learns which genres and content of videos a user has particularly rated highly among those they have watched in the past. Based on the emotion data analyzed by the emotion engine, the server selects and prioritizes the book summaries and videos that best match the user's emotions. Specifically, the server prioritizes books in the user's favorite genres and themes, and generates and distributes summaries and videos based on these.

[0433] Specific examples

[0434] Get book data

[0435] For example, suppose a server retrieves data on bestselling books from an API at midnight every day. For example, suppose a book called "Algorithms for Engineers" is retrieved.

[0436] Summary generation

[0437] The server inputs the data from the above book into a generative AI model and generates a summary like the one below.

[0438] Title: Algorithms for Engineers

[0439] Author: Example author

[0440] Summary: This book introduces fundamental concepts and practical applications of algorithms.

[0441] Generated summary:

[0442] This book introduces fundamental concepts and practical applications of algorithms.

[0443] Video generation

[0444] The server collects related images based on keywords such as "algorithm" and "practice," and combines the summary text and images to generate a 30-second video.

[0445] Video distribution

[0446] The server uploads the generated video to a video distribution service (e.g., YouTube), obtains the video URL, and stores it in a database.

[0447] User Notifications

[0448] The server sends a push notification to the user's smartphone, and the user clicks on the notification to watch the video.

[0449] Sentiment Analysis and Content Optimization

[0450] If the emotion engine analyzes the user's viewing history and recognizes that the user prefers technical books, for example, it will prioritize generating and delivering summaries and videos of technical books from the next time onwards.

[0451] This allows users to efficiently watch video summaries of books that match their interests and emotions.

[0452] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0453] Step 1:

[0454] The server connects to a book database or API to retrieve new book information. It receives a JSON response containing data such as the book title, author, synopsis, and publication date. For example, it sends a GET request to an API and receives JSON data containing book information in response. It then stores this data in an internal database.

[0455] Input: API endpoint

[0456] Output: Book information (title, author, synopsis, publication date)

[0457] Step 2:

[0458] The server generates a prompt based on the acquired book information and inputs it into the generative AI model. The prompt includes the book title, author, and synopsis.

[0459] Example prompt sentence:

[0460] Title: Algorithms for Engineers

[0461] Author: Example author

[0462] Summary: This book introduces fundamental concepts and practical applications of algorithms.

[0463] As a result, a summary sentence is generated from the AI ​​model.

[0464] Input: Book information (title, author, synopsis)

[0465] Output: Summary

[0466] Step 3:

[0467] The server then validates the generated abstract to ensure it is correct, including checking for any invalid language or inappropriate content. If the abstract passes validation, it is sent to the next step.

[0468] Input: Abstract

[0469] Output: Verified summary

[0470] Step 4:

[0471] The server collects related images based on the verified summary, picks out important words from the summary and synopsis to extract keywords, and collects images from the Internet and internal image databases based on these keywords.

[0472] Input: Verified abstract, keywords

[0473] Output: A list of related images

[0474] Step 5:

[0475] The server combines the summary text with the collected images to generate a slideshow-style video using FFmpeg or other video generation tools. The summary text is overlaid on the images to create a video file.

[0476] Input: Verified summary, related images

[0477] Output: Video file

[0478] Step 6:

[0479] The server uploads the generated video file to a video distribution service. At this time, metadata such as the video title, description, and tags are set. For example, the YouTube API is used to upload the video and obtain its URL.

[0480] Input: Video file, metadata

[0481] Output: Video URL

[0482] Step 7:

[0483] The server saves the URL of the new video in an internal database and sends that information to the user's device via a push notification service (e.g., Firebase Cloud Messaging), informing the user that a new summary video has been released.

[0484] Input: Video URL

[0485] Output: Push notification message

[0486] Step 8:

[0487] The server collects users' viewing history and reaction data, and analyzes their emotions using an emotion engine, which determines what genres and content are most suitable for the user.

[0488] Input: User viewing history, reaction data

[0489] Output: Sentiment analysis data

[0490] Step 9:

[0491] Based on the data analyzed by the emotion engine, the server selects the next best book summary video and starts the process of generating and delivering that video, ensuring that content that matches the user's emotions is provided first.

[0492] Input: Sentiment analysis data

[0493] Output: Recommended book information, summary, video URL

[0494] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0495] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0496] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0497] [Second embodiment]

[0498] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0499] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0500] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0501] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0502] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0503] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0504] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0505] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0506] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0507] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0508] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0509] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0510] A natural language description of the program's operation

[0511] This system provides an automated process to summarize international best-selling books in 200 characters or less and distribute them as 30-second videos. The following explains the roles and processes of the server, terminal, and user in detail.

[0512] Get book data

[0513] The server retrieves the book data.

[0514] First, the server connects to a specific book database or API (Application Program Interface) to retrieve new book information. The server runs a scheduled task at a regular time, for example, retrieving the latest best-selling book list at midnight every day. The retrieved data includes the book title, author, synopsis, publication date, etc.

[0515] Summary generation

[0516] The server generates a summary

[0517] Next, the server generates a summary using a generative AI model (such as GPT-3) based on the acquired book information. The server inputs the book title, author, and synopsis into the generative AI model in a prompt format, and the AI ​​model outputs a summary of up to 200 characters. The generated summary is then verified by the server to ensure it does not contain any inappropriate content.

[0518] Video generation

[0519] The server generates a video from the summary text

[0520] Based on the verified summary, the server automatically generates or collects related images. The server extracts keywords from the summary and, depending on the keywords, collects related images from the Internet or selects them from an internal image database. The server then combines these images with the summary, overlaying text on the images to generate a slideshow-style video. Video generation software such as FFmpeg is used to generate video files of up to 30 seconds.

[0521] Video distribution

[0522] The server uploads the video to the video streaming service.

[0523] The server automatically uploads the generated video file to a video distribution service (such as YouTube or Vimeo). The server uses the video distribution service's API to set the information required when uploading the video (title, description, tags, etc.). Once the upload is complete, the video's URL is obtained from the service and saved in an internal database.

[0524] User Notifications

[0525] The device notifies the user of the video link

[0526] The server sends a message to the user's device to notify them that a new video has been uploaded. The device (smartphone or computer) then displays a push notification to the user with a message such as "A new summary video has been released." After confirming the notification, the user clicks on the provided video link to access the video streaming service and watch the 30-second summary video.

[0527] Specific examples

[0528] 1. Acquiring book data

[0529] The server retrieves the bestselling book data from the API at midnight every day. For example, it retrieves the data for the book "The Silent Patient."

[0530] 2. Summary Generation

[0531] The server inputs the data from "The Silent Patient" into a generative AI model and generates the summary: "Alicia is an artist who goes silent after a mysterious incident, and the story is about a therapist who searches for the truth behind it."

[0532] 3. Video Generation

[0533] The server collects free images based on keywords such as "Alicia" and "therapist," and combines the summary text and images to generate a 30-second video.

[0534] 4. Video distribution

[0535] The server uploads the generated video to YouTube, retrieves the video URL and stores it in the database.

[0536] 5. User Notices

[0537] The server sends a push notification to the user's smartphone, and the user clicks on the notification to watch the video.

[0538] This system allows users to easily view summaries of international best-selling books in a short amount of time, enabling them to acquire information efficiently.

[0539] The processing flow will be explained below.

[0540] Step 1:

[0541] The server connects to the book database or API at midnight every day to retrieve the latest bestselling book data, including the book title, author, synopsis, and publication date.

[0542] Step 2:

[0543] The server stores the acquired book information in an internal database and adds it to a processing queue, so that subsequent processes can retrieve and process the book information sequentially.

[0544] Step 3:

[0545] The server retrieves a book from the processing queue and inputs the book's title, author, and synopsis into the generative AI model, providing appropriate input to the AI ​​model in the form of prompts.

[0546] Step 4:

[0547] The generative AI model outputs a summary of up to 200 characters based on the input information. The server receives the summary and prepares it for the next process.

[0548] Step 5:

[0549] The server then validates the generated summary to check for inappropriate content, for example, by checking for inappropriate language or unclear sentences.

[0550] Step 6:

[0551] The server extracts keywords from the verified abstract and passes the list of keywords, including people's names, places, and events, to the image generation module.

[0552] Step 7:

[0553] The image generation module collects free stock images corresponding to each keyword from the Internet or selects them from an internal image database, and the server downloads or retrieves the images.

[0554] Step 8:

[0555] The server combines the abstract with the acquired images to generate a slideshow-style video by overlaying the text on the images. Video generation software (e.g., FFmpeg) is used to generate a video file of up to 30 seconds.

[0556] Step 9:

[0557] The server uploads the generated video file to a video distribution service (e.g., YouTube or Vimeo), and also sets the video's metadata (title, description, tags, etc.) at the same time.

[0558] Step 10:

[0559] Once the upload is complete, the server retrieves the generated video URL from the video streaming service, which is then stored in an internal database.

[0560] Step 11:

[0561] The server sends a message to the user's device informing them that a new video has been uploaded, which includes sending a push notification.

[0562] Step 12:

[0563] The device displays a notification to the user saying, "A new summary video has been released." The user checks the notification and clicks on the provided video link.

[0564] Step 13:

[0565] Users can easily obtain summaries of international best-selling books by watching 30-second summary videos on their device's browser or YouTube app.

[0566] This detailed processing flow automates and efficiently executes the entire process of obtaining book information, generating a summary and video, and finally notifying the user.

[0567] Example 1

[0568] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0569] Current book information systems do not allow users to quickly obtain summaries of international best-selling books, making it difficult for users to grasp the content in a short time. Furthermore, when summaries are provided only in text, the lack of visual elements makes it difficult for users to understand or engage with the information. Therefore, there is a need for a method to provide book information efficiently and in a visually understandable way.

[0570] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0571] In this invention, the server includes means for acquiring book information from a book database or API, means for creating a prompt based on the acquired book information and sending it to the generative AI model, means for verifying the summary returned from the generative AI model, means for extracting keywords from the summary and collecting related images, means for generating a slideshow video based on the collected images and summary, means for uploading the generated video to a video distribution service, and means for sending a notification message including a link to the video to the user terminal. This allows users to quickly and visually understand summaries of international best-selling books.

[0572] A "book database" is a system that systematically stores information about books and has a data structure that allows for searching and extraction.

[0573] An "API" is an interface that allows various pieces of software to communicate with each other and share functions, and is often provided in the form of a web service.

[0574] A "prompt sentence" is an input sentence provided to a generative AI model, a series of text that is generated in anticipation of specific conditions or responses.

[0575] A "generative AI model" is an artificial intelligence algorithm that automatically generates text and information from large amounts of data.

[0576] "Validation" is the process of verifying whether the generated data is appropriate and checking that it does not contain any inappropriate content.

[0577] "Keywords" are important words extracted from summaries and texts, and are used for information retrieval and classification.

[0578] "Image gathering" is the process of acquiring relevant image data based on specific criteria or keywords.

[0579] "Slideshow format" is a visual presentation format that displays multiple images and text sequentially.

[0580] A "video distribution service" is an online platform for uploading videos over the Internet and making them available to viewers.

[0581] A "notification message" is a message sent by the system to convey specific information to the user.

[0582] A "user terminal" is a device such as a smartphone or computer that allows a user to access the Internet and applications.

[0583] To implement this invention, a system is required that automates the process of acquiring book information from a book database or API, generating a summary using a generative AI model based on that information, creating a video using that summary, and finally notifying the user. The following describes in detail the hardware and software used in each processing step, as well as the details of data processing and data calculation.

[0584] Get book data

[0585] Server Roles

[0586] The server connects to a specific book database or API to retrieve new book information. For example, it can use the Goodreads or Google Books API. The server runs a periodic task every day at midnight, sending an HTTP request to retrieve the latest best-selling book list. The retrieved data is provided in JSON format and includes the book title, author, synopsis, publication date, etc.

[0587] Summary generation

[0588] Server Roles

[0589] The server creates a prompt based on the acquired book information and sends it to the generative AI model. Specifically, it combines the book title, author, and synopsis to create a prompt to be input to the generative AI model (e.g., GPT-3). The following is an example of a prompt:

[0590] Example prompt sentence:

[0591] "Book title: 'XXX', author: XXX, summary: XXX. Please write a summary in 200 characters or less."

[0592] The server receives the summary returned by the generative AI model and uses an internal validation algorithm to check for inappropriate content. If there are no problems, the summary is stored in the database.

[0593] Video generation

[0594] Server Roles

[0595] The server extracts keywords from the summary and collects related images. It uses a natural language processing algorithm to extract key keywords from the summary and collects images via HTTP requests from free image services (such as Unsplash and Pexels). It then combines the collected images with the summary to generate a slideshow-style video using video generation software such as FFmpeg. The video must be no longer than 30 seconds.

[0596] Video distribution

[0597] Server Roles

[0598] The server uploads the generated video to a video distribution service (such as YouTube or Vimeo). The server uses the API of the video distribution service to set the necessary information such as title, description, and tags when uploading the video. If the upload is successful, the server obtains the URL of the video returned by the video service and saves it in the database.

[0599] User Notifications

[0600] Device Role

[0601] The server generates a notification message containing the URL of the new video and sends a push notification to the user's device. The device (smartphone or computer) then displays a notification to the user saying, "A new summary video has been released." The user can click the notification to access the video streaming service and watch a 30-second summary video.

[0602] Specific examples

[0603] For example, for the book "Silent Patient," the server sends an HTTP request to the Goodreads API at midnight every day to retrieve information about the book. Next, the following prompt is sent to the generative AI model: "Book Title: 'Silent Patient,' Author: Alex Michaelides, Synopsis: Alicia is an artist who falls silent after a mysterious incident, and the story of a therapist who searches for the truth behind this. Please provide a summary of 200 characters or less." Based on the summary returned by the generative AI model, "Alicia is an artist who falls silent after a mysterious incident, and the story of a therapist who searches for the truth behind this," the server extracts keywords such as "Alicia" and "therapist" and collects related free images using the Unsplash API. These images and the summary are then processed into a slideshow video using FFmpeg. The server then uploads this video through the YouTube API and gives it a title such as "Silent Patient Summary Video." The user's device receives a notification stating, "A new summary video of 'Silent Patient' has been uploaded," and they can click the notification to watch the video.

[0604] This system allows users to quickly obtain summaries of international best-selling books in a visually easy-to-understand manner.

[0605] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0606] Step 1:

[0607] Get book data

[0608] The server connects to the book database or API to retrieve new book information. The server sends an HTTP request at midnight every day to retrieve the latest bestselling book list. The retrieved data is provided in JSON format and includes the book title, author, synopsis, publication date, etc. The input is the book information retrieved from the book database or API, and the output is the parsed book information. The server parses the retrieved data and saves it in its internal database.

[0609] Step 2:

[0610] Creating and sending prompts

[0611] The server creates a prompt based on the acquired book information and sends it to the generative AI model. Specifically, it combines the book title, author, and synopsis to create a prompt to be input to the generative AI model (e.g., GPT-3). The input is the analyzed book information, and the output is the generated prompt. Below is an example of a prompt:

[0612] Example prompt sentence:

[0613] "Book title: 'XXX', author: XXX, summary: XXX. Please write a summary in 200 characters or less."

[0614] Step 3:

[0615] Summary generation and verification

[0616] The server receives the summary returned by the generative AI model. The server checks the generated summary with an internal verification algorithm to ensure that it does not contain inappropriate content. The input is the summary returned by the generative AI model, and the output is the summary after verification. If it is confirmed that it does not contain inappropriate content, the summary is saved in the database.

[0617] Step 4:

[0618] Keyword extraction and image collection

[0619] The server extracts keywords from the abstract and collects related images. It uses a natural language processing algorithm to extract key keywords from the abstract and collects images via HTTP requests from services that provide free images (e.g., Unsplash and Pexels). The input is the verified abstract, and the output is the extracted keywords and collected images.

[0620] Step 5:

[0621] Video generation

[0622] The server generates a slideshow-style video based on the collected images and summaries. The collected images are concatenated and the summaries are overlaid on the images as text. Video generation software such as FFmpeg is used to generate videos of up to 30 seconds. The input is the collected images and summaries, and the output is the generated video file.

[0623] Step 6:

[0624] Uploading videos

[0625] The server uploads the generated video to a video distribution service. Using the API of the video distribution service (for example, YouTube or Vimeo), necessary information such as title, description, and tags is set when uploading the video. The input is the generated video file, and the output is the URL of the uploaded video. If the upload is successful, the server saves the video URL returned by the video service in a database.

[0626] Step 7:

[0627] User Notifications

[0628] The server generates a notification message containing the URL of the new video and sends a push notification to the user's device. The device (smartphone or computer) receives the notification and displays a message to the user saying, "A new summary video has been released." The input is the URL of the uploaded video, and the output is the notification message displayed on the user's device. The user can click the notification to access the video streaming service and watch a 30-second summary video.

[0629] (Application example 1)

[0630] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0631] In today's world, busy people want to obtain information efficiently within a limited amount of time, but it is often difficult to find the time to read long books. Another issue is that simply reading a book summary can weaken comprehension of the content and the impact of the information. Furthermore, there are limited ways for users to easily view summarized book information, and this needs to be addressed.

[0632] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0633] In this invention, the server includes a means for acquiring book information from a book database or an API, a generation AI model means for generating a summary using the acquired book information, and a means for generating a video based on the generated summary, thereby enabling users to efficiently view summarized book information in video format.

[0634] "Book database" refers to one or more data collections that store information about books.

[0635] "API" refers to an interface that allows application programs to communicate with each other.

[0636] A "generative AI model" refers to artificial intelligence that automatically generates text, images, videos, etc. based on specified input data.

[0637] A "summary" is a short summary of the book's contents.

[0638] "Video distribution service" refers to a service that provides video content to users via the Internet.

[0639] "User" refers to an entity that uses the system to receive information or services.

[0640] "Smart devices" refer to mobile information devices with internet connectivity, such as smartphones and tablets.

[0641] MODE FOR CARRYING OUT THE INVENTION

[0642] System Program

[0643] This system retrieves book information from a book database or API, and generates a summary based on that information using a generative AI model. It consists of a series of processes: generating a video based on the generated summary, uploading the video to a video distribution service, and sending a notification to the user. Users can also watch the summary video on their smart devices. The program design for realizing this system is as follows:

[0644] Hardware and software used

[0645] Server hardware: Amazon EC2

[0646] Book data acquisition: Google Books API

[0647] Summary sentence generation: OpenAI GPT-3

[0648] Video generation: FFmpeg

[0649] Video streaming: YouTube Data API

[0650] Push notifications: Firebase Cloud Messaging (FCM)

[0651] Smart devices: smartphones, tablets, etc.

[0652] Data processing and calculation

[0653] The details of how this system works are as follows:

[0654] 1. Acquiring book data

[0655] The server runs a scheduled task at regular intervals, connecting to a book database or API (e.g., Google Books API) to retrieve the latest best-selling book list, including book title, author, synopsis, publication date, etc.

[0656] 2. Summary Generation

[0657] Based on the acquired book information, the server uses a generative AI model (e.g., OpenAI GPT-3) to generate a summary. Specifically, the book title, author, and synopsis are input into the generative AI model in the form of a prompt, and the AI ​​model outputs a summary of up to 200 characters. This summary is then verified by the server to ensure it does not contain any inappropriate content.

[0658] 3. Video Generation

[0659] Based on the verified summary, the server automatically generates or collects related images. The server extracts keywords from the summary and, based on those keywords, collects related images from the Internet or selects them from an internal image database. It then combines these images with the summary, overlaying text on the images to generate a slideshow-style video. It uses video generation software such as FFmpeg to generate video files of up to 30 seconds.

[0660] 4. Video distribution

[0661] The server automatically uploads the generated video file using the API of a video distribution service (such as YouTube or Vimeo). The necessary information (title, description, tags, etc.) is also set when uploading. Once the upload is complete, the video URL is obtained from the service and saved in an internal database.

[0662] 5. User Notices

[0663] The server sends a message to the user's smart device to notify them that a new video has been uploaded. A push notification displays a message to the user, such as "A new summary video has been uploaded." The user checks the notification and clicks on the provided video link to access the video streaming service and watch the 30-second summary video.

[0664] Specific examples

[0665] 1. Acquiring book data

[0666] Book: "The Silent Patient"

[0667] Author: "Alex Michaelides"

[0668] Synopsis: "Alicia Belson is a celebrated artist..."

[0669] 2. Summary Generation

[0670] Prompt statement:

[0671] Title: The Silent Patient

[0672] Author: Alex Michaelides

[0673] Summary: Alicia Belson is a renowned artist whose work is celebrated around the world. One day, she shoots and kills her husband in their home and never speaks again. This story is told from the perspective of a therapist who tries to unravel this mystery.

[0674] Please summarize in 200 characters or less.

[0675] Generated summary:

[0676] Alicia is an artist who goes silent after a mysterious incident, and the story revolves around a therapist who searches for the truth.

[0677] 3. Video Generation

[0678] The generated summary is combined with images related to keywords such as "Alicia" and "therapist," and a video of less than 30 seconds is generated using FFmpeg.

[0679] 4. Video distribution

[0680] Upload a video to YouTube, get the URL and save it in the database.

[0681] 5. User Notices

[0682] Use Firebase Cloud Messaging to send notifications of new video streams to users' smart devices.

[0683] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0684] Step 1:

[0685] The server periodically connects to a book database or API to retrieve the latest best-selling book information, including book title, author, synopsis, and publication date. This allows you to always have access to the latest book information.

[0686] Input: Scheduled task, book database or API

[0687] Output: Retrieved book information (title, author, synopsis, publication date)

[0688] Step 2:

[0689] The server generates a summary using a generative AI model (e.g., OpenAI GPT-3) based on the acquired book information. The book title, author, and synopsis are input into the generative AI model in the form of a prompt, and the AI ​​outputs a summary of up to 200 characters. At this time, the summary is verified to ensure that it does not contain any inappropriate content.

[0690] Input: Book information (title, author, synopsis), generative AI model

[0691] Output: Summary of up to 200 characters

[0692] Step 3:

[0693] The server extracts keywords from the verified summary and uses them to gather relevant images from the internet or select them from an internal image database. It then combines the summary with these images and generates a slideshow-style video by overlaying text on the images. It then uses video generation software such as FFmpeg to generate a video file of up to 30 seconds.

[0694] Input: Abstract, keywords, related images

[0695] Output: Video file up to 30 seconds long

[0696] Step 4:

[0697] The server uploads the generated video via the API of a video distribution service (e.g., YouTube). Information such as the title, description, and tags are also set. Once the upload is complete, the video URL is obtained from the service and saved in an internal database.

[0698] Input: Video file, video streaming service API

[0699] Output: Video URL

[0700] Step 5:

[0701] The server sends a push notification message via Firebase Cloud Messaging to the user's smart device to notify them that a new video has been uploaded, and the user receives the notification and clicks the video link to watch the video.

[0702] Input: Video URL, Firebase Cloud Messaging

[0703] Output: User notification

[0704] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0705] A natural language description of the program's operation

[0706] This system acquires book information from a book database or API, generates a summary based on the acquired information, and generates a video based on the summary. Furthermore, this system incorporates an emotion engine that recognizes the user's emotions, and provides the user with an optimal book summary video based on the acquired emotion information.

[0707] Get book data

[0708] The server retrieves the book data.

[0709] The server connects to a specific book database or API to retrieve data on new bestselling books, including the book's title, author, synopsis, and publication date. The server runs a scheduled task at a regular time to retrieve the data and store it in an internal database.

[0710] Summary generation

[0711] The server generates a summary

[0712] The server uses a generative AI model to generate a summary of up to 200 characters based on the acquired book information. By inputting the book title, author, and synopsis into the generative AI model, the AI ​​model outputs a summary. This summary is then verified by the server to ensure it does not contain any inappropriate content.

[0713] Video generation

[0714] The server generates a video from the summary text

[0715] Based on the verified summary, the server automatically generates or collects relevant images. Depending on the extracted keywords, appropriate images are collected from the Internet or selected from an internal image database. The server then combines these images with the summary, overlaying text on the images to generate a slideshow-style video. Video generation software (e.g., FFmpeg) is used to generate a video file of up to 30 seconds.

[0716] Video distribution

[0717] The server uploads the video to the video streaming service.

[0718] The server automatically uploads the generated video file to a video distribution service (for example, YouTube or Vimeo). The server uses the API of the video distribution service to set the information required when uploading the video (title, description, tags, etc.). Once the upload is complete, the server obtains the URL of the generated video from the video distribution service and saves it in its internal database.

[0719] User Notifications

[0720] The device notifies the user of the video link

[0721] The server sends a message to the user's device to notify them that a new video has been uploaded. The device (smartphone or computer) then displays a push notification to the user, saying something like, "A new summary video has been released." After checking the notification, the user clicks on the provided video link to access the video streaming service and watch the 30-second summary video.

[0722] Sentiment Analysis and Content Optimization

[0723] The server analyzes the user's emotions

[0724] The emotion engine analyzes user emotions based on their viewing history and reaction data. For example, it learns the genres and content that users have particularly rated among the videos they have watched in the past.

[0725] The server delivers the best content

[0726] Based on the emotional data analyzed by the emotion engine, the server selects and prioritizes the book summaries and videos that best fit the user's emotions. Specifically, it prioritizes books in the user's favorite genres and themes, and generates and distributes summaries and videos based on those.

[0727] Specific examples

[0728] 1. Acquiring book data

[0729] The server retrieves data on bestselling books from the API at midnight every day. For example, it retrieves data on the book "Introduction to Algorithms for Engineers."

[0730] 2. Summary Generation

[0731] The server inputs data from "Introduction to Algorithms for Engineers" into a generative AI model and generates a summary statement that reads, "This book introduces the basic concepts and practical applications of algorithms."

[0732] 3. Video Generation

[0733] The server collects related images based on keywords such as "algorithm" and "application," and combines the summary text and images to generate a 30-second video.

[0734] 4. Video distribution

[0735] The server uploads the generated video to YouTube, retrieves the video URL and stores it in the database.

[0736] 5. User Notices

[0737] The server sends a push notification to the user's smartphone, and the user clicks on the notification to watch the video.

[0738] 6. Sentiment Analysis and Content Optimization

[0739] The emotion engine analyzes the user's viewing history, and if, for example, the user often watches technical books, from the next time onwards, it will prioritize summarizing technical books and generating and delivering videos.

[0740] This detailed process allows users to efficiently watch video summaries of books that suit their preferences and emotions.

[0741] The processing flow will be explained below.

[0742] Step 1:

[0743] The server connects to the book database or API at midnight every day to retrieve the latest bestselling book data, including the book title, author, synopsis, and publication date.

[0744] Step 2:

[0745] The server stores the acquired book information in an internal database and adds it to a processing queue, so that subsequent processes can retrieve and process the book information sequentially.

[0746] Step 3:

[0747] The server retrieves a book from the processing queue and inputs the book's title, author, and synopsis into the generative AI model, providing appropriate input to the AI ​​model in the form of prompts.

[0748] Step 4:

[0749] The generative AI model outputs a summary of up to 200 characters based on the input information. The server receives the summary and prepares it for the next process.

[0750] Step 5:

[0751] The server then validates the generated summary to check for inappropriate content, for example, by checking for inappropriate language or unclear sentences.

[0752] Step 6:

[0753] The server extracts keywords from the verified abstract and passes the list of keywords, including people's names, places, and events, to the image generation module.

[0754] Step 7:

[0755] The image generation module collects free stock images corresponding to each keyword from the Internet or selects them from an internal image database, and the server downloads or retrieves the images.

[0756] Step 8:

[0757] The server combines the abstract with the acquired images to generate a slideshow-style video by overlaying the text on the images. Video generation software (e.g., FFmpeg) is used to generate a video file of up to 30 seconds.

[0758] Step 9:

[0759] The server uploads the generated video file to a video distribution service (e.g., YouTube or Vimeo), and also sets the video's metadata (title, description, tags, etc.) at the same time.

[0760] Step 10:

[0761] Once the upload is complete, the server retrieves the generated video URL from the video streaming service, which is then stored in an internal database.

[0762] Step 11:

[0763] The server sends a message to the user's device informing them that a new video has been uploaded, which includes sending a push notification.

[0764] Step 12:

[0765] The device displays a notification to the user saying, "A new summary video has been released." The user checks the notification and clicks on the provided video link.

[0766] Step 13:

[0767] Users can easily obtain summaries of international best-selling books by watching 30-second summary videos on their device's browser or YouTube app.

[0768] Sentiment Analysis and Content Optimization

[0769] Step 14:

[0770] The server collects the user's viewing history and reaction data and sends it to the emotion engine, which analyzes the user's viewing history and reaction data to estimate the user's emotions.

[0771] Step 15:

[0772] Based on the emotional data analyzed by the emotion engine, the server selects the most suitable book summary and video for the user. Specifically, it prioritizes books in genres and themes that the user has given high ratings to, and generates summaries and videos based on those.

[0773] Step 16:

[0774] The server uploads the generated optimal book summary video to a video distribution service and obtains a distribution link.

[0775] Step 17:

[0776] The server sends push notifications to the user's device at the appropriate time, allowing the user to receive a summary video that matches their emotions.

[0777] Step 18:

[0778] The device receives a notification and displays a message to the user saying, "A summary video has been delivered that is recommended for you." The user clicks the link to watch the video.

[0779] This allows the emotional engine to be used to provide summarized videos based on the user's emotions, resulting in even higher satisfaction. To give a concrete example, for example, a user who frequently reads technical books could be given preferential access to videos summarizing new best-selling books in the same genre. In this way, more personalized content is provided to users through sentiment analysis and content optimization.

[0780] Example 2

[0781] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0782] Conventional book recommendation systems have the problem that the process of generating a book summary and distributing it as a video incorporating visuals is complicated and difficult to automate. It is also difficult to provide personalized content that takes into account the user's preferences and emotions. This makes it difficult to attract the user's interest, resulting in a poor content consumption experience.

[0783] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for acquiring book information from a book database or API, a generation AI model means for generating a summary using the acquired book information, means for verifying the generated summary, means for automatically generating or collecting related images based on the verified summary, means for generating a video by combining the generated images and summary, means for uploading the generated video to a video distribution service, means for acquiring and saving the URL of the uploaded video, means for notifying the user, means for analyzing the user's emotions, and means for providing an optimal book summary video based on the analysis results. This automates the process from generating book summaries to distributing videos and providing content optimized for users, enabling the provision of an efficient and high-quality user experience.

[0784] "Book information" refers to basic data such as the book's title, author, synopsis, and publication date.

[0785] A "generative AI model" refers to an algorithm or software that automatically generates a summary from input data using natural language processing technology.

[0786] "Video Streaming Service" means a platform for hosting generated video content online and for users to stream or download it.

[0787] A "means" refers to a specific method, device, or technique for performing a particular function.

[0788] "Means for analyzing emotions" refers to algorithms and software for estimating and analyzing a user's emotional state based on their viewing history and reaction data.

[0789] A "summary" is a text that compresses the original book information into 200 characters or less and succinctly presents the main content of the book.

[0790] "Verification measures" refers to processes or algorithms that check and evaluate the accuracy and appropriateness of generated summaries and other data.

[0791] "Means of notification" refers to the technology and process used to communicate new video distribution information, etc. to users' devices.

[0792] "Optimal book summary video" refers to the book summary video that is determined to be most relevant based on the user's interests and emotional state.

[0793] "Means for automatically generating or collecting relevant images" refers to algorithms that generate relevant images from text, or technologies that collect appropriate images from existing databases or the internet.

[0794] This invention relates to a system that acquires book information from a book database or API, generates a summary using that information, and then generates a video based on the summary. The system includes an emotion engine that recognizes a user's emotions and provides optimal content based on those emotions. The following describes in detail the embodiments of the invention.

[0795] Get book data

[0796] The server retrieves the book data.

[0797] The server connects to a book database or API (for example, a common book database API) and periodically retrieves data on new best-selling books. This data includes the book's title, author, synopsis, and publication date. The server runs a scheduled task and stores the retrieved data in an internal database. Specifically, the server retrieves the data from the API at midnight every day and stores it in an internal database (for example, MySQL or MongoDB).

[0798] Summary generation

[0799] The server generates a summary

[0800] The server uses a generative AI model (e.g., GPT-4) to generate a summary of up to 200 characters based on the acquired book information. The input includes the book title, author, and synopsis. For example, data from the book "Introduction to Algorithms for Engineers" is input into the generative AI model, and the resulting summary is, "This book introduces the basic concepts and practical applications of algorithms."

[0801] Video generation

[0802] The server generates a video from the summary text

[0803] The server automatically generates or collects related images based on the verified summary. Depending on the extracted keywords, appropriate images are collected from the Internet or selected from an internal image database. For example, related images are collected based on keywords such as "algorithm" and "application." The server then combines the images with the summary, overlaying text on the images to generate a slideshow-style video. Specifically, video generation software such as FFmpeg is used to generate video files of up to 30 seconds.

[0804] Video distribution

[0805] The server uploads the video to the video streaming service.

[0806] The server automatically uploads the generated video file using the API of a video distribution service (for example, YouTube or Vimeo). When uploading, necessary information (title, description, tags, etc.) is set, and once the upload is complete, the server obtains the URL of the generated video from the video distribution service and saves it in an internal database. For example, the video title could be "Introduction to Algorithms for Engineers - Summary Video" and the description "Introducing the basic concepts of algorithms in 30 seconds."

[0807] User Notifications

[0808] The device notifies the user of the video link

[0809] The server sends a message to the user's device to notify them that a new video has been uploaded. The device (smartphone or computer) then displays a push notification to the user, saying something like, "A new summary video has been released." After checking the notification, the user clicks on the provided video link to access the video streaming service and watch the 30-second summary video.

[0810] Sentiment Analysis and Content Optimization

[0811] The server analyzes the user's emotions

[0812] The emotion engine analyzes user emotions based on their viewing history and reaction data, for example, learning the genres and content of videos they have particularly rated.

[0813] The server delivers the best content

[0814] Based on the analysis results, the server selects and delivers the book summaries and videos that best suit the user's emotions. Specifically, it prioritizes books in the user's favorite genres and themes, and generates and delivers summaries and videos based on that. For example, if a user frequently reads technical books, the server will continue to generate summaries and videos for technical books from the next time.

[0815] Specific examples

[0816] Specific examples of book data acquisition

[0817] The server retrieves data for "Introduction to Algorithms for Engineers" from the book database API at midnight every day. This data includes the book title, author, synopsis, and publication date.

[0818] A concrete example of summary generation

[0819] The server inputs data from "Introduction to Algorithms for Engineers" into a generative AI model and generates a summary statement that reads, "This book introduces the basic concepts and practical applications of algorithms."

[0820] Example of video generation

[0821] The server collects relevant images based on keywords such as "algorithm" and "application," and uses FFmpeg to generate a 30-second video that combines the summary text and images.

[0822] Specific examples of video distribution

[0823] The server uploads the generated video to YouTube, obtains the video URL "https: / / www.youtube.com / watch?v=example", and stores it in the database.

[0824] Example of user notification

[0825] The server sends a push notification to the user's smartphone saying, "A new summary video has been released." The user clicks the notification to watch the video.

[0826] Examples of sentiment analysis and content optimization

[0827] The emotion engine analyzes the user's viewing history, and if they frequently watch technical books, it will prioritize similar genres for summarization and generate videos from them next time.

[0828] Prompt Sentence Examples

[0829] "Book Title: Introduction to Algorithms for Engineers Author: Taro Yamada Summary: This book introduces the basic concepts and practical applications of algorithms."

[0830] This invention allows users to efficiently watch video summaries of books that match their emotions and interests.

[0831] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0832] Step 1:

[0833] Get book data

[0834] The server connects to a book database or API and periodically retrieves data on new bestselling books. First, the server sends a request to a specific API (e.g., a general book database API) at midnight every day. In response to this request, the API returns book information (title, author, synopsis, publication date). The server stores the retrieved data in an internal database (e.g., MySQL or MongoDB).

[0835] Input: API request / response

[0836] Output: Book information stored in the internal database

[0837] Step 2:

[0838] Summary generation

[0839] The server generates a summary using book information retrieved from an internal database. Specifically, it inputs the book title, author, and synopsis into a generative AI model (e.g., GPT-4). Example prompt:

[0840] "Book Title: Introduction to Algorithms for Engineers Author: Taro Yamada Summary: This book introduces the basic concepts and practical applications of algorithms."

[0841] The generative AI model outputs a summary of up to 200 characters, which is then verified by the server to ensure it does not contain inappropriate content.

[0842] Input: Book information, prompt for the generative AI model

[0843] Output: Summary of up to 200 characters

[0844] Step 3:

[0845] Abstract verification

[0846] The server verifies the generated summary. It uses an automatic verification algorithm to check whether the summary contains any prohibited words. It also checks for grammatical errors and inappropriate content. Once the verification is complete, the summary proceeds to the next process.

[0847] Input: Generated summary

[0848] Output: Verified summary

[0849] Step 4:

[0850] Collection of related images

[0851] The server collects related images based on keywords extracted from the abstract. For example, it searches for and retrieves appropriate images from image databases on the Internet or internal image databases based on keywords such as "algorithm" or "application."

[0852] Input: Keywords extracted from the abstract

[0853] Output: Associated image data

[0854] Step 5:

[0855] Combining images and abstracts

[0856] The server combines the collected images with the verified summaries to generate a slideshow-style video. Specifically, it uses video generation software such as FFmpeg to overlay the summaries on the images and create a video file of up to 30 seconds in length.

[0857] Input: Related images, verified summary

[0858] Output: Generated video file

[0859] Step 6:

[0860] Uploading videos

[0861] The server uploads the generated video file to a video distribution service (e.g., YouTube or Vimeo). At this time, information such as the video title, description, and tags are set. Once the upload is complete, the server obtains the video URL from the video distribution service and saves it in an internal database.

[0862] Input: Generated video file, video information (title, description, tags)

[0863] Output: Video URL, saved to internal database

[0864] Step 7:

[0865] User Notifications

[0866] The server sends a message to the user's device to notify them that a new video has been uploaded. The device (smartphone or computer) then displays the message to the user via a push notification. For example, a notification could be sent to the user saying, "A new summary video has been uploaded."

[0867] Input: Video URL, notification message

[0868] Output: Notification displayed on the user's device

[0869] Step 8:

[0870] sentiment analysis

[0871] The server collects user viewing history and reaction data and uses an emotion engine to analyze the user's emotions. For example, it infers the user's emotional state based on data such as ratings, viewing time, and comments on videos viewed in the past.

[0872] Input: Viewing history, reaction data

[0873] Output: User emotion data

[0874] Step 9:

[0875] Content Optimization

[0876] The server selects the most suitable book summary video for the user based on the analysis results of the emotion engine. It prioritizes books in the user's favorite genres and themes, and generates the summary text and video again based on that. This makes it possible to provide personalized content to the user.

[0877] Input: User emotion data, book information

[0878] Output: Optimized book summary video

[0879] (Application example 2)

[0880] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0881] When viewing book summaries and videos, it is difficult to provide optimal content that matches the user's interests. Another issue is that the content presented to the user does not necessarily match the user's preferences. Furthermore, there is a need for a system that can efficiently generate summaries and videos from a vast amount of book information and optimize them based on the user's emotions.

[0882] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0883] In this invention, the server includes means for acquiring book information from a book database or API, a generation AI model means for generating summaries using the acquired book information, means for generating videos based on the generated summaries, means for uploading the generated videos to a video distribution service, means for notifying the user, an emotion engine means for analyzing the user's emotions, and means for optimizing and distributing book summary videos to the user based on the emotion engine means. This makes it possible to provide optimal book summary videos based on the user's emotions and improve the viewing experience.

[0884] A "book database" is a database that accumulates and organizes information about books.

[0885] "API" stands for Application Programming Interface, an interface for exchanging functions and data between different software programs.

[0886] A "generative AI model" is a model that uses artificial intelligence to generate output data from specific input data. In this invention, it refers to a model that generates a summary from book information.

[0887] The "emotion engine" is an engine that analyzes users' emotions and reactions and provides optimal content based on that data.

[0888] A "video distribution service" is a service that distributes video content over the Internet.

[0889] A "summary" is a sentence that summarizes and shortens a long sentence, and in the present invention, it is intended to express the main content of a book concisely.

[0890] "Optimization" means adjusting and improving to best suit a specific purpose or condition. In this invention, it refers to optimizing the book summary video based on user sentiment.

[0891] "Notification" is the act of informing a user of a certain fact or information, and in the present invention, it is intended to notify the user of the distribution of a new digest video.

[0892] This system acquires book information from a book database or API, generates a summary based on the acquired information, and generates a video based on the summary. Furthermore, this system incorporates an emotion engine that recognizes the user's emotions, and provides the user with an optimal book summary video based on the acquired emotion information.

[0893] Get book data

[0894] The server connects to a specific book database or API to retrieve data on new bestselling books, including the book's title, author, synopsis, and publication date. The server runs a scheduled task at a regular time to retrieve the data and store it in an internal database.

[0895] Summary generation

[0896] The server uses a generative AI model to generate a summary of up to 200 characters based on the acquired book information. By inputting the book title, author, and synopsis into the generative AI model, the AI ​​model outputs a summary. This summary is then verified by the server to ensure it does not contain any inappropriate content.

[0897] Video generation

[0898] The server automatically generates or collects relevant images based on the verified summary. Depending on the extracted keywords, appropriate images are collected from the Internet or selected from an internal image database. The server then combines these images with the summary, overlaying text on the images to generate a slideshow-style video. Video generation software (e.g., FFmpeg) is used to generate a video file of up to 30 seconds.

[0899] Video distribution

[0900] The server automatically uploads the generated video file to a video distribution service (for example, YouTube or Vimeo). The server uses the API of the video distribution service to set the information required when uploading the video (title, description, tags, etc.). Once the upload is complete, the server obtains the URL of the generated video from the video distribution service and saves it in its internal database.

[0901] User Notifications

[0902] The server sends a message to the user's device to notify them that a new video has been uploaded. The device (smartphone or computer) then displays a push notification to the user, saying something like, "A new summary video has been released." After checking the notification, the user clicks on the provided video link to access the video streaming service and watch the 30-second summary video.

[0903] Sentiment Analysis and Content Optimization

[0904] The emotion engine analyzes a user's emotions based on their viewing history and reaction data. For example, the emotion engine learns which genres and content of videos a user has particularly rated highly among those they have watched in the past. Based on the emotion data analyzed by the emotion engine, the server selects and prioritizes the book summaries and videos that best match the user's emotions. Specifically, the server prioritizes books in the user's favorite genres and themes, and generates and distributes summaries and videos based on these.

[0905] Specific examples

[0906] Get book data

[0907] For example, suppose a server retrieves data on bestselling books from an API at midnight every day. For example, suppose a book called "Algorithms for Engineers" is retrieved.

[0908] Summary generation

[0909] The server inputs the data from the above book into a generative AI model and generates a summary like the one below.

[0910] Title: Algorithms for Engineers

[0911] Author: Example author

[0912] Summary: This book introduces fundamental concepts and practical applications of algorithms.

[0913] Generated summary:

[0914] This book introduces fundamental concepts and practical applications of algorithms.

[0915] Video generation

[0916] The server collects related images based on keywords such as "algorithm" and "practice," and combines the summary text and images to generate a 30-second video.

[0917] Video distribution

[0918] The server uploads the generated video to a video distribution service (e.g., YouTube), obtains the video URL, and stores it in a database.

[0919] User Notifications

[0920] The server sends a push notification to the user's smartphone, and the user clicks on the notification to watch the video.

[0921] Sentiment Analysis and Content Optimization

[0922] If the emotion engine analyzes the user's viewing history and recognizes that the user prefers technical books, for example, it will prioritize generating and delivering summaries and videos of technical books from the next time onwards.

[0923] This allows users to efficiently watch video summaries of books that match their interests and emotions.

[0924] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0925] Step 1:

[0926] The server connects to a book database or API to retrieve new book information. It receives a JSON response containing data such as the book title, author, synopsis, and publication date. For example, it sends a GET request to an API and receives JSON data containing book information in response. It then stores this data in an internal database.

[0927] Input: API endpoint

[0928] Output: Book information (title, author, synopsis, publication date)

[0929] Step 2:

[0930] The server generates a prompt based on the acquired book information and inputs it into the generative AI model. The prompt includes the book title, author, and synopsis.

[0931] Example prompt sentence:

[0932] Title: Algorithms for Engineers

[0933] Author: Example author

[0934] Summary: This book introduces fundamental concepts and practical applications of algorithms.

[0935] As a result, a summary sentence is generated from the AI ​​model.

[0936] Input: Book information (title, author, synopsis)

[0937] Output: Summary

[0938] Step 3:

[0939] The server then validates the generated abstract to ensure it is correct, including checking for any invalid language or inappropriate content. If the abstract passes validation, it is sent to the next step.

[0940] Input: Abstract

[0941] Output: Verified summary

[0942] Step 4:

[0943] The server collects related images based on the verified summary, picks out important words from the summary and synopsis to extract keywords, and collects images from the Internet and internal image databases based on these keywords.

[0944] Input: Verified abstract, keywords

[0945] Output: A list of related images

[0946] Step 5:

[0947] The server combines the summary text with the collected images to generate a slideshow-style video using FFmpeg or other video generation tools. The summary text is overlaid on the images to create a video file.

[0948] Input: Verified summary, related images

[0949] Output: Video file

[0950] Step 6:

[0951] The server uploads the generated video file to a video distribution service. At this time, metadata such as the video title, description, and tags are set. For example, the YouTube API is used to upload the video and obtain its URL.

[0952] Input: Video file, metadata

[0953] Output: Video URL

[0954] Step 7:

[0955] The server saves the URL of the new video in an internal database and sends that information to the user's device via a push notification service (e.g., Firebase Cloud Messaging), informing the user that a new summary video has been released.

[0956] Input: Video URL

[0957] Output: Push notification message

[0958] Step 8:

[0959] The server collects users' viewing history and reaction data, and analyzes their emotions using an emotion engine, which determines what genres and content are most suitable for the user.

[0960] Input: User viewing history, reaction data

[0961] Output: Sentiment analysis data

[0962] Step 9:

[0963] Based on the data analyzed by the emotion engine, the server selects the next best book summary video and starts the process of generating and delivering that video, ensuring that content that matches the user's emotions is provided first.

[0964] Input: Sentiment analysis data

[0965] Output: Recommended book information, summary, video URL

[0966] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0967] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0968] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0969] [Third embodiment]

[0970] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0971] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[0972] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0973] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0974] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0975] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0976] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0977] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0978] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0979] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0980] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0981] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0982] A natural language description of the program's operation

[0983] This system provides an automated process to summarize international best-selling books in 200 characters or less and distribute them as 30-second videos. The following explains the roles and processes of the server, terminal, and user in detail.

[0984] Get book data

[0985] The server retrieves the book data.

[0986] First, the server connects to a specific book database or API (Application Program Interface) to retrieve new book information. The server runs a scheduled task at a regular time, for example, retrieving the latest best-selling book list at midnight every day. The retrieved data includes the book title, author, synopsis, publication date, etc.

[0987] Summary generation

[0988] The server generates a summary

[0989] Next, the server generates a summary using a generative AI model (such as GPT-3) based on the acquired book information. The server inputs the book title, author, and synopsis into the generative AI model in a prompt format, and the AI ​​model outputs a summary of up to 200 characters. The generated summary is then verified by the server to ensure it does not contain any inappropriate content.

[0990] Video generation

[0991] The server generates a video from the summary text

[0992] Based on the verified summary, the server automatically generates or collects related images. The server extracts keywords from the summary and, depending on the keywords, collects related images from the Internet or selects them from an internal image database. The server then combines these images with the summary, overlaying text on the images to generate a slideshow-style video. Video generation software such as FFmpeg is used to generate video files of up to 30 seconds.

[0993] Video distribution

[0994] The server uploads the video to the video streaming service.

[0995] The server automatically uploads the generated video file to a video distribution service (such as YouTube or Vimeo). The server uses the video distribution service's API to set the information required when uploading the video (title, description, tags, etc.). Once the upload is complete, the video's URL is obtained from the service and saved in an internal database.

[0996] User Notifications

[0997] The device notifies the user of the video link

[0998] The server sends a message to the user's device to notify them that a new video has been uploaded. The device (smartphone or computer) then displays a push notification to the user with a message such as "A new summary video has been released." After confirming the notification, the user clicks on the provided video link to access the video streaming service and watch the 30-second summary video.

[0999] Specific examples

[1000] 1. Acquiring book data

[1001] The server retrieves the bestselling book data from the API at midnight every day. For example, it retrieves the data for the book "The Silent Patient."

[1002] 2. Summary Generation

[1003] The server inputs the data from "The Silent Patient" into a generative AI model and generates the summary: "Alicia is an artist who goes silent after a mysterious incident, and the story is about a therapist who searches for the truth behind it."

[1004] 3. Video Generation

[1005] The server collects free images based on keywords such as "Alicia" and "therapist," and combines the summary text and images to generate a 30-second video.

[1006] 4. Video distribution

[1007] The server uploads the generated video to YouTube, retrieves the video URL and stores it in the database.

[1008] 5. User Notices

[1009] The server sends a push notification to the user's smartphone, and the user clicks on the notification to watch the video.

[1010] This system allows users to easily view summaries of international best-selling books in a short amount of time, enabling them to acquire information efficiently.

[1011] The processing flow will be explained below.

[1012] Step 1:

[1013] The server connects to the book database or API at midnight every day to retrieve the latest bestselling book data, including the book title, author, synopsis, and publication date.

[1014] Step 2:

[1015] The server stores the acquired book information in an internal database and adds it to a processing queue, so that subsequent processes can retrieve and process the book information sequentially.

[1016] Step 3:

[1017] The server retrieves a book from the processing queue and inputs the book's title, author, and synopsis into the generative AI model, providing appropriate input to the AI ​​model in the form of prompts.

[1018] Step 4:

[1019] The generative AI model outputs a summary of up to 200 characters based on the input information. The server receives the summary and prepares it for the next process.

[1020] Step 5:

[1021] The server then validates the generated summary to check for inappropriate content, for example, by checking for inappropriate language or unclear sentences.

[1022] Step 6:

[1023] The server extracts keywords from the verified abstract and passes the list of keywords, including people's names, places, and events, to the image generation module.

[1024] Step 7:

[1025] The image generation module collects free stock images corresponding to each keyword from the Internet or selects them from an internal image database, and the server downloads or retrieves the images.

[1026] Step 8:

[1027] The server combines the abstract with the acquired images to generate a slideshow-style video by overlaying the text on the images. Video generation software (e.g., FFmpeg) is used to generate a video file of up to 30 seconds.

[1028] Step 9:

[1029] The server uploads the generated video file to a video distribution service (e.g., YouTube or Vimeo), and also sets the video's metadata (title, description, tags, etc.) at the same time.

[1030] Step 10:

[1031] Once the upload is complete, the server retrieves the generated video URL from the video streaming service, which is then stored in an internal database.

[1032] Step 11:

[1033] The server sends a message to the user's device informing them that a new video has been uploaded, which includes sending a push notification.

[1034] Step 12:

[1035] The device displays a notification to the user saying, "A new summary video has been released." The user checks the notification and clicks on the provided video link.

[1036] Step 13:

[1037] Users can easily obtain summaries of international best-selling books by watching 30-second summary videos on their device's browser or YouTube app.

[1038] This detailed processing flow automates and efficiently executes the entire process of obtaining book information, generating a summary and video, and finally notifying the user.

[1039] Example 1

[1040] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1041] Current book information systems do not allow users to quickly obtain summaries of international best-selling books, making it difficult for users to grasp the content in a short time. Furthermore, when summaries are provided only in text, the lack of visual elements makes it difficult for users to understand or engage with the information. Therefore, there is a need for a method to provide book information efficiently and in a visually understandable way.

[1042] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1043] In this invention, the server includes means for acquiring book information from a book database or API, means for creating a prompt based on the acquired book information and sending it to the generative AI model, means for verifying the summary returned from the generative AI model, means for extracting keywords from the summary and collecting related images, means for generating a slideshow video based on the collected images and summary, means for uploading the generated video to a video distribution service, and means for sending a notification message including a link to the video to the user terminal. This allows users to quickly and visually understand summaries of international best-selling books.

[1044] A "book database" is a system that systematically stores information about books and has a data structure that allows for searching and extraction.

[1045] An "API" is an interface that allows various pieces of software to communicate with each other and share functions, and is often provided in the form of a web service.

[1046] A "prompt sentence" is an input sentence provided to a generative AI model, a series of text that is generated in anticipation of specific conditions or responses.

[1047] A "generative AI model" is an artificial intelligence algorithm that automatically generates text and information from large amounts of data.

[1048] "Validation" is the process of verifying whether the generated data is appropriate and checking that it does not contain any inappropriate content.

[1049] "Keywords" are important words extracted from summaries and texts, and are used for information retrieval and classification.

[1050] "Image gathering" is the process of acquiring relevant image data based on specific criteria or keywords.

[1051] "Slideshow format" is a visual presentation format that displays multiple images and text sequentially.

[1052] A "video distribution service" is an online platform for uploading videos over the Internet and making them available to viewers.

[1053] A "notification message" is a message sent by the system to convey specific information to the user.

[1054] A "user terminal" is a device such as a smartphone or computer that allows a user to access the Internet and applications.

[1055] To implement this invention, a system is required that automates the process of acquiring book information from a book database or API, generating a summary using a generative AI model based on that information, creating a video using that summary, and finally notifying the user. The following describes in detail the hardware and software used in each processing step, as well as the details of data processing and data calculation.

[1056] Get book data

[1057] Server Roles

[1058] The server connects to a specific book database or API to retrieve new book information. For example, it can use the Goodreads or Google Books API. The server runs a periodic task every day at midnight, sending an HTTP request to retrieve the latest best-selling book list. The retrieved data is provided in JSON format and includes the book title, author, synopsis, publication date, etc.

[1059] Summary generation

[1060] Server Roles

[1061] The server creates a prompt based on the acquired book information and sends it to the generative AI model. Specifically, it combines the book title, author, and synopsis to create a prompt to be input to the generative AI model (e.g., GPT-3). The following is an example of a prompt:

[1062] Example prompt sentence:

[1063] "Book title: 'XXX', author: XXX, summary: XXX. Please write a summary in 200 characters or less."

[1064] The server receives the summary returned by the generative AI model and uses an internal validation algorithm to check for inappropriate content. If there are no problems, the summary is stored in the database.

[1065] Video generation

[1066] Server Roles

[1067] The server extracts keywords from the summary and collects related images. It uses a natural language processing algorithm to extract key keywords from the summary and collects images via HTTP requests from free image services (such as Unsplash and Pexels). It then combines the collected images with the summary to generate a slideshow-style video using video generation software such as FFmpeg. The video must be no longer than 30 seconds.

[1068] Video distribution

[1069] Server Roles

[1070] The server uploads the generated video to a video distribution service (such as YouTube or Vimeo). The server uses the API of the video distribution service to set the necessary information such as title, description, and tags when uploading the video. If the upload is successful, the server obtains the URL of the video returned by the video service and saves it in the database.

[1071] User Notifications

[1072] Device Role

[1073] The server generates a notification message containing the URL of the new video and sends a push notification to the user's device. The device (smartphone or computer) then displays a notification to the user saying, "A new summary video has been released." The user can click the notification to access the video streaming service and watch a 30-second summary video.

[1074] Specific examples

[1075] For example, for the book "Silent Patient," the server sends an HTTP request to the Goodreads API at midnight every day to retrieve information about the book. Next, the following prompt is sent to the generative AI model: "Book Title: 'Silent Patient,' Author: Alex Michaelides, Synopsis: Alicia is an artist who falls silent after a mysterious incident, and the story of a therapist who searches for the truth behind this. Please provide a summary of 200 characters or less." Based on the summary returned by the generative AI model, "Alicia is an artist who falls silent after a mysterious incident, and the story of a therapist who searches for the truth behind this," the server extracts keywords such as "Alicia" and "therapist" and collects related free images using the Unsplash API. These images and the summary are then processed into a slideshow video using FFmpeg. The server then uploads this video through the YouTube API and gives it a title such as "Silent Patient Summary Video." The user's device receives a notification stating, "A new summary video of 'Silent Patient' has been uploaded," and they can click the notification to watch the video.

[1076] This system allows users to quickly obtain summaries of international best-selling books in a visually easy-to-understand manner.

[1077] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1078] Step 1:

[1079] Get book data

[1080] The server connects to the book database or API to retrieve new book information. The server sends an HTTP request at midnight every day to retrieve the latest bestselling book list. The retrieved data is provided in JSON format and includes the book title, author, synopsis, publication date, etc. The input is the book information retrieved from the book database or API, and the output is the parsed book information. The server parses the retrieved data and saves it in its internal database.

[1081] Step 2:

[1082] Creating and sending prompts

[1083] The server creates a prompt based on the acquired book information and sends it to the generative AI model. Specifically, it combines the book title, author, and synopsis to create a prompt to be input to the generative AI model (e.g., GPT-3). The input is the analyzed book information, and the output is the generated prompt. Below is an example of a prompt:

[1084] Example prompt sentence:

[1085] "Book title: 'XXX', author: XXX, summary: XXX. Please write a summary in 200 characters or less."

[1086] Step 3:

[1087] Summary generation and verification

[1088] The server receives the summary returned by the generative AI model. The server checks the generated summary with an internal verification algorithm to ensure that it does not contain inappropriate content. The input is the summary returned by the generative AI model, and the output is the summary after verification. If it is confirmed that it does not contain inappropriate content, the summary is saved in the database.

[1089] Step 4:

[1090] Keyword extraction and image collection

[1091] The server extracts keywords from the abstract and collects related images. It uses a natural language processing algorithm to extract key keywords from the abstract and collects images via HTTP requests from services that provide free images (e.g., Unsplash and Pexels). The input is the verified abstract, and the output is the extracted keywords and collected images.

[1092] Step 5:

[1093] Video generation

[1094] The server generates a slideshow-style video based on the collected images and summaries. The collected images are concatenated and the summaries are overlaid on the images as text. Video generation software such as FFmpeg is used to generate videos of up to 30 seconds. The input is the collected images and summaries, and the output is the generated video file.

[1095] Step 6:

[1096] Uploading videos

[1097] The server uploads the generated video to a video distribution service. Using the API of the video distribution service (for example, YouTube or Vimeo), necessary information such as title, description, and tags is set when uploading the video. The input is the generated video file, and the output is the URL of the uploaded video. If the upload is successful, the server saves the video URL returned by the video service in a database.

[1098] Step 7:

[1099] User Notifications

[1100] The server generates a notification message containing the URL of the new video and sends a push notification to the user's device. The device (smartphone or computer) receives the notification and displays a message to the user saying, "A new summary video has been released." The input is the URL of the uploaded video, and the output is the notification message displayed on the user's device. The user can click the notification to access the video streaming service and watch a 30-second summary video.

[1101] (Application example 1)

[1102] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1103] In today's world, busy people want to obtain information efficiently within a limited amount of time, but it is often difficult to find the time to read long books. Another issue is that simply reading a book summary can weaken comprehension of the content and the impact of the information. Furthermore, there are limited ways for users to easily view summarized book information, and this needs to be addressed.

[1104] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1105] In this invention, the server includes a means for acquiring book information from a book database or an API, a generation AI model means for generating a summary using the acquired book information, and a means for generating a video based on the generated summary, thereby enabling users to efficiently view summarized book information in video format.

[1106] "Book database" refers to one or more data collections that store information about books.

[1107] "API" refers to an interface that allows application programs to communicate with each other.

[1108] A "generative AI model" refers to artificial intelligence that automatically generates text, images, videos, etc. based on specified input data.

[1109] A "summary" is a short summary of the book's contents.

[1110] "Video distribution service" refers to a service that provides video content to users via the Internet.

[1111] "User" refers to an entity that uses the system to receive information or services.

[1112] "Smart devices" refer to mobile information devices with internet connectivity, such as smartphones and tablets.

[1113] MODE FOR CARRYING OUT THE INVENTION

[1114] System Program

[1115] This system retrieves book information from a book database or API, and generates a summary based on that information using a generative AI model. It consists of a series of processes: generating a video based on the generated summary, uploading the video to a video distribution service, and sending a notification to the user. Users can also watch the summary video on their smart devices. The program design for realizing this system is as follows:

[1116] Hardware and software used

[1117] Server hardware: Amazon EC2

[1118] Book data acquisition: Google Books API

[1119] Summary sentence generation: OpenAI GPT-3

[1120] Video generation: FFmpeg

[1121] Video streaming: YouTube Data API

[1122] Push notifications: Firebase Cloud Messaging (FCM)

[1123] Smart devices: smartphones, tablets, etc.

[1124] Data processing and calculation

[1125] The details of how this system works are as follows:

[1126] 1. Acquiring book data

[1127] The server runs a scheduled task at regular intervals, connecting to a book database or API (e.g., Google Books API) to retrieve the latest best-selling book list, including book title, author, synopsis, publication date, etc.

[1128] 2. Summary Generation

[1129] Based on the acquired book information, the server uses a generative AI model (e.g., OpenAI GPT-3) to generate a summary. Specifically, the book title, author, and synopsis are input into the generative AI model in the form of a prompt, and the AI ​​model outputs a summary of up to 200 characters. This summary is then verified by the server to ensure it does not contain any inappropriate content.

[1130] 3. Video Generation

[1131] Based on the verified summary, the server automatically generates or collects related images. The server extracts keywords from the summary and, based on those keywords, collects related images from the Internet or selects them from an internal image database. It then combines these images with the summary, overlaying text on the images to generate a slideshow-style video. It uses video generation software such as FFmpeg to generate video files of up to 30 seconds.

[1132] 4. Video distribution

[1133] The server automatically uploads the generated video file using the API of a video distribution service (such as YouTube or Vimeo). The necessary information (title, description, tags, etc.) is also set when uploading. Once the upload is complete, the video URL is obtained from the service and saved in an internal database.

[1134] 5. User Notices

[1135] The server sends a message to the user's smart device to notify them that a new video has been uploaded. A push notification displays a message to the user, such as "A new summary video has been uploaded." The user checks the notification and clicks on the provided video link to access the video streaming service and watch the 30-second summary video.

[1136] Specific examples

[1137] 1. Acquiring book data

[1138] Book: "The Silent Patient"

[1139] Author: "Alex Michaelides"

[1140] Synopsis: "Alicia Belson is a celebrated artist..."

[1141] 2. Summary Generation

[1142] Prompt statement:

[1143] Title: The Silent Patient

[1144] Author: Alex Michaelides

[1145] Summary: Alicia Belson is a renowned artist whose work is celebrated around the world. One day, she shoots and kills her husband in their home and never speaks again. This story is told from the perspective of a therapist who tries to unravel this mystery.

[1146] Please summarize in 200 characters or less.

[1147] Generated summary:

[1148] Alicia is an artist who goes silent after a mysterious incident, and the story revolves around a therapist who searches for the truth.

[1149] 3. Video Generation

[1150] The generated summary is combined with images related to keywords such as "Alicia" and "therapist," and a video of less than 30 seconds is generated using FFmpeg.

[1151] 4. Video distribution

[1152] Upload a video to YouTube, get the URL and save it in the database.

[1153] 5. User Notices

[1154] Use Firebase Cloud Messaging to send notifications of new video streams to users' smart devices.

[1155] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1156] Step 1:

[1157] The server periodically connects to a book database or API to retrieve the latest best-selling book information, including book title, author, synopsis, and publication date. This allows you to always have access to the latest book information.

[1158] Input: Scheduled task, book database or API

[1159] Output: Retrieved book information (title, author, synopsis, publication date)

[1160] Step 2:

[1161] The server generates a summary using a generative AI model (e.g., OpenAI GPT-3) based on the acquired book information. The book title, author, and synopsis are input into the generative AI model in the form of a prompt, and the AI ​​outputs a summary of up to 200 characters. At this time, the summary is verified to ensure that it does not contain any inappropriate content.

[1162] Input: Book information (title, author, synopsis), generative AI model

[1163] Output: Summary of up to 200 characters

[1164] Step 3:

[1165] The server extracts keywords from the verified summary and uses them to gather relevant images from the internet or select them from an internal image database. It then combines the summary with these images and generates a slideshow-style video by overlaying text on the images. It then uses video generation software such as FFmpeg to generate a video file of up to 30 seconds.

[1166] Input: Abstract, keywords, related images

[1167] Output: Video file up to 30 seconds long

[1168] Step 4:

[1169] The server uploads the generated video via the API of a video distribution service (e.g., YouTube). Information such as the title, description, and tags are also set. Once the upload is complete, the video URL is obtained from the service and saved in an internal database.

[1170] Input: Video file, video streaming service API

[1171] Output: Video URL

[1172] Step 5:

[1173] The server sends a push notification message via Firebase Cloud Messaging to the user's smart device to notify them that a new video has been uploaded, and the user receives the notification and clicks the video link to watch the video.

[1174] Input: Video URL, Firebase Cloud Messaging

[1175] Output: User notification

[1176] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1177] A natural language description of the program's operation

[1178] This system acquires book information from a book database or API, generates a summary based on the acquired information, and generates a video based on the summary. Furthermore, this system incorporates an emotion engine that recognizes the user's emotions, and provides the user with an optimal book summary video based on the acquired emotion information.

[1179] Get book data

[1180] The server retrieves the book data.

[1181] The server connects to a specific book database or API to retrieve data on new bestselling books, including the book's title, author, synopsis, and publication date. The server runs a scheduled task at a regular time to retrieve the data and store it in an internal database.

[1182] Summary generation

[1183] The server generates a summary

[1184] The server uses a generative AI model to generate a summary of up to 200 characters based on the acquired book information. By inputting the book title, author, and synopsis into the generative AI model, the AI ​​model outputs a summary. This summary is then verified by the server to ensure it does not contain any inappropriate content.

[1185] Video generation

[1186] The server generates a video from the summary text

[1187] Based on the verified summary, the server automatically generates or collects relevant images. Depending on the extracted keywords, appropriate images are collected from the Internet or selected from an internal image database. The server then combines these images with the summary, overlaying text on the images to generate a slideshow-style video. Video generation software (e.g., FFmpeg) is used to generate a video file of up to 30 seconds.

[1188] Video distribution

[1189] The server uploads the video to the video streaming service.

[1190] The server automatically uploads the generated video file to a video distribution service (for example, YouTube or Vimeo). The server uses the API of the video distribution service to set the information required when uploading the video (title, description, tags, etc.). Once the upload is complete, the server obtains the URL of the generated video from the video distribution service and saves it in its internal database.

[1191] User Notifications

[1192] The device notifies the user of the video link

[1193] The server sends a message to the user's device to notify them that a new video has been uploaded. The device (smartphone or computer) then displays a push notification to the user, saying something like, "A new summary video has been released." After checking the notification, the user clicks on the provided video link to access the video streaming service and watch the 30-second summary video.

[1194] Sentiment Analysis and Content Optimization

[1195] The server analyzes the user's emotions

[1196] The emotion engine analyzes user emotions based on their viewing history and reaction data. For example, it learns the genres and content that users have particularly rated among the videos they have watched in the past.

[1197] The server delivers the best content

[1198] Based on the emotional data analyzed by the emotion engine, the server selects and prioritizes the book summaries and videos that best fit the user's emotions. Specifically, it prioritizes books in the user's favorite genres and themes, and generates and distributes summaries and videos based on those.

[1199] Specific examples

[1200] 1. Acquiring book data

[1201] The server retrieves data on bestselling books from the API at midnight every day. For example, it retrieves data on the book "Introduction to Algorithms for Engineers."

[1202] 2. Summary Generation

[1203] The server inputs data from "Introduction to Algorithms for Engineers" into a generative AI model and generates a summary statement that reads, "This book introduces the basic concepts and practical applications of algorithms."

[1204] 3. Video Generation

[1205] The server collects related images based on keywords such as "algorithm" and "application," and combines the summary text and images to generate a 30-second video.

[1206] 4. Video distribution

[1207] The server uploads the generated video to YouTube, retrieves the video URL and stores it in the database.

[1208] 5. User Notices

[1209] The server sends a push notification to the user's smartphone, and the user clicks on the notification to watch the video.

[1210] 6. Sentiment Analysis and Content Optimization

[1211] The emotion engine analyzes the user's viewing history, and if, for example, the user often watches technical books, from the next time onwards, it will prioritize summarizing technical books and generating and delivering videos.

[1212] This detailed process allows users to efficiently watch video summaries of books that suit their preferences and emotions.

[1213] The processing flow will be explained below.

[1214] Step 1:

[1215] The server connects to the book database or API at midnight every day to retrieve the latest bestselling book data, including the book title, author, synopsis, and publication date.

[1216] Step 2:

[1217] The server stores the acquired book information in an internal database and adds it to a processing queue, so that subsequent processes can retrieve and process the book information sequentially.

[1218] Step 3:

[1219] The server retrieves a book from the processing queue and inputs the book's title, author, and synopsis into the generative AI model, providing appropriate input to the AI ​​model in the form of prompts.

[1220] Step 4:

[1221] The generative AI model outputs a summary of up to 200 characters based on the input information. The server receives the summary and prepares it for the next process.

[1222] Step 5:

[1223] The server then validates the generated summary to check for inappropriate content, for example, by checking for inappropriate language or unclear sentences.

[1224] Step 6:

[1225] The server extracts keywords from the verified abstract and passes the list of keywords, including people's names, places, and events, to the image generation module.

[1226] Step 7:

[1227] The image generation module collects free stock images corresponding to each keyword from the Internet or selects them from an internal image database, and the server downloads or retrieves the images.

[1228] Step 8:

[1229] The server combines the abstract with the acquired images to generate a slideshow-style video by overlaying the text on the images. Video generation software (e.g., FFmpeg) is used to generate a video file of up to 30 seconds.

[1230] Step 9:

[1231] The server uploads the generated video file to a video distribution service (e.g., YouTube or Vimeo), and also sets the video's metadata (title, description, tags, etc.) at the same time.

[1232] Step 10:

[1233] Once the upload is complete, the server retrieves the generated video URL from the video streaming service, which is then stored in an internal database.

[1234] Step 11:

[1235] The server sends a message to the user's device informing them that a new video has been uploaded, which includes sending a push notification.

[1236] Step 12:

[1237] The device displays a notification to the user saying, "A new summary video has been released." The user checks the notification and clicks on the provided video link.

[1238] Step 13:

[1239] Users can easily obtain summaries of international best-selling books by watching 30-second summary videos on their device's browser or YouTube app.

[1240] Sentiment Analysis and Content Optimization

[1241] Step 14:

[1242] The server collects the user's viewing history and reaction data and sends it to the emotion engine, which analyzes the user's viewing history and reaction data to estimate the user's emotions.

[1243] Step 15:

[1244] Based on the emotional data analyzed by the emotion engine, the server selects the most suitable book summary and video for the user. Specifically, it prioritizes books in genres and themes that the user has given high ratings to, and generates summaries and videos based on those.

[1245] Step 16:

[1246] The server uploads the generated optimal book summary video to a video distribution service and obtains a distribution link.

[1247] Step 17:

[1248] The server sends push notifications to the user's device at the appropriate time, allowing the user to receive a summary video that matches their emotions.

[1249] Step 18:

[1250] The device receives a notification and displays a message to the user saying, "A summary video has been delivered that is recommended for you." The user clicks the link to watch the video.

[1251] This allows the emotional engine to be used to provide summarized videos based on the user's emotions, resulting in even higher satisfaction. To give a concrete example, for example, a user who frequently reads technical books could be given preferential access to videos summarizing new best-selling books in the same genre. In this way, more personalized content is provided to users through sentiment analysis and content optimization.

[1252] Example 2

[1253] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1254] Conventional book recommendation systems have the problem that the process of generating a book summary and distributing it as a video incorporating visuals is complicated and difficult to automate. It is also difficult to provide personalized content that takes into account the user's preferences and emotions. This makes it difficult to attract the user's interest, resulting in a poor content consumption experience.

[1255] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for acquiring book information from a book database or API, a generation AI model means for generating a summary using the acquired book information, means for verifying the generated summary, means for automatically generating or collecting related images based on the verified summary, means for generating a video by combining the generated images and summary, means for uploading the generated video to a video distribution service, means for acquiring and saving the URL of the uploaded video, means for notifying the user, means for analyzing the user's emotions, and means for providing an optimal book summary video based on the analysis results. This automates the process from generating book summaries to distributing videos and providing content optimized for users, enabling the provision of an efficient and high-quality user experience.

[1256] "Book information" refers to basic data such as the book's title, author, synopsis, and publication date.

[1257] A "generative AI model" refers to an algorithm or software that automatically generates a summary from input data using natural language processing technology.

[1258] "Video Streaming Service" means a platform for hosting generated video content online and for users to stream or download it.

[1259] A "means" refers to a specific method, device, or technique for performing a particular function.

[1260] "Means for analyzing emotions" refers to algorithms and software for estimating and analyzing a user's emotional state based on their viewing history and reaction data.

[1261] A "summary" is a text that compresses the original book information into 200 characters or less and succinctly presents the main content of the book.

[1262] "Verification measures" refers to processes or algorithms that check and evaluate the accuracy and appropriateness of generated summaries and other data.

[1263] "Means of notification" refers to the technology and process used to communicate new video distribution information, etc. to users' devices.

[1264] "Optimal book summary video" refers to the book summary video that is determined to be most relevant based on the user's interests and emotional state.

[1265] "Means for automatically generating or collecting relevant images" refers to algorithms that generate relevant images from text, or technologies that collect appropriate images from existing databases or the internet.

[1266] This invention relates to a system that acquires book information from a book database or API, generates a summary using that information, and then generates a video based on the summary. The system includes an emotion engine that recognizes a user's emotions and provides optimal content based on those emotions. The following describes in detail the embodiments of the invention.

[1267] Get book data

[1268] The server retrieves the book data.

[1269] The server connects to a book database or API (for example, a common book database API) and periodically retrieves data on new best-selling books. This data includes the book's title, author, synopsis, and publication date. The server runs a scheduled task and stores the retrieved data in an internal database. Specifically, the server retrieves the data from the API at midnight every day and stores it in an internal database (for example, MySQL or MongoDB).

[1270] Summary generation

[1271] The server generates a summary

[1272] The server uses a generative AI model (e.g., GPT-4) to generate a summary of up to 200 characters based on the acquired book information. The input includes the book title, author, and synopsis. For example, data from the book "Introduction to Algorithms for Engineers" is input into the generative AI model, and the resulting summary is, "This book introduces the basic concepts and practical applications of algorithms."

[1273] Video generation

[1274] The server generates a video from the summary text

[1275] The server automatically generates or collects related images based on the verified summary. Depending on the extracted keywords, appropriate images are collected from the Internet or selected from an internal image database. For example, related images are collected based on keywords such as "algorithm" and "application." The server then combines the images with the summary, overlaying text on the images to generate a slideshow-style video. Specifically, video generation software such as FFmpeg is used to generate video files of up to 30 seconds.

[1276] Video distribution

[1277] The server uploads the video to the video streaming service.

[1278] The server automatically uploads the generated video file using the API of a video distribution service (for example, YouTube or Vimeo). When uploading, necessary information (title, description, tags, etc.) is set, and once the upload is complete, the server obtains the URL of the generated video from the video distribution service and saves it in an internal database. For example, the video title could be "Introduction to Algorithms for Engineers - Summary Video" and the description "Introducing the basic concepts of algorithms in 30 seconds."

[1279] User Notifications

[1280] The device notifies the user of the video link

[1281] The server sends a message to the user's device to notify them that a new video has been uploaded. The device (smartphone or computer) then displays a push notification to the user, saying something like, "A new summary video has been released." After checking the notification, the user clicks on the provided video link to access the video streaming service and watch the 30-second summary video.

[1282] Sentiment Analysis and Content Optimization

[1283] The server analyzes the user's emotions

[1284] The emotion engine analyzes user emotions based on their viewing history and reaction data, for example, learning the genres and content of videos they have particularly rated.

[1285] The server delivers the best content

[1286] Based on the analysis results, the server selects and delivers the book summaries and videos that best suit the user's emotions. Specifically, it prioritizes books in the user's favorite genres and themes, and generates and delivers summaries and videos based on that. For example, if a user frequently reads technical books, the server will continue to generate summaries and videos for technical books from the next time.

[1287] Specific examples

[1288] Specific examples of book data acquisition

[1289] The server retrieves data for "Introduction to Algorithms for Engineers" from the book database API at midnight every day. This data includes the book title, author, synopsis, and publication date.

[1290] A concrete example of summary generation

[1291] The server inputs data from "Introduction to Algorithms for Engineers" into a generative AI model and generates a summary statement that reads, "This book introduces the basic concepts and practical applications of algorithms."

[1292] Example of video generation

[1293] The server collects relevant images based on keywords such as "algorithm" and "application," and uses FFmpeg to generate a 30-second video that combines the summary text and images.

[1294] Specific examples of video distribution

[1295] The server uploads the generated video to YouTube, obtains the video URL "https: / / www.youtube.com / watch?v=example", and stores it in the database.

[1296] Example of user notification

[1297] The server sends a push notification to the user's smartphone saying, "A new summary video has been released." The user clicks the notification to watch the video.

[1298] Examples of sentiment analysis and content optimization

[1299] The emotion engine analyzes the user's viewing history, and if they frequently watch technical books, it will prioritize similar genres for summarization and generate videos from them next time.

[1300] Prompt Sentence Examples

[1301] "Book Title: Introduction to Algorithms for Engineers Author: Taro Yamada Summary: This book introduces the basic concepts and practical applications of algorithms."

[1302] This invention allows users to efficiently watch video summaries of books that match their emotions and interests.

[1303] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1304] Step 1:

[1305] Get book data

[1306] The server connects to a book database or API and periodically retrieves data on new bestselling books. First, the server sends a request to a specific API (e.g., a general book database API) at midnight every day. In response to this request, the API returns book information (title, author, synopsis, publication date). The server stores the retrieved data in an internal database (e.g., MySQL or MongoDB).

[1307] Input: API request / response

[1308] Output: Book information stored in the internal database

[1309] Step 2:

[1310] Summary generation

[1311] The server generates a summary using book information retrieved from an internal database. Specifically, it inputs the book title, author, and synopsis into a generative AI model (e.g., GPT-4). Example prompt:

[1312] "Book Title: Introduction to Algorithms for Engineers Author: Taro Yamada Summary: This book introduces the basic concepts and practical applications of algorithms."

[1313] The generative AI model outputs a summary of up to 200 characters, which is then verified by the server to ensure it does not contain inappropriate content.

[1314] Input: Book information, prompt for the generative AI model

[1315] Output: Summary of up to 200 characters

[1316] Step 3:

[1317] Abstract verification

[1318] The server verifies the generated summary. It uses an automatic verification algorithm to check whether the summary contains any prohibited words. It also checks for grammatical errors and inappropriate content. Once the verification is complete, the summary proceeds to the next process.

[1319] Input: Generated summary

[1320] Output: Verified summary

[1321] Step 4:

[1322] Collection of related images

[1323] The server collects related images based on keywords extracted from the abstract. For example, it searches for and retrieves appropriate images from image databases on the Internet or internal image databases based on keywords such as "algorithm" or "application."

[1324] Input: Keywords extracted from the abstract

[1325] Output: Associated image data

[1326] Step 5:

[1327] Combining images and abstracts

[1328] The server combines the collected images with the verified summaries to generate a slideshow-style video. Specifically, it uses video generation software such as FFmpeg to overlay the summaries on the images and create a video file of up to 30 seconds in length.

[1329] Input: Related images, verified summary

[1330] Output: Generated video file

[1331] Step 6:

[1332] Uploading videos

[1333] The server uploads the generated video file to a video distribution service (e.g., YouTube or Vimeo). At this time, information such as the video title, description, and tags are set. Once the upload is complete, the server obtains the video URL from the video distribution service and saves it in an internal database.

[1334] Input: Generated video file, video information (title, description, tags)

[1335] Output: Video URL, saved to internal database

[1336] Step 7:

[1337] User Notifications

[1338] The server sends a message to the user's device to notify them that a new video has been uploaded. The device (smartphone or computer) then displays the message to the user via a push notification. For example, a notification could be sent to the user saying, "A new summary video has been uploaded."

[1339] Input: Video URL, notification message

[1340] Output: Notification displayed on the user's device

[1341] Step 8:

[1342] sentiment analysis

[1343] The server collects user viewing history and reaction data and uses an emotion engine to analyze the user's emotions. For example, it infers the user's emotional state based on data such as ratings, viewing time, and comments on videos viewed in the past.

[1344] Input: Viewing history, reaction data

[1345] Output: User emotion data

[1346] Step 9:

[1347] Content Optimization

[1348] The server selects the most suitable book summary video for the user based on the analysis results of the emotion engine. It prioritizes books in the user's favorite genres and themes, and generates the summary text and video again based on that. This makes it possible to provide personalized content to the user.

[1349] Input: User emotion data, book information

[1350] Output: Optimized book summary video

[1351] (Application example 2)

[1352] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1353] When viewing book summaries and videos, it is difficult to provide optimal content that matches the user's interests. Another issue is that the content presented to the user does not necessarily match the user's preferences. Furthermore, there is a need for a system that can efficiently generate summaries and videos from a vast amount of book information and optimize them based on the user's emotions.

[1354] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1355] In this invention, the server includes means for acquiring book information from a book database or API, a generation AI model means for generating summaries using the acquired book information, means for generating videos based on the generated summaries, means for uploading the generated videos to a video distribution service, means for notifying the user, an emotion engine means for analyzing the user's emotions, and means for optimizing and distributing book summary videos to the user based on the emotion engine means. This makes it possible to provide optimal book summary videos based on the user's emotions and improve the viewing experience.

[1356] A "book database" is a database that accumulates and organizes information about books.

[1357] "API" stands for Application Programming Interface, an interface for exchanging functions and data between different software programs.

[1358] A "generative AI model" is a model that uses artificial intelligence to generate output data from specific input data. In this invention, it refers to a model that generates a summary from book information.

[1359] The "emotion engine" is an engine that analyzes users' emotions and reactions and provides optimal content based on that data.

[1360] A "video distribution service" is a service that distributes video content over the Internet.

[1361] A "summary" is a sentence that summarizes and shortens a long sentence, and in the present invention, it is intended to express the main content of a book concisely.

[1362] "Optimization" means adjusting and improving to best suit a specific purpose or condition. In this invention, it refers to optimizing the book summary video based on user sentiment.

[1363] "Notification" is the act of informing a user of a certain fact or information, and in the present invention, it is intended to notify the user of the distribution of a new digest video.

[1364] This system acquires book information from a book database or API, generates a summary based on the acquired information, and generates a video based on the summary. Furthermore, this system incorporates an emotion engine that recognizes the user's emotions, and provides the user with an optimal book summary video based on the acquired emotion information.

[1365] Get book data

[1366] The server connects to a specific book database or API to retrieve data on new bestselling books, including the book's title, author, synopsis, and publication date. The server runs a scheduled task at a regular time to retrieve the data and store it in an internal database.

[1367] Summary generation

[1368] The server uses a generative AI model to generate a summary of up to 200 characters based on the acquired book information. By inputting the book title, author, and synopsis into the generative AI model, the AI ​​model outputs a summary. This summary is then verified by the server to ensure it does not contain any inappropriate content.

[1369] Video generation

[1370] The server automatically generates or collects relevant images based on the verified summary. Depending on the extracted keywords, appropriate images are collected from the Internet or selected from an internal image database. The server then combines these images with the summary, overlaying text on the images to generate a slideshow-style video. Video generation software (e.g., FFmpeg) is used to generate a video file of up to 30 seconds.

[1371] Video distribution

[1372] The server automatically uploads the generated video file to a video distribution service (for example, YouTube or Vimeo). The server uses the API of the video distribution service to set the information required when uploading the video (title, description, tags, etc.). Once the upload is complete, the server obtains the URL of the generated video from the video distribution service and saves it in its internal database.

[1373] User Notifications

[1374] The server sends a message to the user's device to notify them that a new video has been uploaded. The device (smartphone or computer) then displays a push notification to the user, saying something like, "A new summary video has been released." After checking the notification, the user clicks on the provided video link to access the video streaming service and watch the 30-second summary video.

[1375] Sentiment Analysis and Content Optimization

[1376] The emotion engine analyzes a user's emotions based on their viewing history and reaction data. For example, the emotion engine learns which genres and content of videos a user has particularly rated highly among those they have watched in the past. Based on the emotion data analyzed by the emotion engine, the server selects and prioritizes the book summaries and videos that best match the user's emotions. Specifically, the server prioritizes books in the user's favorite genres and themes, and generates and distributes summaries and videos based on these.

[1377] Specific examples

[1378] Get book data

[1379] For example, suppose a server retrieves data on bestselling books from an API at midnight every day. For example, suppose a book called "Algorithms for Engineers" is retrieved.

[1380] Summary generation

[1381] The server inputs the data from the above book into a generative AI model and generates a summary like the one below.

[1382] Title: Algorithms for Engineers

[1383] Author: Example author

[1384] Summary: This book introduces fundamental concepts and practical applications of algorithms.

[1385] Generated summary:

[1386] This book introduces fundamental concepts and practical applications of algorithms.

[1387] Video generation

[1388] The server collects related images based on keywords such as "algorithm" and "practice," and combines the summary text and images to generate a 30-second video.

[1389] Video distribution

[1390] The server uploads the generated video to a video distribution service (e.g., YouTube), obtains the video URL, and stores it in a database.

[1391] User Notifications

[1392] The server sends a push notification to the user's smartphone, and the user clicks on the notification to watch the video.

[1393] Sentiment Analysis and Content Optimization

[1394] If the emotion engine analyzes the user's viewing history and recognizes that the user prefers technical books, for example, it will prioritize generating and delivering summaries and videos of technical books from the next time onwards.

[1395] This allows users to efficiently watch video summaries of books that match their interests and emotions.

[1396] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1397] Step 1:

[1398] The server connects to a book database or API to retrieve new book information. It receives a JSON response containing data such as the book title, author, synopsis, and publication date. For example, it sends a GET request to an API and receives JSON data containing book information in response. It then stores this data in an internal database.

[1399] Input: API endpoint

[1400] Output: Book information (title, author, synopsis, publication date)

[1401] Step 2:

[1402] The server generates a prompt based on the acquired book information and inputs it into the generative AI model. The prompt includes the book title, author, and synopsis.

[1403] Example prompt sentence:

[1404] Title: Algorithms for Engineers

[1405] Author: Example author

[1406] Summary: This book introduces fundamental concepts and practical applications of algorithms.

[1407] As a result, a summary sentence is generated from the AI ​​model.

[1408] Input: Book information (title, author, synopsis)

[1409] Output: Summary

[1410] Step 3:

[1411] The server then validates the generated abstract to ensure it is correct, including checking for any invalid language or inappropriate content. If the abstract passes validation, it is sent to the next step.

[1412] Input: Abstract

[1413] Output: Verified summary

[1414] Step 4:

[1415] The server collects related images based on the verified summary, picks out important words from the summary and synopsis to extract keywords, and collects images from the Internet and internal image databases based on these keywords.

[1416] Input: Verified abstract, keywords

[1417] Output: A list of related images

[1418] Step 5:

[1419] The server combines the summary text with the collected images to generate a slideshow-style video using FFmpeg or other video generation tools. The summary text is overlaid on the images to create a video file.

[1420] Input: Verified summary, related images

[1421] Output: Video file

[1422] Step 6:

[1423] The server uploads the generated video file to a video distribution service. At this time, metadata such as the video title, description, and tags are set. For example, the YouTube API is used to upload the video and obtain its URL.

[1424] Input: Video file, metadata

[1425] Output: Video URL

[1426] Step 7:

[1427] The server saves the URL of the new video in an internal database and sends that information to the user's device via a push notification service (e.g., Firebase Cloud Messaging), informing the user that a new summary video has been released.

[1428] Input: Video URL

[1429] Output: Push notification message

[1430] Step 8:

[1431] The server collects users' viewing history and reaction data, and analyzes their emotions using an emotion engine, which determines what genres and content are most suitable for the user.

[1432] Input: User viewing history, reaction data

[1433] Output: Sentiment analysis data

[1434] Step 9:

[1435] Based on the data analyzed by the emotion engine, the server selects the next best book summary video and starts the process of generating and delivering that video, ensuring that content that matches the user's emotions is provided first.

[1436] Input: Sentiment analysis data

[1437] Output: Recommended book information, summary, video URL

[1438] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1439] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1440] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1441] [Fourth embodiment]

[1442] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1443] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1444] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1445] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1446] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1447] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1448] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1449] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1450] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1451] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1452] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1453] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1454] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1455] A natural language description of the program's operation

[1456] This system provides an automated process to summarize international best-selling books in 200 characters or less and distribute them as 30-second videos. The following explains the roles and processes of the server, terminal, and user in detail.

[1457] Get book data

[1458] The server retrieves the book data.

[1459] First, the server connects to a specific book database or API (Application Program Interface) to retrieve new book information. The server runs a scheduled task at a regular time, for example, retrieving the latest best-selling book list at midnight every day. The retrieved data includes the book title, author, synopsis, publication date, etc.

[1460] Summary generation

[1461] The server generates a summary

[1462] Next, the server generates a summary using a generative AI model (such as GPT-3) based on the acquired book information. The server inputs the book title, author, and synopsis into the generative AI model in a prompt format, and the AI ​​model outputs a summary of up to 200 characters. The generated summary is then verified by the server to ensure it does not contain any inappropriate content.

[1463] Video generation

[1464] The server generates a video from the summary text

[1465] Based on the verified summary, the server automatically generates or collects related images. The server extracts keywords from the summary and, depending on the keywords, collects related images from the Internet or selects them from an internal image database. The server then combines these images with the summary, overlaying text on the images to generate a slideshow-style video. Video generation software such as FFmpeg is used to generate video files of up to 30 seconds.

[1466] Video distribution

[1467] The server uploads the video to the video streaming service.

[1468] The server automatically uploads the generated video file to a video distribution service (such as YouTube or Vimeo). The server uses the video distribution service's API to set the information required when uploading the video (title, description, tags, etc.). Once the upload is complete, the video's URL is obtained from the service and saved in an internal database.

[1469] User Notifications

[1470] The device notifies the user of the video link

[1471] The server sends a message to the user's device to notify them that a new video has been uploaded. The device (smartphone or computer) then displays a push notification to the user with a message such as "A new summary video has been released." After confirming the notification, the user clicks on the provided video link to access the video streaming service and watch the 30-second summary video.

[1472] Specific examples

[1473] 1. Acquiring book data

[1474] The server retrieves the bestselling book data from the API at midnight every day. For example, it retrieves the data for the book "The Silent Patient."

[1475] 2. Summary Generation

[1476] The server inputs the data from "The Silent Patient" into a generative AI model and generates the summary: "Alicia is an artist who goes silent after a mysterious incident, and the story is about a therapist who searches for the truth behind it."

[1477] 3. Video Generation

[1478] The server collects free images based on keywords such as "Alicia" and "therapist," and combines the summary text and images to generate a 30-second video.

[1479] 4. Video distribution

[1480] The server uploads the generated video to YouTube, retrieves the video URL and stores it in the database.

[1481] 5. User Notices

[1482] The server sends a push notification to the user's smartphone, and the user clicks on the notification to watch the video.

[1483] This system allows users to easily view summaries of international best-selling books in a short amount of time, enabling them to acquire information efficiently.

[1484] The processing flow will be explained below.

[1485] Step 1:

[1486] The server connects to the book database or API at midnight every day to retrieve the latest bestselling book data, including the book title, author, synopsis, and publication date.

[1487] Step 2:

[1488] The server stores the acquired book information in an internal database and adds it to a processing queue, so that subsequent processes can retrieve and process the book information sequentially.

[1489] Step 3:

[1490] The server retrieves a book from the processing queue and inputs the book's title, author, and synopsis into the generative AI model, providing appropriate input to the AI ​​model in the form of prompts.

[1491] Step 4:

[1492] The generative AI model outputs a summary of up to 200 characters based on the input information. The server receives the summary and prepares it for the next process.

[1493] Step 5:

[1494] The server then validates the generated summary to check for inappropriate content, for example, by checking for inappropriate language or unclear sentences.

[1495] Step 6:

[1496] The server extracts keywords from the verified abstract and passes the list of keywords, including people's names, places, and events, to the image generation module.

[1497] Step 7:

[1498] The image generation module collects free stock images corresponding to each keyword from the Internet or selects them from an internal image database, and the server downloads or retrieves the images.

[1499] Step 8:

[1500] The server combines the abstract with the acquired images to generate a slideshow-style video by overlaying the text on the images. Video generation software (e.g., FFmpeg) is used to generate a video file of up to 30 seconds.

[1501] Step 9:

[1502] The server uploads the generated video file to a video distribution service (e.g., YouTube or Vimeo), and also sets the video's metadata (title, description, tags, etc.) at the same time.

[1503] Step 10:

[1504] Once the upload is complete, the server retrieves the generated video URL from the video streaming service, which is then stored in an internal database.

[1505] Step 11:

[1506] The server sends a message to the user's device informing them that a new video has been uploaded, which includes sending a push notification.

[1507] Step 12:

[1508] The device displays a notification to the user saying, "A new summary video has been released." The user checks the notification and clicks on the provided video link.

[1509] Step 13:

[1510] Users can easily obtain summaries of international best-selling books by watching 30-second summary videos on their device's browser or YouTube app.

[1511] This detailed processing flow automates and efficiently executes the entire process of obtaining book information, generating a summary and video, and finally notifying the user.

[1512] Example 1

[1513] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1514] Current book information systems do not allow users to quickly obtain summaries of international best-selling books, making it difficult for users to grasp the content in a short time. Furthermore, when summaries are provided only in text, the lack of visual elements makes it difficult for users to understand or engage with the information. Therefore, there is a need for a method to provide book information efficiently and in a visually understandable way.

[1515] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1516] In this invention, the server includes means for acquiring book information from a book database or API, means for creating a prompt based on the acquired book information and sending it to the generative AI model, means for verifying the summary returned from the generative AI model, means for extracting keywords from the summary and collecting related images, means for generating a slideshow video based on the collected images and summary, means for uploading the generated video to a video distribution service, and means for sending a notification message including a link to the video to the user terminal. This allows users to quickly and visually understand summaries of international best-selling books.

[1517] A "book database" is a system that systematically stores information about books and has a data structure that allows for searching and extraction.

[1518] An "API" is an interface that allows various pieces of software to communicate with each other and share functions, and is often provided in the form of a web service.

[1519] A "prompt sentence" is an input sentence provided to a generative AI model, a series of text that is generated in anticipation of specific conditions or responses.

[1520] A "generative AI model" is an artificial intelligence algorithm that automatically generates text and information from large amounts of data.

[1521] "Validation" is the process of verifying whether the generated data is appropriate and checking that it does not contain any inappropriate content.

[1522] "Keywords" are important words extracted from summaries and texts, and are used for information retrieval and classification.

[1523] "Image gathering" is the process of acquiring relevant image data based on specific criteria or keywords.

[1524] "Slideshow format" is a visual presentation format that displays multiple images and text sequentially.

[1525] A "video distribution service" is an online platform for uploading videos over the Internet and making them available to viewers.

[1526] A "notification message" is a message sent by the system to convey specific information to the user.

[1527] A "user terminal" is a device such as a smartphone or computer that allows a user to access the Internet and applications.

[1528] To implement this invention, a system is required that automates the process of acquiring book information from a book database or API, generating a summary using a generative AI model based on that information, creating a video using that summary, and finally notifying the user. The following describes in detail the hardware and software used in each processing step, as well as the details of data processing and data calculation.

[1529] Get book data

[1530] Server Roles

[1531] The server connects to a specific book database or API to retrieve new book information. For example, it can use the Goodreads or Google Books API. The server runs a periodic task every day at midnight, sending an HTTP request to retrieve the latest best-selling book list. The retrieved data is provided in JSON format and includes the book title, author, synopsis, publication date, etc.

[1532] Summary generation

[1533] Server Roles

[1534] The server creates a prompt based on the acquired book information and sends it to the generative AI model. Specifically, it combines the book title, author, and synopsis to create a prompt to be input to the generative AI model (e.g., GPT-3). The following is an example of a prompt:

[1535] Example prompt sentence:

[1536] "Book title: 'XXX', author: XXX, summary: XXX. Please write a summary in 200 characters or less."

[1537] The server receives the summary returned by the generative AI model and uses an internal validation algorithm to check for inappropriate content. If there are no problems, the summary is stored in the database.

[1538] Video generation

[1539] Server Roles

[1540] The server extracts keywords from the summary and collects related images. It uses a natural language processing algorithm to extract key keywords from the summary and collects images via HTTP requests from free image services (such as Unsplash and Pexels). It then combines the collected images with the summary to generate a slideshow-style video using video generation software such as FFmpeg. The video must be no longer than 30 seconds.

[1541] Video distribution

[1542] Server Roles

[1543] The server uploads the generated video to a video distribution service (such as YouTube or Vimeo). The server uses the API of the video distribution service to set the necessary information such as title, description, and tags when uploading the video. If the upload is successful, the server obtains the URL of the video returned by the video service and saves it in the database.

[1544] User Notifications

[1545] Device Role

[1546] The server generates a notification message containing the URL of the new video and sends a push notification to the user's device. The device (smartphone or computer) then displays a notification to the user saying, "A new summary video has been released." The user can click the notification to access the video streaming service and watch a 30-second summary video.

[1547] Specific examples

[1548] For example, for the book "Silent Patient," the server sends an HTTP request to the Goodreads API at midnight every day to retrieve information about the book. Next, the following prompt is sent to the generative AI model: "Book Title: 'Silent Patient,' Author: Alex Michaelides, Synopsis: Alicia is an artist who falls silent after a mysterious incident, and the story of a therapist who searches for the truth behind this. Please provide a summary of 200 characters or less." Based on the summary returned by the generative AI model, "Alicia is an artist who falls silent after a mysterious incident, and the story of a therapist who searches for the truth behind this," the server extracts keywords such as "Alicia" and "therapist" and collects related free images using the Unsplash API. These images and the summary are then processed into a slideshow video using FFmpeg. The server then uploads this video through the YouTube API and gives it a title such as "Silent Patient Summary Video." The user's device receives a notification stating, "A new summary video of 'Silent Patient' has been uploaded," and they can click the notification to watch the video.

[1549] This system allows users to quickly obtain summaries of international best-selling books in a visually easy-to-understand manner.

[1550] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1551] Step 1:

[1552] Get book data

[1553] The server connects to the book database or API to retrieve new book information. The server sends an HTTP request at midnight every day to retrieve the latest bestselling book list. The retrieved data is provided in JSON format and includes the book title, author, synopsis, publication date, etc. The input is the book information retrieved from the book database or API, and the output is the parsed book information. The server parses the retrieved data and saves it in its internal database.

[1554] Step 2:

[1555] Creating and sending prompts

[1556] The server creates a prompt based on the acquired book information and sends it to the generative AI model. Specifically, it combines the book title, author, and synopsis to create a prompt to be input to the generative AI model (e.g., GPT-3). The input is the analyzed book information, and the output is the generated prompt. Below is an example of a prompt:

[1557] Example prompt sentence:

[1558] "Book title: 'XXX', author: XXX, summary: XXX. Please write a summary in 200 characters or less."

[1559] Step 3:

[1560] Summary generation and verification

[1561] The server receives the summary returned by the generative AI model. The server checks the generated summary with an internal verification algorithm to ensure that it does not contain inappropriate content. The input is the summary returned by the generative AI model, and the output is the summary after verification. If it is confirmed that it does not contain inappropriate content, the summary is saved in the database.

[1562] Step 4:

[1563] Keyword extraction and image collection

[1564] The server extracts keywords from the abstract and collects related images. It uses a natural language processing algorithm to extract key keywords from the abstract and collects images via HTTP requests from services that provide free images (e.g., Unsplash and Pexels). The input is the verified abstract, and the output is the extracted keywords and collected images.

[1565] Step 5:

[1566] Video generation

[1567] The server generates a slideshow-style video based on the collected images and summaries. The collected images are concatenated and the summaries are overlaid on the images as text. Video generation software such as FFmpeg is used to generate videos of up to 30 seconds. The input is the collected images and summaries, and the output is the generated video file.

[1568] Step 6:

[1569] Uploading videos

[1570] The server uploads the generated video to a video distribution service. Using the API of the video distribution service (for example, YouTube or Vimeo), necessary information such as title, description, and tags is set when uploading the video. The input is the generated video file, and the output is the URL of the uploaded video. If the upload is successful, the server saves the video URL returned by the video service in a database.

[1571] Step 7:

[1572] User Notifications

[1573] The server generates a notification message containing the URL of the new video and sends a push notification to the user's device. The device (smartphone or computer) receives the notification and displays a message to the user saying, "A new summary video has been released." The input is the URL of the uploaded video, and the output is the notification message displayed on the user's device. The user can click the notification to access the video streaming service and watch a 30-second summary video.

[1574] (Application example 1)

[1575] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1576] In today's world, busy people want to obtain information efficiently within a limited amount of time, but it is often difficult to find the time to read long books. Another issue is that simply reading a book summary can weaken comprehension of the content and the impact of the information. Furthermore, there are limited ways for users to easily view summarized book information, and this needs to be addressed.

[1577] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1578] In this invention, the server includes a means for acquiring book information from a book database or an API, a generation AI model means for generating a summary using the acquired book information, and a means for generating a video based on the generated summary, thereby enabling users to efficiently view summarized book information in video format.

[1579] "Book database" refers to one or more data collections that store information about books.

[1580] "API" refers to an interface that allows application programs to communicate with each other.

[1581] A "generative AI model" refers to artificial intelligence that automatically generates text, images, videos, etc. based on specified input data.

[1582] A "summary" is a short summary of the book's contents.

[1583] "Video distribution service" refers to a service that provides video content to users via the Internet.

[1584] "User" refers to an entity that uses the system to receive information or services.

[1585] "Smart devices" refer to mobile information devices with internet connectivity, such as smartphones and tablets.

[1586] MODE FOR CARRYING OUT THE INVENTION

[1587] System Program

[1588] This system retrieves book information from a book database or API, and generates a summary based on that information using a generative AI model. It consists of a series of processes: generating a video based on the generated summary, uploading the video to a video distribution service, and sending a notification to the user. Users can also watch the summary video on their smart devices. The program design for realizing this system is as follows:

[1589] Hardware and software used

[1590] Server hardware: Amazon EC2

[1591] Book data acquisition: Google Books API

[1592] Summary sentence generation: OpenAI GPT-3

[1593] Video generation: FFmpeg

[1594] Video streaming: YouTube Data API

[1595] Push notifications: Firebase Cloud Messaging (FCM)

[1596] Smart devices: smartphones, tablets, etc.

[1597] Data processing and calculation

[1598] The details of how this system works are as follows:

[1599] 1. Acquiring book data

[1600] The server runs a scheduled task at regular intervals, connecting to a book database or API (e.g., Google Books API) to retrieve the latest best-selling book list, including book title, author, synopsis, publication date, etc.

[1601] 2. Summary Generation

[1602] Based on the acquired book information, the server uses a generative AI model (e.g., OpenAI GPT-3) to generate a summary. Specifically, the book title, author, and synopsis are input into the generative AI model in the form of a prompt, and the AI ​​model outputs a summary of up to 200 characters. This summary is then verified by the server to ensure it does not contain any inappropriate content.

[1603] 3. Video Generation

[1604] Based on the verified summary, the server automatically generates or collects related images. The server extracts keywords from the summary and, based on those keywords, collects related images from the Internet or selects them from an internal image database. It then combines these images with the summary, overlaying text on the images to generate a slideshow-style video. It uses video generation software such as FFmpeg to generate video files of up to 30 seconds.

[1605] 4. Video distribution

[1606] The server automatically uploads the generated video file using the API of a video distribution service (such as YouTube or Vimeo). The necessary information (title, description, tags, etc.) is also set when uploading. Once the upload is complete, the video URL is obtained from the service and saved in an internal database.

[1607] 5. User Notices

[1608] The server sends a message to the user's smart device to notify them that a new video has been uploaded. A push notification displays a message to the user, such as "A new summary video has been uploaded." The user checks the notification and clicks on the provided video link to access the video streaming service and watch the 30-second summary video.

[1609] Specific examples

[1610] 1. Acquiring book data

[1611] Book: "The Silent Patient"

[1612] Author: "Alex Michaelides"

[1613] Synopsis: "Alicia Belson is a celebrated artist..."

[1614] 2. Summary Generation

[1615] Prompt statement:

[1616] Title: The Silent Patient

[1617] Author: Alex Michaelides

[1618] Summary: Alicia Belson is a renowned artist whose work is celebrated around the world. One day, she shoots and kills her husband in their home and never speaks again. This story is told from the perspective of a therapist who tries to unravel this mystery.

[1619] Please summarize in 200 characters or less.

[1620] Generated summary:

[1621] Alicia is an artist who goes silent after a mysterious incident, and the story revolves around a therapist who searches for the truth.

[1622] 3. Video Generation

[1623] The generated summary is combined with images related to keywords such as "Alicia" and "therapist," and a video of less than 30 seconds is generated using FFmpeg.

[1624] 4. Video distribution

[1625] Upload a video to YouTube, get the URL and save it in the database.

[1626] 5. User Notices

[1627] Use Firebase Cloud Messaging to send notifications of new video streams to users' smart devices.

[1628] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1629] Step 1:

[1630] The server periodically connects to a book database or API to retrieve the latest best-selling book information, including book title, author, synopsis, and publication date. This allows you to always have access to the latest book information.

[1631] Input: Scheduled task, book database or API

[1632] Output: Retrieved book information (title, author, synopsis, publication date)

[1633] Step 2:

[1634] The server generates a summary using a generative AI model (e.g., OpenAI GPT-3) based on the acquired book information. The book title, author, and synopsis are input into the generative AI model in the form of a prompt, and the AI ​​outputs a summary of up to 200 characters. At this time, the summary is verified to ensure that it does not contain any inappropriate content.

[1635] Input: Book information (title, author, synopsis), generative AI model

[1636] Output: Summary of up to 200 characters

[1637] Step 3:

[1638] The server extracts keywords from the verified summary and uses them to gather relevant images from the internet or select them from an internal image database. It then combines the summary with these images and generates a slideshow-style video by overlaying text on the images. It then uses video generation software such as FFmpeg to generate a video file of up to 30 seconds.

[1639] Input: Abstract, keywords, related images

[1640] Output: Video file up to 30 seconds long

[1641] Step 4:

[1642] The server uploads the generated video via the API of a video distribution service (e.g., YouTube). Information such as the title, description, and tags are also set. Once the upload is complete, the video URL is obtained from the service and saved in an internal database.

[1643] Input: Video file, video streaming service API

[1644] Output: Video URL

[1645] Step 5:

[1646] The server sends a push notification message via Firebase Cloud Messaging to the user's smart device to notify them that a new video has been uploaded, and the user receives the notification and clicks the video link to watch the video.

[1647] Input: Video URL, Firebase Cloud Messaging

[1648] Output: User notification

[1649] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1650] A natural language description of the program's operation

[1651] This system acquires book information from a book database or API, generates a summary based on the acquired information, and generates a video based on the summary. Furthermore, this system incorporates an emotion engine that recognizes the user's emotions, and provides the user with an optimal book summary video based on the acquired emotion information.

[1652] Get book data

[1653] The server retrieves the book data.

[1654] The server connects to a specific book database or API to retrieve data on new bestselling books, including the book's title, author, synopsis, and publication date. The server runs a scheduled task at a regular time to retrieve the data and store it in an internal database.

[1655] Summary generation

[1656] The server generates a summary

[1657] The server uses a generative AI model to generate a summary of up to 200 characters based on the acquired book information. By inputting the book title, author, and synopsis into the generative AI model, the AI ​​model outputs a summary. This summary is then verified by the server to ensure it does not contain any inappropriate content.

[1658] Video generation

[1659] The server generates a video from the summary text

[1660] Based on the verified summary, the server automatically generates or collects relevant images. Depending on the extracted keywords, appropriate images are collected from the Internet or selected from an internal image database. The server then combines these images with the summary, overlaying text on the images to generate a slideshow-style video. Video generation software (e.g., FFmpeg) is used to generate a video file of up to 30 seconds.

[1661] Video distribution

[1662] The server uploads the video to the video streaming service.

[1663] The server automatically uploads the generated video file to a video distribution service (for example, YouTube or Vimeo). The server uses the API of the video distribution service to set the information required when uploading the video (title, description, tags, etc.). Once the upload is complete, the server obtains the URL of the generated video from the video distribution service and saves it in its internal database.

[1664] User Notifications

[1665] The device notifies the user of the video link

[1666] The server sends a message to the user's device to notify them that a new video has been uploaded. The device (smartphone or computer) then displays a push notification to the user, saying something like, "A new summary video has been released." After checking the notification, the user clicks on the provided video link to access the video streaming service and watch the 30-second summary video.

[1667] Sentiment Analysis and Content Optimization

[1668] The server analyzes the user's emotions

[1669] The emotion engine analyzes user emotions based on their viewing history and reaction data. For example, it learns the genres and content that users have particularly rated among the videos they have watched in the past.

[1670] The server delivers the best content

[1671] Based on the emotional data analyzed by the emotion engine, the server selects and prioritizes the book summaries and videos that best fit the user's emotions. Specifically, it prioritizes books in the user's favorite genres and themes, and generates and distributes summaries and videos based on those.

[1672] Specific examples

[1673] 1. Acquiring book data

[1674] The server retrieves data on bestselling books from the API at midnight every day. For example, it retrieves data on the book "Introduction to Algorithms for Engineers."

[1675] 2. Summary Generation

[1676] The server inputs data from "Introduction to Algorithms for Engineers" into a generative AI model and generates a summary statement that reads, "This book introduces the basic concepts and practical applications of algorithms."

[1677] 3. Video Generation

[1678] The server collects related images based on keywords such as "algorithm" and "application," and combines the summary text and images to generate a 30-second video.

[1679] 4. Video distribution

[1680] The server uploads the generated video to YouTube, retrieves the video URL and stores it in the database.

[1681] 5. User Notices

[1682] The server sends a push notification to the user's smartphone, and the user clicks on the notification to watch the video.

[1683] 6. Sentiment Analysis and Content Optimization

[1684] The emotion engine analyzes the user's viewing history, and if, for example, the user often watches technical books, from the next time onwards, it will prioritize summarizing technical books and generating and delivering videos.

[1685] This detailed process allows users to efficiently watch video summaries of books that suit their preferences and emotions.

[1686] The processing flow will be explained below.

[1687] Step 1:

[1688] The server connects to the book database or API at midnight every day to retrieve the latest bestselling book data, including the book title, author, synopsis, and publication date.

[1689] Step 2:

[1690] The server stores the acquired book information in an internal database and adds it to a processing queue, so that subsequent processes can retrieve and process the book information sequentially.

[1691] Step 3:

[1692] The server retrieves a book from the processing queue and inputs the book's title, author, and synopsis into the generative AI model, providing appropriate input to the AI ​​model in the form of prompts.

[1693] Step 4:

[1694] The generative AI model outputs a summary of up to 200 characters based on the input information. The server receives the summary and prepares it for the next process.

[1695] Step 5:

[1696] The server then validates the generated summary to check for inappropriate content, for example, by checking for inappropriate language or unclear sentences.

[1697] Step 6:

[1698] The server extracts keywords from the verified abstract and passes the list of keywords, including people's names, places, and events, to the image generation module.

[1699] Step 7:

[1700] The image generation module collects free stock images corresponding to each keyword from the Internet or selects them from an internal image database, and the server downloads or retrieves the images.

[1701] Step 8:

[1702] The server combines the abstract with the acquired images to generate a slideshow-style video by overlaying the text on the images. Video generation software (e.g., FFmpeg) is used to generate a video file of up to 30 seconds.

[1703] Step 9:

[1704] The server uploads the generated video file to a video distribution service (e.g., YouTube or Vimeo), and also sets the video's metadata (title, description, tags, etc.) at the same time.

[1705] Step 10:

[1706] Once the upload is complete, the server retrieves the generated video URL from the video streaming service, which is then stored in an internal database.

[1707] Step 11:

[1708] The server sends a message to the user's device informing them that a new video has been uploaded, which includes sending a push notification.

[1709] Step 12:

[1710] The device displays a notification to the user saying, "A new summary video has been released." The user checks the notification and clicks on the provided video link.

[1711] Step 13:

[1712] Users can easily obtain summaries of international best-selling books by watching 30-second summary videos on their device's browser or YouTube app.

[1713] Sentiment Analysis and Content Optimization

[1714] Step 14:

[1715] The server collects the user's viewing history and reaction data and sends it to the emotion engine, which analyzes the user's viewing history and reaction data to estimate the user's emotions.

[1716] Step 15:

[1717] Based on the emotional data analyzed by the emotion engine, the server selects the most suitable book summary and video for the user. Specifically, it prioritizes books in genres and themes that the user has given high ratings to, and generates summaries and videos based on those.

[1718] Step 16:

[1719] The server uploads the generated optimal book summary video to a video distribution service and obtains a distribution link.

[1720] Step 17:

[1721] The server sends push notifications to the user's device at the appropriate time, allowing the user to receive a summary video that matches their emotions.

[1722] Step 18:

[1723] The device receives a notification and displays a message to the user saying, "A summary video has been delivered that is recommended for you." The user clicks the link to watch the video.

[1724] This allows the emotional engine to be used to provide summarized videos based on the user's emotions, resulting in even higher satisfaction. To give a concrete example, for example, a user who frequently reads technical books could be given preferential access to videos summarizing new best-selling books in the same genre. In this way, more personalized content is provided to users through sentiment analysis and content optimization.

[1725] Example 2

[1726] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1727] Conventional book recommendation systems have the problem that the process of generating a book summary and distributing it as a video incorporating visuals is complicated and difficult to automate. It is also difficult to provide personalized content that takes into account the user's preferences and emotions. This makes it difficult to attract the user's interest, resulting in a poor content consumption experience.

[1728] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for acquiring book information from a book database or API, a generation AI model means for generating a summary using the acquired book information, means for verifying the generated summary, means for automatically generating or collecting related images based on the verified summary, means for generating a video by combining the generated images and summary, means for uploading the generated video to a video distribution service, means for acquiring and saving the URL of the uploaded video, means for notifying the user, means for analyzing the user's emotions, and means for providing an optimal book summary video based on the analysis results. This automates the process from generating book summaries to distributing videos and providing content optimized for users, enabling the provision of an efficient and high-quality user experience.

[1729] "Book information" refers to basic data such as the book's title, author, synopsis, and publication date.

[1730] A "generative AI model" refers to an algorithm or software that automatically generates a summary from input data using natural language processing technology.

[1731] "Video Streaming Service" means a platform for hosting generated video content online and for users to stream or download it.

[1732] A "means" refers to a specific method, device, or technique for performing a particular function.

[1733] "Means for analyzing emotions" refers to algorithms and software for estimating and analyzing a user's emotional state based on their viewing history and reaction data.

[1734] A "summary" is a text that compresses the original book information into 200 characters or less and succinctly presents the main content of the book.

[1735] "Verification measures" refers to processes or algorithms that check and evaluate the accuracy and appropriateness of generated summaries and other data.

[1736] "Means of notification" refers to the technology and process used to communicate new video distribution information, etc. to users' devices.

[1737] "Optimal book summary video" refers to the book summary video that is determined to be most relevant based on the user's interests and emotional state.

[1738] "Means for automatically generating or collecting relevant images" refers to algorithms that generate relevant images from text, or technologies that collect appropriate images from existing databases or the internet.

[1739] This invention relates to a system that acquires book information from a book database or API, generates a summary using that information, and then generates a video based on the summary. The system includes an emotion engine that recognizes a user's emotions and provides optimal content based on those emotions. The following describes in detail the embodiments of the invention.

[1740] Get book data

[1741] The server retrieves the book data.

[1742] The server connects to a book database or API (for example, a common book database API) and periodically retrieves data on new best-selling books. This data includes the book's title, author, synopsis, and publication date. The server runs a scheduled task and stores the retrieved data in an internal database. Specifically, the server retrieves the data from the API at midnight every day and stores it in an internal database (for example, MySQL or MongoDB).

[1743] Summary generation

[1744] The server generates a summary

[1745] The server uses a generative AI model (e.g., GPT-4) to generate a summary of up to 200 characters based on the acquired book information. The input includes the book title, author, and synopsis. For example, data from the book "Introduction to Algorithms for Engineers" is input into the generative AI model, and the resulting summary is, "This book introduces the basic concepts and practical applications of algorithms."

[1746] Video generation

[1747] The server generates a video from the summary text

[1748] The server automatically generates or collects related images based on the verified summary. Depending on the extracted keywords, appropriate images are collected from the Internet or selected from an internal image database. For example, related images are collected based on keywords such as "algorithm" and "application." The server then combines the images with the summary, overlaying text on the images to generate a slideshow-style video. Specifically, video generation software such as FFmpeg is used to generate video files of up to 30 seconds.

[1749] Video distribution

[1750] The server uploads the video to the video streaming service.

[1751] The server automatically uploads the generated video file using the API of a video distribution service (for example, YouTube or Vimeo). When uploading, necessary information (title, description, tags, etc.) is set, and once the upload is complete, the server obtains the URL of the generated video from the video distribution service and saves it in an internal database. For example, the video title could be "Introduction to Algorithms for Engineers - Summary Video" and the description "Introducing the basic concepts of algorithms in 30 seconds."

[1752] User Notifications

[1753] The device notifies the user of the video link

[1754] The server sends a message to the user's device to notify them that a new video has been uploaded. The device (smartphone or computer) then displays a push notification to the user, saying something like, "A new summary video has been released." After checking the notification, the user clicks on the provided video link to access the video streaming service and watch the 30-second summary video.

[1755] Sentiment Analysis and Content Optimization

[1756] The server analyzes the user's emotions

[1757] The emotion engine analyzes user emotions based on their viewing history and reaction data, for example, learning the genres and content of videos they have particularly rated.

[1758] The server delivers the best content

[1759] Based on the analysis results, the server selects and delivers the book summaries and videos that best suit the user's emotions. Specifically, it prioritizes books in the user's favorite genres and themes, and generates and delivers summaries and videos based on that. For example, if a user frequently reads technical books, the server will continue to generate summaries and videos for technical books from the next time.

[1760] Specific examples

[1761] Specific examples of book data acquisition

[1762] The server retrieves data for "Introduction to Algorithms for Engineers" from the book database API at midnight every day. This data includes the book title, author, synopsis, and publication date.

[1763] A concrete example of summary generation

[1764] The server inputs data from "Introduction to Algorithms for Engineers" into a generative AI model and generates a summary statement that reads, "This book introduces the basic concepts and practical applications of algorithms."

[1765] Example of video generation

[1766] The server collects relevant images based on keywords such as "algorithm" and "application," and uses FFmpeg to generate a 30-second video that combines the summary text and images.

[1767] Specific examples of video distribution

[1768] The server uploads the generated video to YouTube, obtains the video URL "https: / / www.youtube.com / watch?v=example", and stores it in the database.

[1769] Example of user notification

[1770] The server sends a push notification to the user's smartphone saying, "A new summary video has been released." The user clicks the notification to watch the video.

[1771] Examples of sentiment analysis and content optimization

[1772] The emotion engine analyzes the user's viewing history, and if they frequently watch technical books, it will prioritize similar genres for summarization and generate videos from them next time.

[1773] Prompt Sentence Examples

[1774] "Book Title: Introduction to Algorithms for Engineers Author: Taro Yamada Summary: This book introduces the basic concepts and practical applications of algorithms."

[1775] This invention allows users to efficiently watch video summaries of books that match their emotions and interests.

[1776] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1777] Step 1:

[1778] Get book data

[1779] The server connects to a book database or API and periodically retrieves data on new bestselling books. First, the server sends a request to a specific API (e.g., a general book database API) at midnight every day. In response to this request, the API returns book information (title, author, synopsis, publication date). The server stores the retrieved data in an internal database (e.g., MySQL or MongoDB).

[1780] Input: API request / response

[1781] Output: Book information stored in the internal database

[1782] Step 2:

[1783] Summary generation

[1784] The server generates a summary using book information retrieved from an internal database. Specifically, it inputs the book title, author, and synopsis into a generative AI model (e.g., GPT-4). Example prompt:

[1785] "Book Title: Introduction to Algorithms for Engineers Author: Taro Yamada Summary: This book introduces the basic concepts and practical applications of algorithms."

[1786] The generative AI model outputs a summary of up to 200 characters, which is then verified by the server to ensure it does not contain inappropriate content.

[1787] Input: Book information, prompt for the generative AI model

[1788] Output: Summary of up to 200 characters

[1789] Step 3:

[1790] Abstract verification

[1791] The server verifies the generated summary. It uses an automatic verification algorithm to check whether the summary contains any prohibited words. It also checks for grammatical errors and inappropriate content. Once the verification is complete, the summary proceeds to the next process.

[1792] Input: Generated summary

[1793] Output: Verified summary

[1794] Step 4:

[1795] Collection of related images

[1796] The server collects related images based on keywords extracted from the abstract. For example, it searches for and retrieves appropriate images from image databases on the Internet or internal image databases based on keywords such as "algorithm" or "application."

[1797] Input: Keywords extracted from the abstract

[1798] Output: Associated image data

[1799] Step 5:

[1800] Combining images and abstracts

[1801] The server combines the collected images with the verified summaries to generate a slideshow-style video. Specifically, it uses video generation software such as FFmpeg to overlay the summaries on the images and create a video file of up to 30 seconds in length.

[1802] Input: Related images, verified summary

[1803] Output: Generated video file

[1804] Step 6:

[1805] Uploading videos

[1806] The server uploads the generated video file to a video distribution service (e.g., YouTube or Vimeo). At this time, information such as the video title, description, and tags are set. Once the upload is complete, the server obtains the video URL from the video distribution service and saves it in an internal database.

[1807] Input: Generated video file, video information (title, description, tags)

[1808] Output: Video URL, saved to internal database

[1809] Step 7:

[1810] User Notifications

[1811] The server sends a message to the user's device to notify them that a new video has been uploaded. The device (smartphone or computer) then displays the message to the user via a push notification. For example, a notification could be sent to the user saying, "A new summary video has been uploaded."

[1812] Input: Video URL, notification message

[1813] Output: Notification displayed on the user's device

[1814] Step 8:

[1815] sentiment analysis

[1816] The server collects user viewing history and reaction data and uses an emotion engine to analyze the user's emotions. For example, it infers the user's emotional state based on data such as ratings, viewing time, and comments on videos viewed in the past.

[1817] Input: Viewing history, reaction data

[1818] Output: User emotion data

[1819] Step 9:

[1820] Content Optimization

[1821] The server selects the most suitable book summary video for the user based on the analysis results of the emotion engine. It prioritizes books in the user's favorite genres and themes, and generates the summary text and video again based on that. This makes it possible to provide personalized content to the user.

[1822] Input: User emotion data, book information

[1823] Output: Optimized book summary video

[1824] (Application example 2)

[1825] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1826] When viewing book summaries and videos, it is difficult to provide optimal content that matches the user's interests. Another issue is that the content presented to the user does not necessarily match the user's preferences. Furthermore, there is a need for a system that can efficiently generate summaries and videos from a vast amount of book information and optimize them based on the user's emotions.

[1827] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1828] In this invention, the server includes means for acquiring book information from a book database or API, a generation AI model means for generating summaries using the acquired book information, means for generating videos based on the generated summaries, means for uploading the generated videos to a video distribution service, means for notifying the user, an emotion engine means for analyzing the user's emotions, and means for optimizing and distributing book summary videos to the user based on the emotion engine means. This makes it possible to provide optimal book summary videos based on the user's emotions and improve the viewing experience.

[1829] A "book database" is a database that accumulates and organizes information about books.

[1830] "API" stands for Application Programming Interface, an interface for exchanging functions and data between different software programs.

[1831] A "generative AI model" is a model that uses artificial intelligence to generate output data from specific input data. In this invention, it refers to a model that generates a summary from book information.

[1832] The "emotion engine" is an engine that analyzes users' emotions and reactions and provides optimal content based on that data.

[1833] A "video distribution service" is a service that distributes video content over the Internet.

[1834] A "summary" is a sentence that summarizes and shortens a long sentence, and in the present invention, it is intended to express the main content of a book concisely.

[1835] "Optimization" means adjusting and improving to best suit a specific purpose or condition. In this invention, it refers to optimizing the book summary video based on user sentiment.

[1836] "Notification" is the act of informing a user of a certain fact or information, and in the present invention, it is intended to notify the user of the distribution of a new digest video.

[1837] This system acquires book information from a book database or API, generates a summary based on the acquired information, and generates a video based on the summary. Furthermore, this system incorporates an emotion engine that recognizes the user's emotions, and provides the user with an optimal book summary video based on the acquired emotion information.

[1838] Get book data

[1839] The server connects to a specific book database or API to retrieve data on new bestselling books, including the book's title, author, synopsis, and publication date. The server runs a scheduled task at a regular time to retrieve the data and store it in an internal database.

[1840] Summary generation

[1841] The server uses a generative AI model to generate a summary of up to 200 characters based on the acquired book information. By inputting the book title, author, and synopsis into the generative AI model, the AI ​​model outputs a summary. This summary is then verified by the server to ensure it does not contain any inappropriate content.

[1842] Video generation

[1843] The server automatically generates or collects relevant images based on the verified summary. Depending on the extracted keywords, appropriate images are collected from the Internet or selected from an internal image database. The server then combines these images with the summary, overlaying text on the images to generate a slideshow-style video. Video generation software (e.g., FFmpeg) is used to generate a video file of up to 30 seconds.

[1844] Video distribution

[1845] The server automatically uploads the generated video file to a video distribution service (for example, YouTube or Vimeo). The server uses the API of the video distribution service to set the information required when uploading the video (title, description, tags, etc.). Once the upload is complete, the server obtains the URL of the generated video from the video distribution service and saves it in its internal database.

[1846] User Notifications

[1847] The server sends a message to the user's device to notify them that a new video has been uploaded. The device (smartphone or computer) then displays a push notification to the user, saying something like, "A new summary video has been released." After checking the notification, the user clicks on the provided video link to access the video streaming service and watch the 30-second summary video.

[1848] Sentiment Analysis and Content Optimization

[1849] The emotion engine analyzes a user's emotions based on their viewing history and reaction data. For example, the emotion engine learns which genres and content of videos a user has particularly rated highly among those they have watched in the past. Based on the emotion data analyzed by the emotion engine, the server selects and prioritizes the book summaries and videos that best match the user's emotions. Specifically, the server prioritizes books in the user's favorite genres and themes, and generates and distributes summaries and videos based on these.

[1850] Specific examples

[1851] Get book data

[1852] For example, suppose a server retrieves data on bestselling books from an API at midnight every day. For example, suppose a book called "Algorithms for Engineers" is retrieved.

[1853] Summary generation

[1854] The server inputs the data from the above book into a generative AI model and generates a summary like the one below.

[1855] Title: Algorithms for Engineers

[1856] Author: Example author

[1857] Summary: This book introduces fundamental concepts and practical applications of algorithms.

[1858] Generated summary:

[1859] This book introduces fundamental concepts and practical applications of algorithms.

[1860] Video generation

[1861] The server collects related images based on keywords such as "algorithm" and "practice," and combines the summary text and images to generate a 30-second video.

[1862] Video distribution

[1863] The server uploads the generated video to a video distribution service (e.g., YouTube), obtains the video URL, and stores it in a database.

[1864] User Notifications

[1865] The server sends a push notification to the user's smartphone, and the user clicks on the notification to watch the video.

[1866] Sentiment Analysis and Content Optimization

[1867] If the emotion engine analyzes the user's viewing history and recognizes that the user prefers technical books, for example, it will prioritize generating and delivering summaries and videos of technical books from the next time onwards.

[1868] This allows users to efficiently watch video summaries of books that match their interests and emotions.

[1869] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1870] Step 1:

[1871] The server connects to a book database or API to retrieve new book information. It receives a JSON response containing data such as the book title, author, synopsis, and publication date. For example, it sends a GET request to an API and receives JSON data containing book information in response. It then stores this data in an internal database.

[1872] Input: API endpoint

[1873] Output: Book information (title, author, synopsis, publication date)

[1874] Step 2:

[1875] The server generates a prompt based on the acquired book information and inputs it into the generative AI model. The prompt includes the book title, author, and synopsis.

[1876] Example prompt sentence:

[1877] Title: Algorithms for Engineers

[1878] Author: Example author

[1879] Summary: This book introduces fundamental concepts and practical applications of algorithms.

[1880] As a result, a summary sentence is generated from the AI ​​model.

[1881] Input: Book information (title, author, synopsis)

[1882] Output: Summary

[1883] Step 3:

[1884] The server then validates the generated abstract to ensure it is correct, including checking for any invalid language or inappropriate content. If the abstract passes validation, it is sent to the next step.

[1885] Input: Abstract

[1886] Output: Verified summary

[1887] Step 4:

[1888] The server collects related images based on the verified summary, picks out important words from the summary and synopsis to extract keywords, and collects images from the Internet and internal image databases based on these keywords.

[1889] Input: Verified abstract, keywords

[1890] Output: A list of related images

[1891] Step 5:

[1892] The server combines the summary text with the collected images to generate a slideshow-style video using FFmpeg or other video generation tools. The summary text is overlaid on the images to create a video file.

[1893] Input: Verified summary, related images

[1894] Output: Video file

[1895] Step 6:

[1896] The server uploads the generated video file to a video distribution service. At this time, metadata such as the video title, description, and tags are set. For example, the YouTube API is used to upload the video and obtain its URL.

[1897] Input: Video file, metadata

[1898] Output: Video URL

[1899] Step 7:

[1900] The server saves the URL of the new video in an internal database and sends that information to the user's device via a push notification service (e.g., Firebase Cloud Messaging), informing the user that a new summary video has been released.

[1901] Input: Video URL

[1902] Output: Push notification message

[1903] Step 8:

[1904] The server collects users' viewing history and reaction data, and analyzes their emotions using an emotion engine, which determines what genres and content are most suitable for the user.

[1905] Input: User viewing history, reaction data

[1906] Output: Sentiment analysis data

[1907] Step 9:

[1908] Based on the data analyzed by the emotion engine, the server selects the next best book summary video and starts the process of generating and delivering that video, ensuring that content that matches the user's emotions is provided first.

[1909] Input: Sentiment analysis data

[1910] Output: Recommended book information, summary, video URL

[1911] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1912] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1913] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1914] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1915] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1916] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1917] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1918] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1919] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1920] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1921] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1922] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1923] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1924] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1925] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1926] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1927] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1928] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1929] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1930] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1931] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1932] The following is further disclosed regarding the above embodiment.

[1933] (Claim 1)

[1934] A means for obtaining book information from a book database or API;

[1935] A generation AI model means for generating a summary using the acquired book information;

[1936] A means for generating a video based on the generated summary sentence;

[1937] A means for uploading the generated video to a video distribution service;

[1938] a means of notifying the user; and

[1939] A system including:

[1940] (Claim 2)

[1941] 2. The system according to claim 1, wherein the book information includes data on best-selling books.

[1942] (Claim 3)

[1943] The system of claim 1, wherein the generated summary is limited to 200 characters or less.

[1944] (Claim 4)

[1945] The system of claim 1, wherein the length of the generated video is limited to 30 seconds or less.

[1946] (Claim 5)

[1947] The system of claim 1, further comprising means for automatically generating related images based on the summary text.

[1948] "Example 1"

[1949] (Claim 1)

[1950] A means for obtaining book information from a book database or API;

[1951] A means for creating a prompt sentence based on the acquired book information and sending it to the generative AI model;

[1952] A means for validating the summary returned by the generative AI model; and

[1953] A means for extracting keywords from the abstract and collecting related images;

[1954] A means for generating a slideshow-style video based on the collected images and summaries;

[1955] A means for uploading the generated video to a video distribution service;

[1956] means for sending a notification message including a video link to a user terminal;

[1957] A system including:

[1958] (Claim 2)

[1959] 2. The system according to claim 1, wherein the book information includes data on best-selling books.

[1960] (Claim 3)

[1961] The system of claim 1, wherein the generated summary is limited to 200 characters or less.

[1962] "Application Example 1"

[1963] (Claim 1)

[1964] A means for obtaining book information from a book database or API;

[1965] A generation AI model means for generating a summary using the acquired book information;

[1966] A means for generating a video based on the generated summary sentence;

[1967] A means for uploading the generated video to a video distribution service;

[1968] a means of notifying the user; and

[1969] A means for the user to watch the summary video on a smart device;

[1970] A system including:

[1971] (Claim 2)

[1972] 2. The system according to claim 1, wherein the book information includes data on best-selling books.

[1973] (Claim 3)

[1974] The system of claim 1, wherein the generated summary is limited to 200 characters or less.

[1975] "Example 2: Combining Emotion Engines"

[1976] (Claim 1)

[1977] A means for obtaining book information from a book database or API;

[1978] A generation AI model means for generating a summary using the acquired book information;

[1979] a means for verifying the generated summary;

[1980] means for automatically generating or collecting relevant images based on the verified abstract sentences;

[1981] A means for generating a video by combining the generated images and summary text;

[1982] A means for uploading the generated video to a video distribution service;

[1983] A means to retrieve and save the URL of the uploaded video, and

[1984] a means of notifying the user; and

[1985] A means of analyzing user sentiment;

[1986] A means for providing optimal book summary videos based on the analysis results;

[1987] A system including:

[1988] (Claim 2)

[1989] 2. The system according to claim 1, wherein the book information includes data on best-selling books.

[1990] (Claim 3)

[1991] The system of claim 1, wherein the generated summary is limited to 200 characters or less.

[1992] "Application example 2 when combining emotion engines"

[1993] (Claim 1)

[1994] A means for obtaining book information from a book database or API;

[1995] A generation AI model means for generating a summary using the acquired book information;

[1996] A means for generating a video based on the generated summary sentence;

[1997] A means for uploading the generated video to a video distribution service;

[1998] a means of notifying the user; and

[1999] an emotion engine means for analyzing user emotions;

[2000] A means for optimizing and delivering a book summary video to a user based on an emotion engine means;

[2001] A system including:

[2002] (Claim 2)

[2003] 2. The system according to claim 1, wherein the book information includes data on best-selling books.

[2004] (Claim 3)

[2005] The system of claim 1, wherein the generated summary is limited to 200 characters or less. [Explanation of symbols]

[2006] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. A means for obtaining book information from a book database or API; A generation AI model means for generating a summary using the acquired book information; A means for generating a video based on the generated summary sentence; A means for uploading the generated video to a video distribution service; a means of notifying the user; and A system including:

2. 2. The system according to claim 1, wherein the book information includes data on best-selling books.

3. The system of claim 1 , wherein the generated summary is limited to 200 characters or less.

4. The system of claim 1 , wherein the length of the generated video is limited to 30 seconds or less.

5. The system of claim 1 further comprising means for automatically generating related images based on the abstract.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A