System

The system addresses the challenge of inefficient personalized content distribution by using generative AI to create and deliver videos tailored to user interests and emotions, improving user satisfaction and engagement.

JP2026014271APending Publication Date: 2026-01-29SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024115268
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-18
Publication Date
2026-01-29

AI Technical Summary

Technical Problem

Conventional information distribution systems struggle to automatically generate and efficiently distribute personalized content for each user, leading to low user satisfaction and high costs due to manual content creation and distribution.

Method used

A system that utilizes a server to acquire user interests from past behavioral data and registration information, generates a script using generative AI, collects or generates image materials, combines them with text-to-speech narration to create videos, and distributes them to platforms for user devices, optimizing content for individual user preferences.

Benefits of technology

This system enables automatic generation and distribution of personalized video content, enhancing user satisfaction by providing timely and relevant information tailored to individual user interests and emotional states.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026014271000001_ABST
    Figure 2026014271000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: An information distribution system comprising: means for generating a script suitable for each user using generation-related AI on the basis of the acquired region of interest of the user; means for collecting or generating related image materials on the basis of the generated script; means for combining the script and the image materials and adding narration using a reading program to generate a moving image; and means for distributing the generated moving image to a platform and notifying the moving image to terminals of users.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] In recent years, with the spread of the Internet, the types and amounts of information that users want to obtain have increased dramatically. However, because a large amount of information is provided uniformly, it is difficult for users to efficiently obtain information that matches their individual interests. In addition, there is a lack of systems that regularly deliver information that matches users' interests in video format, and there is a need for a means to accurately meet individual requests. [Means for solving the problem]

[0005] The present invention provides a means for acquiring a user's areas of interest from the user's past behavioral data and registered information. It also provides a means for generating a script tailored to each user using generative AI based on the acquired information. It also includes a means for collecting or generating related image materials based on the script. It further provides a means for combining the script and image materials, adding narration using a text-to-speech program, and generating a video. Finally, it provides a means for distributing the generated video to a platform and notifying the user's device. This series of means allows users to regularly obtain content optimized for their interests.

[0006] "User past behavioral data" refers to information that includes the history of various actions that a user has taken on an online platform in the past.

[0007] "Registration Information" means the personal information and interest data provided by a User to the Platform.

[0008] An "interest area" is a range of information related to a particular topic or category that interests a user.

[0009] "Generative AI" is an artificial intelligence technology that automatically generates content such as text and images based on the user's areas of interest.

[0010] A "script" is text content generated based on specific information for reading aloud or generating video.

[0011] "Image Material" means photographs, graphics, charts, and other visual elements used to create the Video Content.

[0012] A "reading program" is software that converts text data into audio data and automatically reads it aloud.

[0013] "Video generation" is the process of combining the original data, such as a script and image materials, to create a single piece of video content.

[0014] "Platform" refers to the digital service infrastructure for delivering generated videos to users.

[0015] A "notification" is a means of communication that notifies a user's device that new content is available. [Brief explanation of the drawings]

[0016] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12]FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0017] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0018] First, the terms used in the following description will be explained.

[0019] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0020] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0021] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0022] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0023] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0024] [First embodiment]

[0025] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0026] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0027] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0028] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0029] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0030] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0031] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0032] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0033] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0034] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0035] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0036] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0037] The present invention provides a system for automatically generating and distributing video content optimized for each user based on the user's past behavioral data and registration information. Hereinafter, a specific embodiment of the present invention will be described.

[0038] Retrieving User Information

[0039] The server obtains the user's past behavioral data and registration information, including the articles the user has viewed on the online platform, the portfolio stocks registered, search history, and categories of interest.

[0040] Script generation

[0041] Based on the acquired user information, the server uses generative AI to generate a script tailored to each user, specifically generating news, market trends, product descriptions, and other information related to the user's areas of interest in natural language.

[0042] example:

[0043] If User A is interested in "technology-related stock information," the server generates scripts about "today's technology stock market trends" and "earnings reports for specific companies."

[0044] Image material generation and acquisition

[0045] Based on the generated script, the server collects or generates the necessary image materials, such as graphs and charts showing market trends, photos related to news articles, etc. It can also use generative AI to generate new image materials as needed.

[0046] example:

[0047] If the script for user A is "about the rise in stock prices of ABC Company," the logo of ABC Company and a graph showing fluctuations in stock prices will be collected as image materials.

[0048] Video generation

[0049] The server combines the generated script with the collected image materials to generate a video. A text-to-speech program converts the script into audio and synchronizes it with the image materials.

[0050] example:

[0051] A script is generated saying "ABC Company's stock price has risen by 5%" and a video with narration is generated along with the corresponding stock price graph.

[0052] Video distribution

[0053] The server uploads the generated video to a designated platform (e.g., a social media app) and distributes it to users. The notification function notifies users of the availability of new content on their devices.

[0054] example:

[0055] User A receives a notification on his device saying, "A new video about technology stocks has been added." User A receives the notification and can watch the video.

[0056] User Viewing

[0057] Users receive notifications and can watch videos on their devices. Content tailored to their interests and preferences is available, resulting in high levels of satisfaction.

[0058] example:

[0059] User A opens the LINE app and watches a video streamed from the server. The video includes information on "technology stock market trends" and "details on the rise in ABC Company's stock price."

[0060] Through this series of processes, the system effectively provides content that users are interested in, increasing user satisfaction. In addition, by automatically generating and delivering video content customized for each user, the way information is received becomes more personalized.

[0061] The processing flow will be explained below.

[0062] Step 1:

[0063] The server retrieves the user's past behavioral data and registration information, which includes collecting data such as the articles the user has viewed, the portfolio holdings they have registered, their search history, and categories of interest.

[0064] Step 2:

[0065] The server analyzes the acquired data to identify the user's areas of interest, including the categories the user frequently browses and extracting related keywords.

[0066] Step 3:

[0067] The server uses generative AI to generate a script tailored to each user, which generates news, market trends, product descriptions, and other information in natural language based on the user's areas of interest.

[0068] Step 4:

[0069] Based on the generated script, the server collects or generates the necessary image materials, such as graphs and charts showing market trends and photos related to the news article, and, if necessary, uses generative AI to create new image materials.

[0070] Step 5:

[0071] The server combines the image material with the generated script to create a video, using a text-to-speech program to convert the script into audio and synchronize it with the image material.

[0072] Step 6:

[0073] The server uploads the generated video file to the specified platform and prepares it for distribution. As soon as the video is published on the distribution platform, a notification is sent to the user's device.

[0074] Step 7:

[0075] Users receive a notification sent to their device to watch the video. Clicking on the notification opens the application and plays the generated video. Users can review the video content and find out the information they are interested in.

[0076] Example 1

[0077] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0078] Conventional information distribution systems have struggled to automatically generate and efficiently distribute personalized content for each user. As a result, it has been difficult to provide information tailored to the user's interests, often resulting in low user satisfaction. Furthermore, manual content creation and distribution is time-consuming and costly, so efficient operation is required.

[0079] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0080] In this invention, the server includes means for acquiring a user's area of ​​interest from the user's past behavioral data and registration information, means for generating a script suitable for each user using a generative AI based on the acquired user's area of ​​interest, means for collecting or generating related image materials based on the generated script, means for combining the script and image materials and generating audio and video using a text-to-speech program, means for distributing the generated video to a specific platform and notifying the user's device, and means for viewing the notified content on the user's device. This makes it possible to automatically generate and distribute video content personalized for each user, thereby achieving high user satisfaction.

[0081] "User past behavioral data" refers to data such as the actions a user takes on an online platform, browsing history, search history, and purchase history.

[0082] "Registration Information" means the personal information and interest category information provided by a User when registering for the Service.

[0083] "User interest areas" are the range of categories or topics in which a user is interested, identified based on the user's behavioral data and registration information.

[0084] "Generative AI" is an artificial intelligence model that generates natural language text, images, and other content based on data and prompts provided.

[0085] A "script" is text containing narration and content instructions for a video, generated based on information related to the user's area of ​​interest.

[0086] "Image material" refers to visual content used in a video, such as photographs, graphs, charts, logos, etc.

[0087] A "read-aloud program" is software that converts text into synthetic speech and plays it back as narration.

[0088] "Video" means dynamic visual content consisting of multiple frames and may include audio and text.

[0089] "Platform" is a general term for websites and applications that allow users to upload generated videos and make them viewable.

[0090] "Notifications" are messages that inform users when newly generated videos or content is available.

[0091] "Device" means a device on which a user receives notifications and views content, including a smartphone, tablet, or PC.

[0092] The present invention provides an information distribution system that automatically generates and distributes video content optimized for each user. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS The following describes in detail an embodiment of the present invention.

[0093] Retrieving User Information

[0094] The server retrieves the user's past behavioral data and registration information from the database, including browsing history, registered portfolio stocks, search history, category information of interest, etc. When the server retrieves this data, it uses the user ID to efficiently collect related data.

[0095] As a concrete example, the server issues an SQL query to retrieve data as follows:

[0096] SELECT FROM user_behavior WHERE user_id = 'userA';

[0097] Script generation

[0098] The server uses generative AI (e.g., GPT-4) to generate a script tailored to each user based on the acquired user information, and sends prompts to generate news, market trends, and product descriptions related to the user's areas of interest in natural language.

[0099] As a concrete example, if user A is interested in technology stock information, the server might use a prompt like this:

[0100] User A is interested in technology stocks. Create a news script for this user.

[0101] In response to this prompt, the generative AI outputs a script that reads, "Today's technology market trends: ABC Company's stock price rose 5%."

[0102] Image material generation and acquisition

[0103] The server collects or generates the necessary image materials based on the generated script, such as photos, graphs, charts, etc., using external data sources or generative AI (e.g., DALL-E).

[0104] You can use the API to obtain image materials. For example, issue the following API request to obtain image materials.

[0105] GET / stock_images?company=ABC

[0106] Video generation

[0107] The server combines the generated script with image materials and generates a video using automatic video generation software (e.g., Adobe Premiere Pro API), and also converts the script into audio using a speech synthesis API (e.g., Google Text-to-Speech) and embeds it in the video.

[0108] As a concrete example, an API request is issued as follows to generate a video.

[0109] POST / generate_video

[0110] {

[0111] "script": "Today's technology market trends: ABC Company's stock price rose 5%.",

[0112] "images": ["abc_logo.png", "stock_chart.png"],

[0113] "voice": "synthesized voice file"

[0114] }

[0115] Video distribution

[0116] The server uploads the generated video to a specific platform (e.g., YouTube) and notifies the user's device. The server then uses the platform's API to upload the video and obtain its URL.

[0117] As a concrete example, you can upload a video by issuing an API request as follows:

[0118] POST / upload_video

[0119] {

[0120] "video_file": "generated_video.mp4",

[0121] "platform": "YouTube"

[0122] }

[0123] After uploading, a notification message will be sent to the user's device.

[0124] Notification: "New technology stock info video added. Watch it here: [URL]"

[0125] User Viewing

[0126] The user receives a notification on their device, opens an application (e.g., a messaging app) to watch the video, and can tap the notification to go directly to the viewing page.

[0127] For example, when a user taps on a notification on their device, a messaging app opens with the URL of the specified video, which the user can click to watch the video.

[0128] Through these steps, the present invention realizes the automatic generation and distribution of video content optimized for each user, and builds an information distribution system that provides high levels of user satisfaction.

[0129] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0130] Step 1: Get user information

[0131] The server retrieves the user's past behavioral data and registration information from the database. As input, the user ID is provided to the server, and the server executes a query against the database based on this user ID. Specifically, the server issues an SQL query as follows:

[0132] SELECT FROM user_behavior WHERE user_id = 'userA';

[0133] This query outputs data related to the user's areas of interest (browsing history, registered portfolio stocks, search history, interest categories, etc.).

[0134] Step 2: Generate the script

[0135] The server uses the acquired user information as input and sends a prompt to the generative AI (e.g., GPT-4) to generate a script. Specifically, the server sends the following prompt to the generative AI:

[0136] User A is interested in technology stocks. Create a news script for this user.

[0137] Based on this prompt, the generative AI outputs text related to the user's area of ​​interest (e.g., "Today's technology market trends: ABC Company's stock price rose 5%.").

[0138] Step 3: Creating and acquiring image materials

[0139] The server uses the generated script as input to collect or generate the necessary image materials. Specifically, the server issues an API request to an external data source.

[0140] GET / stock_images?company=ABC

[0141] This request will output relevant image material (e.g., ABC company logo, stock price graph).

[0142] Step 4: Generate the video

[0143] The server generates a video using automatic video generation software (e.g., Adobe Premiere Pro API) using the generated script and collected image materials as input. Specifically, the server issues the following API request:

[0144] POST / generate_video

[0145] {

[0146] "script": "Today's technology market trends: ABC Company's stock price rose 5%.",

[0147] "images": ["abc_logo.png", "stock_chart.png"],

[0148] "voice": "synthesized voice file"

[0149] }

[0150] This request causes the generated video to be output.

[0151] Step 5: Publish your video

[0152] The server takes the generated video as input, uploads it to a specific platform (e.g., YouTube), and notifies the user's device. Specifically, the server issues the following API request:

[0153] POST / upload_video

[0154] {

[0155] "video_file": "generated_video.mp4",

[0156] "platform": "YouTube"

[0157] }

[0158] This action will output a video URL from the platform, which the server will then use to send a notification to the user's device.

[0159] Notification: "New technology stock info video added. Watch it here: [URL]"

[0160] Step 6: Watch the video

[0161] The user receives a notification on their device and opens an application (e.g., a messaging app) to watch the video. Specifically, the user taps the notification, which opens the application and displays the video URL. Clicking on this URL starts watching the video.

[0162] This series of steps realizes a system that can automatically generate and efficiently deliver video content optimized for each user.

[0163] (Application example 1)

[0164] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0165] With conventional content distribution systems, it was difficult to automatically generate and distribute video content optimized for user interests. Furthermore, content individualization was insufficient, limiting improvements in user satisfaction. Furthermore, notifications of generated video content were not provided in a timely manner, resulting in users missing viewing opportunities.

[0166] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0167] In this invention, the server includes means for acquiring a user's area of ​​interest from the user's past behavioral data and registration information, means for generating a script appropriate for each user using a generative AI based on the acquired user's area of ​​interest, means for collecting or generating related image materials based on the generated script, means for combining the script and image materials and adding narration using a text-to-speech program to generate a video, means for distributing the generated video to a platform and notifying the user's device, and means for notifying the user that the generated video is available. This makes it possible to automatically generate and effectively distribute video content optimized for user interests.

[0168] "User past behavioral data" refers to records of a user's browsing history, search history, click history, etc. on online platforms.

[0169] "Registration Information" refers to personal information, areas of interest, settings information, etc. provided by a user when registering for the service.

[0170] "Interest areas" refer to categories, themes, or topics that a user is particularly interested in.

[0171] "Generative AI" refers to an artificial intelligence model that automatically generates new content and information based on acquired data.

[0172] "Script" refers to a screenplay written in narration or text format created by generative AI.

[0173] "Image material" refers to visual materials that complement the generated script, including photographs, illustrations, graphs, etc.

[0174] A "reading program" refers to software that converts a text script into audio and reads it aloud as narration.

[0175] "Video" refers to visual and audio content created by combining a generated script with image material and narration.

[0176] "Platform" refers to online services and applications for delivering generated videos to users.

[0177] "Device" refers to the electronic device, such as a smartphone or tablet, that a User uses to receive and watch videos.

[0178] "Notification" means a message or alert that notifies the user that the generated video is available for viewing.

[0179] The present invention provides a system for automatically generating and distributing individually optimized video content based on a user's past behavioral data and registration information. Hereinafter, an embodiment of the present invention will be described in detail.

[0180] First, the server obtains the user's past behavioral data and registration information, including the articles the user has viewed on the online platform, search history, categories of interest, etc. Based on this information, the server identifies the user's areas of interest.

[0181] The server then uses a generative AI model (e.g., GPT-4) to generate a script based on the identified areas of interest. This script is tailored to the user's interests and may include specific topics such as "latest technology stock market trends." Examples of prompts include:

[0182] "Users are interested in: Technology stock information

[0183] Generation Objective: Generate a script about the latest technology stock market trends.

[0184] Based on the generated script, the server collects or generates relevant image materials, such as photos related to the news article or graphs and charts showing market trends. Generative AI can also be used to generate new image materials if necessary.

[0185] The server then combines the script with the collected image material to generate a video, adding narration using a text-to-speech program that reads the script aloud. The resulting video combines specific visual content and narration that correspond to the user's interests.

[0186] The generated video is uploaded to the specified platform (e.g., a social media app), and the server notifies the user's device that the generated video is available, allowing the user to immediately know that new video content is available for viewing.

[0187] Users receive a notification on their device and can watch the generated video. The video contains content optimized for the user's interests, resulting in a high level of satisfaction.

[0188] The specific hardware and software used to implement the above process include servers, user devices (smartphones, tablets, etc.), generative AI models using Python (e.g., GPT-4), image processing libraries (e.g., Pillow), and video editing libraries (e.g., MoviePy). This enables the automatic generation and distribution of individually optimized video content.

[0189] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0190] Step 1:

[0191] The server retrieves the user's past behavioral data and registration information. This information includes the articles the user has viewed on the online platform, their search history, and categories of interest. The user ID is used as input. The retrieved data is used to identify the user's areas of interest. Specifically, the server retrieves user information from the database using a REST API.

[0192] Step 2:

[0193] The server generates a script using a generative AI model (e.g., GPT-4) based on the acquired user's areas of interest. In this process, the server uses the user's areas of interest as input and provides a prompt sentence to the natural language generation AI. Specifically, the server inputs the areas of interest in text format into the AI ​​model and obtains the generated script text. The output is a script customized for each user.

[0194] Step 3:

[0195] The server collects or generates related image materials based on the generated script. The generated script text is used as input. Specifically, it collects images through a web API or generates new images using generative AI. The output is image materials that correspond to the script content.

[0196] Step 4:

[0197] The server combines the script with the collected image material to generate a video. The script text and image material are used as input. Specifically, it converts the script content into audio using a text-to-speech program, synchronizes it with the image material, and generates a video using a video editing library (e.g., MoviePy). The output is a video file with narration.

[0198] Step 5:

[0199] The server distributes the generated video to the specified platform. The generated video file and information about the distribution platform are used as input. The specific operation is to upload the video file to the platform using the REST API. The output is the status of distribution completion.

[0200] Step 6:

[0201] The server notifies the user's device that the generated video is available. The inputs are the delivery completion status and the user's contact information. The specific operation is to send a notification to the user's smartphone using a push notification service. The output is the notification sending status.

[0202] Step 7:

[0203] The user receives the notification on their device and watches the generated video. The notification content and the URL of the video distribution destination are used as input. The specific operation is to open a video viewing application through the smartphone's notification system and play the video. The output is log data of the video playback status.

[0204] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0205] The present invention provides an information distribution system that combines a user's past behavioral data and registered information with an emotion engine that recognizes the user's emotions. Hereinafter, embodiments of the present invention will be described in detail.

[0206] Obtaining user information and sentiment

[0207] The server retrieves the user's past behavioral data and registration information, and further recognizes the user's emotional state using an emotion engine, including the articles the user has viewed, the portfolio stocks they have registered, their search history, categories of interest, and emotional data through facial recognition and voice analysis.

[0208] example:

[0209] Let's say User B is interested in "fitness-related information" and is feeling "good" today.

[0210] Script generation

[0211] The server uses generative AI to generate a script suited to each user based on the acquired user information and emotional data. By incorporating the emotional data, a more personalized script is created.

[0212] example:

[0213] For user B, a script about "effective fitness training methods" is generated in a positive tone that reflects the emotional data of "energetic."

[0214] Image material generation and acquisition

[0215] The server collects or generates the necessary image materials based on the generated script. These materials include visuals and graphs that demonstrate the training method. If necessary, new image materials can be generated using generative AI.

[0216] example:

[0217] Based on User B's script, images showing the steps of the training method and visuals to increase motivation are collected.

[0218] Video generation

[0219] The server then combines the generated script with the collected image materials to generate a video. It uses a text-to-speech program to convert the script into audio and synchronize it with the image materials. It also adjusts the tone and content of the narration based on the user's emotions.

[0220] example:

[0221] A video is created in an energetic tone that explains to User B a training method that is highly effective for fitness.

[0222] Video distribution

[0223] The server uploads the generated video file to the specified platform and prepares it for distribution. As soon as the video is published on the distribution platform, a notification is sent to the user's device.

[0224] example:

[0225] User B receives a notification on their device saying "New fitness video available."

[0226] User Viewing

[0227] Users receive a notification sent to their device to watch the video. Clicking on the notification opens the application and plays the generated video. Users can review the video content and find out the information they are interested in.

[0228] example:

[0229] User B clicks on the notification and watches a fitness video to get training information that suits their energetic mood.

[0230] Through this process, the system can effectively and personalizedly provide content that users are interested in, increasing user satisfaction. In addition, by combining it with an emotion engine, video content can be adapted to the user's current emotional state, resulting in deeper engagement.

[0231] The processing flow will be explained below.

[0232] Step 1:

[0233] The server retrieves the user's past behavioral data and registration information, which includes collecting data such as the articles the user has viewed, the portfolio holdings they have registered, their search history, and categories of interest.

[0234] Step 2:

[0235] The server uses an emotion engine to recognize the user's current emotional state, which includes analyzing the user's facial expressions, tone of voice, and input.

[0236] Step 3:

[0237] The server combines the acquired behavioral data and registration information with the emotional data to identify the user's areas of interest, including categories of interest, keywords, and the user's emotions.

[0238] Step 4:

[0239] The server uses generative AI to generate a script based on the identified areas of interest and emotional data, with content and tone that reflects the user's current emotional state.

[0240] Step 5:

[0241] The server then collects or generates relevant image material based on the generated script, including graphs and charts showing market trends, photos related to news articles, and even uses generative AI to create new image material that matches the emotion.

[0242] Step 6:

[0243] The server combines image material with the generated script to create a video, using a text-to-speech program to convert the script into audio and adjust the tone of the narration based on emotion.

[0244] Step 7:

[0245] The server uploads the generated video file to the specified platform and prepares it for distribution. As soon as the video is published on the distribution platform, a notification is sent to the user's device.

[0246] Step 8:

[0247] Users receive a notification sent to their device to watch the video. Clicking on the notification opens the application and plays the generated video. Users can review the video content and find out the information they are interested in.

[0248] Example 2

[0249] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0250] Conventional information delivery systems have performed personalization by taking into account a user's past behavioral data and registration information, but have not been able to provide content that reflects the user's emotional state. As a result, while they can provide information that matches a user's interests and concerns, it is difficult to provide optimal content that reflects the user's emotional state at the time, which limits the ability to improve user satisfaction and engagement. Another problem is that the time and effort required to generate videos hinders the rapid provision of information.

[0251] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for acquiring a user's past behavioral data and registration information, means for recognizing the user's emotional state using the acquired data, means for generating a script suitable for each user using a generative AI model based on the acquired user information and emotional data, means for collecting or generating related image materials based on the generated script, means for combining the script and image materials and adding narration using a text-to-speech program to generate a video, and means for uploading the generated video to a distribution platform and notifying the user's terminal. This enables rapid generation and distribution of personalized video content based on a user's past behavioral data, registration information, and emotional state.

[0252] "User's past behavioral data" refers to information such as the user's previous web browsing history, search history, browsing history, and registered portfolio stocks.

[0253] "Registration Information" refers to data such as personal attributes, interests, and preferences that a user provides to the system.

[0254] "Emotional state" is data that indicates a user's current emotions and mood, and is obtained using methods such as facial recognition and voice analysis.

[0255] "Means of acquisition" refers to the methods and technologies used to collect users' past behavioral data and registration information from databases, etc.

[0256] "Means for recognizing emotional state" refers to methods or technologies for determining a user's current emotions using facial recognition technology or voice analysis technology.

[0257] A "generative AI model" refers to a machine learning model that uses artificial intelligence to perform tasks such as natural language generation and image generation.

[0258] "Script" refers to the outline or script of content created for each user using a generative AI model.

[0259] "Related image material" refers to the visual content required based on the generated script, including collected images and images newly created by the generative AI.

[0260] "Read-out program" refers to a program for converting text information into speech.

[0261] "Means for generating video" refers to methods and technologies for creating video content by synthesizing a script, image materials, and narration audio.

[0262] A "distribution platform" refers to a web service or application that allows generated video content to be published and delivered to users.

[0263] "Means of notification" refers to the methods and technologies used to notify a user's device that new content is available for viewing.

[0264] The present invention is an information distribution system that combines a user's past behavioral data and registered information with an emotion engine that recognizes the user's emotional state. Details of this system are described below.

[0265] Obtaining user information and sentiment

[0266] The server uses a database management system (e.g., MySQL) to retrieve the user's past behavioral data and registration information. This information includes the articles the user has viewed, their search history, and registered portfolio stocks. The server also uses an emotion engine (e.g., Emotion API) to recognize the user's emotional state. This emotion data includes information obtained through facial recognition and voice analysis.

[0267] Examples:

[0268] If User B has viewed "fitness-related information" multiple times in the past, the server will acquire this historical data. Also, if the emotion engine recognizes User B as "healthy," the server will acquire this emotion data.

[0269] Script generation

[0270] The server generates a personalized script using a generative AI model (e.g., GPT-4) based on the acquired user information and emotion data. It creates a prompt sentence to input into the generative AI model, and the script is generated based on the content of that sentence.

[0271] Specific prompt examples:

[0272] "User B has viewed fitness-related articles multiple times in the past and is currently feeling energetic."

[0273] Examples:

[0274] For User B, a script about "effective fitness training methods" with a positive tone is generated, reflecting the emotional data of "energetic."

[0275] Image material generation and acquisition

[0276] The server collects or generates the necessary image materials based on the generated script. Specifically, it searches and collects the necessary visual materials using an image collection platform (e.g., Unsplash API), and in some cases generates new image materials using a generative AI model (e.g., DALL-E).

[0277] Examples:

[0278] Based on User B's script, images showing training methods and visuals to increase motivation are collected.

[0279] Video generation

[0280] The server combines the generated script with the collected image material to generate a video. It uses narration synthesis software (e.g., Amazon Polly) to convert the script into audio, and video editing software (e.g., Adobe Premiere Pro) to synchronize the audio. It also adjusts the tone and content of the narration based on the user's emotions.

[0281] Examples:

[0282] Videos are created that explain effective fitness training methods in an energetic tone.

[0283] Video distribution

[0284] The server uploads the resulting video file to a distribution platform (e.g., YouTube API) and sends a notification to the user's device. As soon as the video is published, the user's device is notified that a new video is available to watch.

[0285] Examples:

[0286] User B receives a notification on their device saying "New fitness video available."

[0287] User Viewing

[0288] Users receive a notification sent to their device to watch the video. Clicking on the notification launches the video streaming application and plays the generated video, allowing users to immediately consume content of interest.

[0289] Examples:

[0290] User B clicks on the notification and watches a fitness video to get training information that suits their energetic mood.

[0291] This system aims to improve user satisfaction and engagement by quickly generating and delivering personalized content based on users' interest data and emotional state.

[0292] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0293] Step 1: Get user information

[0294] Input: User's ID

[0295] Output: User's past behavior data and registration information

[0296] The server uses a database management system (e.g., MySQL) to retrieve past behavioral data and registration information using the user's ID as a key. The data includes browsing history, search history, registered portfolio stocks, etc.

[0297] Specific behavior:

[0298] sql

[0299] SELECT FROM user_data WHERE user_id = 'B';

[0300] By executing this query, user B's past behavioral data and registration information will be retrieved.

[0301] Step 2: Obtaining emotion data

[0302] Input: User's ID

[0303] Output: User's emotional state data

[0304] The server uses an emotion engine (e.g., Emotion API) to recognize the user's emotional state through facial recognition and voice analysis.

[0305] Specific behavior:

[0306] python

[0307] emotion_data = emotion_api.recognize_emotion(user_id='B')

[0308] This API call obtains User B's current emotional state (e.g., "energetic").

[0309] Step 3: Generate a prompt statement

[0310] Input: User's past behavior data, registration information, emotional state data

[0311] Output: prompt statement

[0312] The server creates prompt sentences to input into the generative AI model based on user information and emotional data.

[0313] Specific behavior:

[0314] python

[0315] prompt = f "User B has viewed fitness-related articles multiple times in the past and is currently feeling energetic."

[0316] This generates a personalized prompt for User B.

[0317] Step 4: Generate the script

[0318] Input: prompt statement

[0319] Output: Script

[0320] The server uses a generative AI model (e.g., GPT-4) to generate a script based on the prompt.

[0321] Specific behavior:

[0322] python

[0323] script = generate_ai_model.generate(prompt)

[0324] Running this generative AI model will generate a fitness-related script that is appropriate for User B.

[0325] Step 5: Identifying image material

[0326] Input: Script

[0327] Output: A list of required image materials

[0328] The server identifies the necessary image material from the generated script.

[0329] Specific behavior:

[0330] python

[0331] required_images = extract_images_from_script(script)

[0332] This identifies a list of image materials required for the script.

[0333] Step 6: Collecting and creating image materials

[0334] Input: List of image materials

[0335] Output: Image material

[0336] The server collects the necessary visual materials using an image collection platform (e.g., Unsplash API) and, in some cases, generates new image materials using a generative AI model (e.g., DALL-E).

[0337] Specific behavior:

[0338] python

[0339] images = unsplash_api.search_images(query=required_images)

[0340] generated_images = dalle_api.generate(query=required_images)

[0341] This allows the necessary image materials to be collected or generated.

[0342] Step 7: Generate narration

[0343] Input: Script

[0344] Output: Narration audio

[0345] The server uses narration synthesis software (e.g., Amazon Polly) to convert the contents of the script into voice.

[0346] Specific behavior:

[0347] python

[0348] narration = polly_synthesize_speech(text=script)

[0349] This generates a narration voice based on the script.

[0350] Step 8: Edit your video

[0351] Input: script, image materials, narration audio

[0352] Output: Finished video

[0353] The server uses video editing software (e.g., Adobe Premiere Pro API) to generate a video that synchronizes the script, image materials, and narration audio.

[0354] Specific behavior:

[0355] python

[0356] video = create_video(narration=narration, images=images, script=script)

[0357] This process generates a video optimized for user B.

[0358] Step 9: Publish your video

[0359] Input: Finished video

[0360] Output: Video streaming link

[0361] The server uploads the generated video file to a distribution platform (e.g., YouTube API) and obtains a distribution link for the video.

[0362] Specific behavior:

[0363] python

[0364] upload_response = youtube_api.upload_video(file_path=video_path)

[0365] This will generate a distribution link for the video.

[0366] Step 10: Sending notifications

[0367] Input: Distribution link, user information

[0368] Output: Notification sent to the user's device

[0369] The server notifies the user's device that a new video is available for viewing.

[0370] Specific behavior:

[0371] python

[0372] send_notification(user_id='B', message='New fitness video available')

[0373] This will send a notification to User B's device.

[0374] Step 11: Watch the video

[0375] Input: Notification

[0376] Output: Video playback

[0377] The user receives a notification sent to their device and can watch the video. Clicking on the notification launches the video streaming application on their device and plays the generated video.

[0378] Specific behavior:

[0379] Click the notification on your device and open the video streaming application.

[0380] Play the new video to find out more.

[0381] This allows user B to view personalized video content.

[0382] (Application example 2)

[0383] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0384] Conventional information distribution systems were able to provide content based on users' past behavioral data and registration information, but it was difficult to provide personalized content that adapted to the user's emotional state. It is necessary to further improve user satisfaction and increase engagement by providing information based on the user's current emotions.

[0385] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0386] In this invention, the server includes: means for acquiring a user's areas of interest from the user's past behavioral data and registration information; means for generating a script appropriate for each user using a generative AI model based on the acquired user's areas of interest; means for collecting or generating related image materials based on the generated script; means for combining the script and image materials and adding narration using a text-to-speech program to generate a video; means for distributing the generated video to a platform and notifying the user's device; means for acquiring user emotional data through facial recognition and voice analysis; and means for adjusting the tone and content of the script based on the acquired emotional data, thereby enabling information distribution that is adapted to the user's current emotional state.

[0387] "User's past behavioral data and registration information" refers to data such as the user's previous actions and choices, browsing history, and personal information and areas of interest provided by the user when registering.

[0388] "Interest areas" are data that indicate specific categories, topics, or fields in which a user is interested.

[0389] A "generative AI model" is an artificial intelligence mechanism or program that generates specific content based on a user's areas of interest and emotional data.

[0390] A "script" is text data generated by a generative AI model to explain the content of an audio narration or video.

[0391] "Image materials" are images, graphs, and visual data that visually support the contents of the script.

[0392] A "reading program" is software that converts the contents of a script into audio and adds it to a video as narration.

[0393] A "platform" is an internet service or application for distributing video content.

[0394] "Facial recognition" is an image processing technology used to analyze emotions from a user's face.

[0395] "Voice analysis" is an acoustic processing technology used to analyze emotions and intentions from a user's speech.

[0396] "Emotional data" is data that represents a user's emotional state, obtained through facial recognition or voice analysis.

[0397] The present invention is an information distribution system that provides individually personalized information based on a user's past behavioral data, registration information, and emotional data. Specific embodiments will be described below.

[0398] Obtaining user information and sentiment

[0399] The server acquires the user's past behavioral data and registration information, and analyzes the user's emotional data using an emotion recognition engine. The specific software used for this is OpenCV and an emotion recognition model (e.g., emotion_model.onnx). The server acquires the user's emotions through facial recognition, and analyzes emotions from the user's speech through voice analysis.

[0400] example:

[0401] Let's say User B is interested in "fitness-related information" and is feeling "good" today.

[0402] Script generation

[0403] The server uses a generative AI model based on the acquired user information and emotional data to generate a script appropriate for the user. By incorporating emotional data, a more personalized script can be created. One example of the AI ​​model used is GPT-3.

[0404] Example prompt for a generative AI model:

[0405] "Generate a cheerful tone script about effective fitness training methods."

[0406] Image material generation and acquisition

[0407] The server collects or generates the necessary image materials based on the generated script. These materials include visuals that demonstrate training methods and visuals to increase motivation. If necessary, new image materials may be generated using generative AI.

[0408] example:

[0409] Based on User B's script, images showing the steps of the training method and visuals to increase motivation are collected.

[0410] Video generation

[0411] The server then combines the generated script with the collected image materials to generate a video. A text-to-speech program converts the script into audio and synchronizes it with the image materials. This program uses tools such as Google Text-to-Speech (gTTS). The server also adjusts the tone and content of the narration based on the user's emotions.

[0412] example:

[0413] A video is created in an energetic tone that explains to User B a training method that is highly effective for fitness.

[0414] Video distribution

[0415] The server uploads the generated video file to the distribution platform and sends a notification to the user's device using the distribution platform's API.

[0416] example:

[0417] User B receives a notification on their device saying "New fitness video available."

[0418] User Viewing

[0419] Users receive a notification sent to their device and can watch the video. Clicking on the notification opens the application and plays the generated video. This allows users to efficiently obtain information that matches their emotional state.

[0420] example:

[0421] User B clicks on the notification and watches a fitness video to get training information that suits their energetic mood.

[0422] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0423] Step 1:

[0424] Obtaining user information and sentiment

[0425] The server obtains the user's past behavioral data and registration information from the API, and collects emotional data using facial recognition and voice analysis. For example, the user's browsing history, search history, and registration information are input, and an emotion recognition engine (OpenCV and emotion recognition model) outputs emotional tags such as "cheerful" or "sad" from facial images and voice. This integrates the user's behavioral data and emotional data.

[0426] Step 2:

[0427] Script generation

[0428] The server generates a script using a generative AI model (e.g., GPT-3) based on the acquired user information and emotional data. The input is the user's area of ​​interest (e.g., fitness) and emotional data (e.g., "energetic"), and the script text generated based on the prompt ("Generate a cheerful tone script about effective fitness training methods.") is output. This generates text that matches the user's emotions.

[0429] Step 3:

[0430] Image material generation and acquisition

[0431] The server collects or generates related image materials based on the generated script. The input is the script text (e.g., "Explanation of effective fitness training methods in a lively tone"), and the output is a corresponding image file (e.g., a diagram showing training steps). This may involve using an image collection service or image generation AI.

[0432] Step 4:

[0433] Video generation

[0434] The server combines the generated script with image materials to generate a video. It then uses a text-to-speech program to convert the script into audio and synchronize it with the image materials. The input is the script text and image files, and the output is a video file with narration. Google Text-to-Speech (gTTS) is used for speech synthesis.

[0435] Step 5:

[0436] Video distribution

[0437] The server uploads the generated video file to the distribution platform and sends a notification to the user's device. The input is the video file and user information, and the output is the video uploaded to the distribution platform and a notification message to the user (e.g., "A new fitness video is available to watch"). This allows the user to be aware of the existence of the video and watch it.

[0438] Step 6:

[0439] User Viewing

[0440] The user receives a notification sent to their device and watches the video. The input is the notification message and the device, and the output is the video being played and the user can watch the content. This allows the user to enjoy personalized content.

[0441] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0442] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0443] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0444] [Second embodiment]

[0445] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0446] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0447] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0448] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0449] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0450] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0451] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0452] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0453] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0454] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0455] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0456] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0457] The present invention provides a system for automatically generating and distributing video content optimized for each user based on the user's past behavioral data and registration information. Hereinafter, a specific embodiment of the present invention will be described.

[0458] Retrieving User Information

[0459] The server obtains the user's past behavioral data and registration information, including the articles the user has viewed on the online platform, the portfolio stocks registered, search history, and categories of interest.

[0460] Script generation

[0461] Based on the acquired user information, the server uses generative AI to generate a script tailored to each user, specifically generating news, market trends, product descriptions, and other information related to the user's areas of interest in natural language.

[0462] example:

[0463] If User A is interested in "technology-related stock information," the server generates scripts about "today's technology stock market trends" and "earnings reports for specific companies."

[0464] Image material generation and acquisition

[0465] Based on the generated script, the server collects or generates the necessary image materials, such as graphs and charts showing market trends, photos related to news articles, etc. It can also use generative AI to generate new image materials as needed.

[0466] example:

[0467] If the script for user A is "about the rise in stock prices of ABC Company," the logo of ABC Company and a graph showing fluctuations in stock prices will be collected as image materials.

[0468] Video generation

[0469] The server combines the generated script with the collected image materials to generate a video. A text-to-speech program converts the script into audio and synchronizes it with the image materials.

[0470] example:

[0471] A script is generated saying "ABC Company's stock price has risen by 5%" and a video with narration is generated along with the corresponding stock price graph.

[0472] Video distribution

[0473] The server uploads the generated video to a designated platform (e.g., a social media app) and distributes it to users. The notification function notifies users of the availability of new content on their devices.

[0474] example:

[0475] User A receives a notification on his device saying, "A new video about technology stocks has been added." User A receives the notification and can watch the video.

[0476] User Viewing

[0477] Users receive notifications and can watch videos on their devices. Content tailored to their interests and preferences is available, resulting in high levels of satisfaction.

[0478] example:

[0479] User A opens the LINE app and watches a video streamed from the server. The video includes information on "technology stock market trends" and "details on the rise in ABC Company's stock price."

[0480] Through this series of processes, the system effectively provides content that users are interested in, increasing user satisfaction. In addition, by automatically generating and delivering video content customized for each user, the way information is received becomes more personalized.

[0481] The processing flow will be explained below.

[0482] Step 1:

[0483] The server retrieves the user's past behavioral data and registration information, which includes collecting data such as the articles the user has viewed, the portfolio holdings they have registered, their search history, and categories of interest.

[0484] Step 2:

[0485] The server analyzes the acquired data to identify the user's areas of interest, including the categories the user frequently browses and extracting related keywords.

[0486] Step 3:

[0487] The server uses generative AI to generate a script tailored to each user, which generates news, market trends, product descriptions, and other information in natural language based on the user's areas of interest.

[0488] Step 4:

[0489] Based on the generated script, the server collects or generates the necessary image materials, such as graphs and charts showing market trends and photos related to the news article, and, if necessary, uses generative AI to create new image materials.

[0490] Step 5:

[0491] The server combines the image material with the generated script to create a video, using a text-to-speech program to convert the script into audio and synchronize it with the image material.

[0492] Step 6:

[0493] The server uploads the generated video file to the specified platform and prepares it for distribution. As soon as the video is published on the distribution platform, a notification is sent to the user's device.

[0494] Step 7:

[0495] Users receive a notification sent to their device to watch the video. Clicking on the notification opens the application and plays the generated video. Users can review the video content and find out the information they are interested in.

[0496] Example 1

[0497] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0498] Conventional information distribution systems have struggled to automatically generate and efficiently distribute personalized content for each user. As a result, it has been difficult to provide information tailored to the user's interests, often resulting in low user satisfaction. Furthermore, manual content creation and distribution is time-consuming and costly, so efficient operation is required.

[0499] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0500] In this invention, the server includes means for acquiring a user's area of ​​interest from the user's past behavioral data and registration information, means for generating a script suitable for each user using a generative AI based on the acquired user's area of ​​interest, means for collecting or generating related image materials based on the generated script, means for combining the script and image materials and generating audio and video using a text-to-speech program, means for distributing the generated video to a specific platform and notifying the user's device, and means for viewing the notified content on the user's device. This makes it possible to automatically generate and distribute video content personalized for each user, thereby achieving high user satisfaction.

[0501] "User past behavioral data" refers to data such as the actions a user takes on an online platform, browsing history, search history, and purchase history.

[0502] "Registration Information" means the personal information and interest category information provided by a User when registering for the Service.

[0503] "User interest areas" are the range of categories or topics in which a user is interested, identified based on the user's behavioral data and registration information.

[0504] "Generative AI" is an artificial intelligence model that generates natural language text, images, and other content based on data and prompts provided.

[0505] A "script" is text containing narration and content instructions for a video, generated based on information related to the user's area of ​​interest.

[0506] "Image material" refers to visual content used in a video, such as photographs, graphs, charts, logos, etc.

[0507] A "read-aloud program" is software that converts text into synthetic speech and plays it back as narration.

[0508] "Video" means dynamic visual content consisting of multiple frames and may include audio and text.

[0509] "Platform" is a general term for websites and applications that allow users to upload generated videos and make them viewable.

[0510] "Notifications" are messages that inform users when newly generated videos or content is available.

[0511] "Device" means a device on which a user receives notifications and views content, including a smartphone, tablet, or PC.

[0512] The present invention provides an information distribution system that automatically generates and distributes video content optimized for each user. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS The following describes in detail an embodiment of the present invention.

[0513] Retrieving User Information

[0514] The server retrieves the user's past behavioral data and registration information from the database, including browsing history, registered portfolio stocks, search history, category information of interest, etc. When the server retrieves this data, it uses the user ID to efficiently collect related data.

[0515] As a concrete example, the server issues an SQL query to retrieve data as follows:

[0516] SELECT FROM user_behavior WHERE user_id = 'userA';

[0517] Script generation

[0518] The server uses generative AI (e.g., GPT-4) to generate a script tailored to each user based on the acquired user information, and sends prompts to generate news, market trends, and product descriptions related to the user's areas of interest in natural language.

[0519] As a concrete example, if user A is interested in technology stock information, the server might use a prompt like this:

[0520] User A is interested in technology stocks. Create a news script for this user.

[0521] In response to this prompt, the generative AI outputs a script that reads, "Today's technology market trends: ABC Company's stock price rose 5%."

[0522] Image material generation and acquisition

[0523] The server collects or generates the necessary image materials based on the generated script, such as photos, graphs, charts, etc., using external data sources or generative AI (e.g., DALL-E).

[0524] You can use the API to obtain image materials. For example, issue the following API request to obtain image materials.

[0525] GET / stock_images?company=ABC

[0526] Video generation

[0527] The server combines the generated script with image materials and generates a video using automatic video generation software (e.g., Adobe Premiere Pro API), and also converts the script into audio using a speech synthesis API (e.g., Google Text-to-Speech) and embeds it in the video.

[0528] As a concrete example, an API request is issued as follows to generate a video.

[0529] POST / generate_video

[0530] {

[0531] "script": "Today's technology market trends: ABC Company's stock price rose 5%.",

[0532] "images": ["abc_logo.png", "stock_chart.png"],

[0533] "voice": "synthesized voice file"

[0534] }

[0535] Video distribution

[0536] The server uploads the generated video to a specific platform (e.g., YouTube) and notifies the user's device. The server then uses the platform's API to upload the video and obtain its URL.

[0537] As a concrete example, you can upload a video by issuing an API request as follows:

[0538] POST / upload_video

[0539] {

[0540] "video_file": "generated_video.mp4",

[0541] "platform": "YouTube"

[0542] }

[0543] After uploading, a notification message will be sent to the user's device.

[0544] Notification: "New technology stock info video added. Watch it here: [URL]"

[0545] User Viewing

[0546] The user receives a notification on their device, opens an application (e.g., a messaging app) to watch the video, and can tap the notification to go directly to the viewing page.

[0547] For example, when a user taps on a notification on their device, a messaging app opens with the URL of the specified video, which the user can click to watch the video.

[0548] Through these steps, the present invention realizes the automatic generation and distribution of video content optimized for each user, and builds an information distribution system that provides high levels of user satisfaction.

[0549] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0550] Step 1: Get user information

[0551] The server retrieves the user's past behavioral data and registration information from the database. As input, the user ID is provided to the server, and the server executes a query against the database based on this user ID. Specifically, the server issues an SQL query as follows:

[0552] SELECT FROM user_behavior WHERE user_id = 'userA';

[0553] This query outputs data related to the user's areas of interest (browsing history, registered portfolio stocks, search history, interest categories, etc.).

[0554] Step 2: Generate the script

[0555] The server uses the acquired user information as input and sends a prompt to the generative AI (e.g., GPT-4) to generate a script. Specifically, the server sends the following prompt to the generative AI:

[0556] User A is interested in technology stocks. Create a news script for this user.

[0557] Based on this prompt, the generative AI outputs text related to the user's area of ​​interest (e.g., "Today's technology market trends: ABC Company's stock price rose 5%.").

[0558] Step 3: Creating and acquiring image materials

[0559] The server uses the generated script as input to collect or generate the necessary image materials. Specifically, the server issues an API request to an external data source.

[0560] GET / stock_images?company=ABC

[0561] This request will output relevant image material (e.g., ABC company logo, stock price graph).

[0562] Step 4: Generate the video

[0563] The server generates a video using automatic video generation software (e.g., Adobe Premiere Pro API) using the generated script and collected image materials as input. Specifically, the server issues the following API request:

[0564] POST / generate_video

[0565] {

[0566] "script": "Today's technology market trends: ABC Company's stock price rose 5%.",

[0567] "images": ["abc_logo.png", "stock_chart.png"],

[0568] "voice": "synthesized voice file"

[0569] }

[0570] This request causes the generated video to be output.

[0571] Step 5: Publish your video

[0572] The server takes the generated video as input, uploads it to a specific platform (e.g., YouTube), and notifies the user's device. Specifically, the server issues the following API request:

[0573] POST / upload_video

[0574] {

[0575] "video_file": "generated_video.mp4",

[0576] "platform": "YouTube"

[0577] }

[0578] This action will output a video URL from the platform, which the server will then use to send a notification to the user's device.

[0579] Notification: "New technology stock info video added. Watch it here: [URL]"

[0580] Step 6: Watch the video

[0581] The user receives a notification on their device and opens an application (e.g., a messaging app) to watch the video. Specifically, the user taps the notification, which opens the application and displays the video URL. Clicking on this URL starts watching the video.

[0582] This series of steps realizes a system that can automatically generate and efficiently deliver video content optimized for each user.

[0583] (Application example 1)

[0584] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0585] With conventional content distribution systems, it was difficult to automatically generate and distribute video content optimized for user interests. Furthermore, content individualization was insufficient, limiting improvements in user satisfaction. Furthermore, notifications of generated video content were not provided in a timely manner, resulting in users missing viewing opportunities.

[0586] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0587] In this invention, the server includes means for acquiring a user's area of ​​interest from the user's past behavioral data and registration information, means for generating a script appropriate for each user using a generative AI based on the acquired user's area of ​​interest, means for collecting or generating related image materials based on the generated script, means for combining the script and image materials and adding narration using a text-to-speech program to generate a video, means for distributing the generated video to a platform and notifying the user's device, and means for notifying the user that the generated video is available. This makes it possible to automatically generate and effectively distribute video content optimized for user interests.

[0588] "User past behavioral data" refers to records of a user's browsing history, search history, click history, etc. on online platforms.

[0589] "Registration Information" refers to personal information, areas of interest, settings information, etc. provided by a user when registering for the service.

[0590] "Interest areas" refer to categories, themes, or topics that a user is particularly interested in.

[0591] "Generative AI" refers to an artificial intelligence model that automatically generates new content and information based on acquired data.

[0592] "Script" refers to a screenplay written in narration or text format created by generative AI.

[0593] "Image material" refers to visual materials that complement the generated script, including photographs, illustrations, graphs, etc.

[0594] A "reading program" refers to software that converts a text script into audio and reads it aloud as narration.

[0595] "Video" refers to visual and audio content created by combining a generated script with image material and narration.

[0596] "Platform" refers to online services and applications for delivering generated videos to users.

[0597] "Device" refers to the electronic device, such as a smartphone or tablet, that a User uses to receive and watch videos.

[0598] "Notification" means a message or alert that notifies the user that the generated video is available for viewing.

[0599] The present invention provides a system for automatically generating and distributing individually optimized video content based on a user's past behavioral data and registration information. Hereinafter, an embodiment of the present invention will be described in detail.

[0600] First, the server obtains the user's past behavioral data and registration information, including the articles the user has viewed on the online platform, search history, categories of interest, etc. Based on this information, the server identifies the user's areas of interest.

[0601] The server then uses a generative AI model (e.g., GPT-4) to generate a script based on the identified areas of interest. This script is tailored to the user's interests and may include specific topics such as "latest technology stock market trends." Examples of prompts include:

[0602] "Users are interested in: Technology stock information

[0603] Generation Objective: Generate a script about the latest technology stock market trends.

[0604] Based on the generated script, the server collects or generates relevant image materials, such as photos related to the news article or graphs and charts showing market trends. Generative AI can also be used to generate new image materials if necessary.

[0605] The server then combines the script with the collected image material to generate a video, adding narration using a text-to-speech program that reads the script aloud. The resulting video combines specific visual content and narration that correspond to the user's interests.

[0606] The generated video is uploaded to the specified platform (e.g., a social media app), and the server notifies the user's device that the generated video is available, allowing the user to immediately know that new video content is available for viewing.

[0607] Users receive a notification on their device and can watch the generated video. The video contains content optimized for the user's interests, resulting in a high level of satisfaction.

[0608] The specific hardware and software used to implement the above process include servers, user devices (smartphones, tablets, etc.), generative AI models using Python (e.g., GPT-4), image processing libraries (e.g., Pillow), and video editing libraries (e.g., MoviePy). This enables the automatic generation and distribution of individually optimized video content.

[0609] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0610] Step 1:

[0611] The server retrieves the user's past behavioral data and registration information. This information includes the articles the user has viewed on the online platform, their search history, and categories of interest. The user ID is used as input. The retrieved data is used to identify the user's areas of interest. Specifically, the server retrieves user information from the database using a REST API.

[0612] Step 2:

[0613] The server generates a script using a generative AI model (e.g., GPT-4) based on the acquired user's areas of interest. In this process, the server uses the user's areas of interest as input and provides a prompt sentence to the natural language generation AI. Specifically, the server inputs the areas of interest in text format into the AI ​​model and obtains the generated script text. The output is a script customized for each user.

[0614] Step 3:

[0615] The server collects or generates related image materials based on the generated script. The generated script text is used as input. Specifically, it collects images through a web API or generates new images using generative AI. The output is image materials that correspond to the script content.

[0616] Step 4:

[0617] The server combines the script with the collected image material to generate a video. The script text and image material are used as input. Specifically, it converts the script content into audio using a text-to-speech program, synchronizes it with the image material, and generates a video using a video editing library (e.g., MoviePy). The output is a video file with narration.

[0618] Step 5:

[0619] The server distributes the generated video to the specified platform. The generated video file and information about the distribution platform are used as input. The specific operation is to upload the video file to the platform using the REST API. The output is the status of distribution completion.

[0620] Step 6:

[0621] The server notifies the user's device that the generated video is available. The inputs are the delivery completion status and the user's contact information. The specific operation is to send a notification to the user's smartphone using a push notification service. The output is the notification sending status.

[0622] Step 7:

[0623] The user receives the notification on their device and watches the generated video. The notification content and the URL of the video distribution destination are used as input. The specific operation is to open a video viewing application through the smartphone's notification system and play the video. The output is log data of the video playback status.

[0624] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0625] The present invention provides an information distribution system that combines a user's past behavioral data and registered information with an emotion engine that recognizes the user's emotions. Hereinafter, embodiments of the present invention will be described in detail.

[0626] Obtaining user information and sentiment

[0627] The server retrieves the user's past behavioral data and registration information, and further recognizes the user's emotional state using an emotion engine, including the articles the user has viewed, the portfolio stocks they have registered, their search history, categories of interest, and emotional data through facial recognition and voice analysis.

[0628] example:

[0629] Let's say User B is interested in "fitness-related information" and is feeling "good" today.

[0630] Script generation

[0631] The server uses generative AI to generate a script suited to each user based on the acquired user information and emotional data. By incorporating the emotional data, a more personalized script is created.

[0632] example:

[0633] For user B, a script about "effective fitness training methods" is generated in a positive tone that reflects the emotional data of "energetic."

[0634] Image material generation and acquisition

[0635] The server collects or generates the necessary image materials based on the generated script. These materials include visuals and graphs that demonstrate the training method. If necessary, new image materials can be generated using generative AI.

[0636] example:

[0637] Based on User B's script, images showing the steps of the training method and visuals to increase motivation are collected.

[0638] Video generation

[0639] The server then combines the generated script with the collected image materials to generate a video. It uses a text-to-speech program to convert the script into audio and synchronize it with the image materials. It also adjusts the tone and content of the narration based on the user's emotions.

[0640] example:

[0641] A video is created in an energetic tone that explains to User B a training method that is highly effective for fitness.

[0642] Video distribution

[0643] The server uploads the generated video file to the specified platform and prepares it for distribution. As soon as the video is published on the distribution platform, a notification is sent to the user's device.

[0644] example:

[0645] User B receives a notification on their device saying "New fitness video available."

[0646] User Viewing

[0647] Users receive a notification sent to their device to watch the video. Clicking on the notification opens the application and plays the generated video. Users can review the video content and find out the information they are interested in.

[0648] example:

[0649] User B clicks on the notification and watches a fitness video to get training information that suits their energetic mood.

[0650] Through this process, the system can effectively and personalizedly provide content that users are interested in, increasing user satisfaction. In addition, by combining it with an emotion engine, video content can be adapted to the user's current emotional state, resulting in deeper engagement.

[0651] The processing flow will be explained below.

[0652] Step 1:

[0653] The server retrieves the user's past behavioral data and registration information, which includes collecting data such as the articles the user has viewed, the portfolio holdings they have registered, their search history, and categories of interest.

[0654] Step 2:

[0655] The server uses an emotion engine to recognize the user's current emotional state, which includes analyzing the user's facial expressions, tone of voice, and input.

[0656] Step 3:

[0657] The server combines the acquired behavioral data and registration information with the emotional data to identify the user's areas of interest, including categories of interest, keywords, and the user's emotions.

[0658] Step 4:

[0659] The server uses generative AI to generate a script based on the identified areas of interest and emotional data, with content and tone that reflects the user's current emotional state.

[0660] Step 5:

[0661] The server then collects or generates relevant image material based on the generated script, including graphs and charts showing market trends, photos related to news articles, and even uses generative AI to create new image material that matches the emotion.

[0662] Step 6:

[0663] The server combines image material with the generated script to create a video, using a text-to-speech program to convert the script into audio and adjust the tone of the narration based on emotion.

[0664] Step 7:

[0665] The server uploads the generated video file to the specified platform and prepares it for distribution. As soon as the video is published on the distribution platform, a notification is sent to the user's device.

[0666] Step 8:

[0667] Users receive a notification sent to their device to watch the video. Clicking on the notification opens the application and plays the generated video. Users can review the video content and find out the information they are interested in.

[0668] Example 2

[0669] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0670] Conventional information delivery systems have performed personalization by taking into account a user's past behavioral data and registration information, but have not been able to provide content that reflects the user's emotional state. As a result, while they can provide information that matches a user's interests and concerns, it is difficult to provide optimal content that reflects the user's emotional state at the time, which limits the ability to improve user satisfaction and engagement. Another problem is that the time and effort required to generate videos hinders the rapid provision of information.

[0671] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for acquiring a user's past behavioral data and registration information, means for recognizing the user's emotional state using the acquired data, means for generating a script suitable for each user using a generative AI model based on the acquired user information and emotional data, means for collecting or generating related image materials based on the generated script, means for combining the script and image materials and adding narration using a text-to-speech program to generate a video, and means for uploading the generated video to a distribution platform and notifying the user's terminal. This enables rapid generation and distribution of personalized video content based on a user's past behavioral data, registration information, and emotional state.

[0672] "User's past behavioral data" refers to information such as the user's previous web browsing history, search history, browsing history, and registered portfolio stocks.

[0673] "Registration Information" refers to data such as personal attributes, interests, and preferences that a user provides to the system.

[0674] "Emotional state" is data that indicates a user's current emotions and mood, and is obtained using methods such as facial recognition and voice analysis.

[0675] "Means of acquisition" refers to the methods and technologies used to collect users' past behavioral data and registration information from databases, etc.

[0676] "Means for recognizing emotional state" refers to methods or technologies for determining a user's current emotions using facial recognition technology or voice analysis technology.

[0677] A "generative AI model" refers to a machine learning model that uses artificial intelligence to perform tasks such as natural language generation and image generation.

[0678] "Script" refers to the outline or script of content created for each user using a generative AI model.

[0679] "Related image material" refers to the visual content required based on the generated script, including collected images and images newly created by the generative AI.

[0680] "Read-out program" refers to a program for converting text information into speech.

[0681] "Means for generating video" refers to methods and technologies for creating video content by synthesizing a script, image materials, and narration audio.

[0682] A "distribution platform" refers to a web service or application that allows generated video content to be published and delivered to users.

[0683] "Means of notification" refers to the methods and technologies used to notify a user's device that new content is available for viewing.

[0684] The present invention is an information distribution system that combines a user's past behavioral data and registered information with an emotion engine that recognizes the user's emotional state. Details of this system are described below.

[0685] Obtaining user information and sentiment

[0686] The server uses a database management system (e.g., MySQL) to retrieve the user's past behavioral data and registration information. This information includes the articles the user has viewed, their search history, and registered portfolio stocks. The server also uses an emotion engine (e.g., Emotion API) to recognize the user's emotional state. This emotion data includes information obtained through facial recognition and voice analysis.

[0687] Examples:

[0688] If User B has viewed "fitness-related information" multiple times in the past, the server will acquire this historical data. Also, if the emotion engine recognizes User B as "healthy," the server will acquire this emotion data.

[0689] Script generation

[0690] The server generates a personalized script using a generative AI model (e.g., GPT-4) based on the acquired user information and emotion data. It creates a prompt sentence to input into the generative AI model, and the script is generated based on the content of that sentence.

[0691] Specific prompt examples:

[0692] "User B has viewed fitness-related articles multiple times in the past and is currently feeling energetic."

[0693] Examples:

[0694] For User B, a script about "effective fitness training methods" with a positive tone is generated, reflecting the emotional data of "energetic."

[0695] Image material generation and acquisition

[0696] The server collects or generates the necessary image materials based on the generated script. Specifically, it searches and collects the necessary visual materials using an image collection platform (e.g., Unsplash API), and in some cases generates new image materials using a generative AI model (e.g., DALL-E).

[0697] Examples:

[0698] Based on User B's script, images showing training methods and visuals to increase motivation are collected.

[0699] Video generation

[0700] The server combines the generated script with the collected image material to generate a video. It uses narration synthesis software (e.g., Amazon Polly) to convert the script into audio, and video editing software (e.g., Adobe Premiere Pro) to synchronize the audio. It also adjusts the tone and content of the narration based on the user's emotions.

[0701] Examples:

[0702] Videos are created that explain effective fitness training methods in an energetic tone.

[0703] Video distribution

[0704] The server uploads the resulting video file to a distribution platform (e.g., YouTube API) and sends a notification to the user's device. As soon as the video is published, the user's device is notified that a new video is available to watch.

[0705] Examples:

[0706] User B receives a notification on their device saying "New fitness video available."

[0707] User Viewing

[0708] Users receive a notification sent to their device to watch the video. Clicking on the notification launches the video streaming application and plays the generated video, allowing users to immediately consume content of interest.

[0709] Examples:

[0710] User B clicks on the notification and watches a fitness video to get training information that suits their energetic mood.

[0711] This system aims to improve user satisfaction and engagement by quickly generating and delivering personalized content based on users' interest data and emotional state.

[0712] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0713] Step 1: Get user information

[0714] Input: User's ID

[0715] Output: User's past behavior data and registration information

[0716] The server uses a database management system (e.g., MySQL) to retrieve past behavioral data and registration information using the user's ID as a key. The data includes browsing history, search history, registered portfolio stocks, etc.

[0717] Specific behavior:

[0718] sql

[0719] SELECT FROM user_data WHERE user_id = 'B';

[0720] By executing this query, user B's past behavioral data and registration information will be retrieved.

[0721] Step 2: Obtaining emotion data

[0722] Input: User's ID

[0723] Output: User's emotional state data

[0724] The server uses an emotion engine (e.g., Emotion API) to recognize the user's emotional state through facial recognition and voice analysis.

[0725] Specific behavior:

[0726] python

[0727] emotion_data = emotion_api.recognize_emotion(user_id='B')

[0728] This API call obtains User B's current emotional state (e.g., "energetic").

[0729] Step 3: Generate a prompt statement

[0730] Input: User's past behavior data, registration information, emotional state data

[0731] Output: prompt statement

[0732] The server creates prompt sentences to input into the generative AI model based on user information and emotional data.

[0733] Specific behavior:

[0734] python

[0735] prompt = f "User B has viewed fitness-related articles multiple times in the past and is currently feeling energetic."

[0736] This generates a personalized prompt for User B.

[0737] Step 4: Generate the script

[0738] Input: prompt statement

[0739] Output: Script

[0740] The server uses a generative AI model (e.g., GPT-4) to generate a script based on the prompt.

[0741] Specific behavior:

[0742] python

[0743] script = generate_ai_model.generate(prompt)

[0744] Running this generative AI model will generate a fitness-related script that is appropriate for User B.

[0745] Step 5: Identifying image material

[0746] Input: Script

[0747] Output: A list of required image materials

[0748] The server identifies the necessary image material from the generated script.

[0749] Specific behavior:

[0750] python

[0751] required_images = extract_images_from_script(script)

[0752] This identifies a list of image materials required for the script.

[0753] Step 6: Collecting and creating image materials

[0754] Input: List of image materials

[0755] Output: Image material

[0756] The server collects the necessary visual materials using an image collection platform (e.g., Unsplash API) and, in some cases, generates new image materials using a generative AI model (e.g., DALL-E).

[0757] Specific behavior:

[0758] python

[0759] images = unsplash_api.search_images(query=required_images)

[0760] generated_images = dalle_api.generate(query=required_images)

[0761] This allows the necessary image materials to be collected or generated.

[0762] Step 7: Generate narration

[0763] Input: Script

[0764] Output: Narration audio

[0765] The server uses narration synthesis software (e.g., Amazon Polly) to convert the contents of the script into voice.

[0766] Specific behavior:

[0767] python

[0768] narration = polly_synthesize_speech(text=script)

[0769] This generates a narration voice based on the script.

[0770] Step 8: Edit your video

[0771] Input: script, image materials, narration audio

[0772] Output: Finished video

[0773] The server uses video editing software (e.g., Adobe Premiere Pro API) to generate a video that synchronizes the script, image materials, and narration audio.

[0774] Specific behavior:

[0775] python

[0776] video = create_video(narration=narration, images=images, script=script)

[0777] This process generates a video optimized for user B.

[0778] Step 9: Publish your video

[0779] Input: Finished video

[0780] Output: Video streaming link

[0781] The server uploads the generated video file to a distribution platform (e.g., YouTube API) and obtains a distribution link for the video.

[0782] Specific behavior:

[0783] python

[0784] upload_response = youtube_api.upload_video(file_path=video_path)

[0785] This will generate a distribution link for the video.

[0786] Step 10: Sending notifications

[0787] Input: Distribution link, user information

[0788] Output: Notification sent to the user's device

[0789] The server notifies the user's device that a new video is available for viewing.

[0790] Specific behavior:

[0791] python

[0792] send_notification(user_id='B', message='New fitness video available')

[0793] This will send a notification to User B's device.

[0794] Step 11: Watch the video

[0795] Input: Notification

[0796] Output: Video playback

[0797] The user receives a notification sent to their device and can watch the video. Clicking on the notification launches the video streaming application on their device and plays the generated video.

[0798] Specific behavior:

[0799] Click the notification on your device and open the video streaming application.

[0800] Play the new video to find out more.

[0801] This allows user B to view personalized video content.

[0802] (Application example 2)

[0803] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0804] Conventional information distribution systems were able to provide content based on users' past behavioral data and registration information, but it was difficult to provide personalized content that adapted to the user's emotional state. It is necessary to further improve user satisfaction and increase engagement by providing information based on the user's current emotions.

[0805] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0806] In this invention, the server includes: means for acquiring a user's areas of interest from the user's past behavioral data and registration information; means for generating a script appropriate for each user using a generative AI model based on the acquired user's areas of interest; means for collecting or generating related image materials based on the generated script; means for combining the script and image materials and adding narration using a text-to-speech program to generate a video; means for distributing the generated video to a platform and notifying the user's device; means for acquiring user emotional data through facial recognition and voice analysis; and means for adjusting the tone and content of the script based on the acquired emotional data, thereby enabling information distribution that is adapted to the user's current emotional state.

[0807] "User's past behavioral data and registration information" refers to data such as the user's previous actions and choices, browsing history, and personal information and areas of interest provided by the user when registering.

[0808] "Interest areas" are data that indicate specific categories, topics, or fields in which a user is interested.

[0809] A "generative AI model" is an artificial intelligence mechanism or program that generates specific content based on a user's areas of interest and emotional data.

[0810] A "script" is text data generated by a generative AI model to explain the content of an audio narration or video.

[0811] "Image materials" are images, graphs, and visual data that visually support the contents of the script.

[0812] A "reading program" is software that converts the contents of a script into audio and adds it to a video as narration.

[0813] A "platform" is an internet service or application for distributing video content.

[0814] "Facial recognition" is an image processing technology used to analyze emotions from a user's face.

[0815] "Voice analysis" is an acoustic processing technology used to analyze emotions and intentions from a user's speech.

[0816] "Emotional data" is data that represents a user's emotional state, obtained through facial recognition or voice analysis.

[0817] The present invention is an information distribution system that provides individually personalized information based on a user's past behavioral data, registration information, and emotional data. Specific embodiments will be described below.

[0818] Obtaining user information and sentiment

[0819] The server acquires the user's past behavioral data and registration information, and analyzes the user's emotional data using an emotion recognition engine. The specific software used for this is OpenCV and an emotion recognition model (e.g., emotion_model.onnx). The server acquires the user's emotions through facial recognition, and analyzes emotions from the user's speech through voice analysis.

[0820] example:

[0821] Let's say User B is interested in "fitness-related information" and is feeling "good" today.

[0822] Script generation

[0823] The server uses a generative AI model based on the acquired user information and emotional data to generate a script appropriate for the user. By incorporating emotional data, a more personalized script can be created. One example of the AI ​​model used is GPT-3.

[0824] Example prompt for a generative AI model:

[0825] "Generate a cheerful tone script about effective fitness training methods."

[0826] Image material generation and acquisition

[0827] The server collects or generates the necessary image materials based on the generated script. These materials include visuals that demonstrate training methods and visuals to increase motivation. If necessary, new image materials may be generated using generative AI.

[0828] example:

[0829] Based on User B's script, images showing the steps of the training method and visuals to increase motivation are collected.

[0830] Video generation

[0831] The server then combines the generated script with the collected image materials to generate a video. A text-to-speech program converts the script into audio and synchronizes it with the image materials. This program uses tools such as Google Text-to-Speech (gTTS). The server also adjusts the tone and content of the narration based on the user's emotions.

[0832] example:

[0833] A video is created in an energetic tone that explains to User B a training method that is highly effective for fitness.

[0834] Video distribution

[0835] The server uploads the generated video file to the distribution platform and sends a notification to the user's device using the distribution platform's API.

[0836] example:

[0837] User B receives a notification on their device saying "New fitness video available."

[0838] User Viewing

[0839] Users receive a notification sent to their device and can watch the video. Clicking on the notification opens the application and plays the generated video. This allows users to efficiently obtain information that matches their emotional state.

[0840] example:

[0841] User B clicks on the notification and watches a fitness video to get training information that suits their energetic mood.

[0842] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0843] Step 1:

[0844] Obtaining user information and sentiment

[0845] The server obtains the user's past behavioral data and registration information from the API, and collects emotional data using facial recognition and voice analysis. For example, the user's browsing history, search history, and registration information are input, and an emotion recognition engine (OpenCV and emotion recognition model) outputs emotional tags such as "cheerful" or "sad" from facial images and voice. This integrates the user's behavioral data and emotional data.

[0846] Step 2:

[0847] Script generation

[0848] The server generates a script using a generative AI model (e.g., GPT-3) based on the acquired user information and emotional data. The input is the user's area of ​​interest (e.g., fitness) and emotional data (e.g., "energetic"), and the script text generated based on the prompt ("Generate a cheerful tone script about effective fitness training methods.") is output. This generates text that matches the user's emotions.

[0849] Step 3:

[0850] Image material generation and acquisition

[0851] The server collects or generates related image materials based on the generated script. The input is the script text (e.g., "Explanation of effective fitness training methods in a lively tone"), and the output is a corresponding image file (e.g., a diagram showing training steps). This may involve using an image collection service or image generation AI.

[0852] Step 4:

[0853] Video generation

[0854] The server combines the generated script with image materials to generate a video. It then uses a text-to-speech program to convert the script into audio and synchronize it with the image materials. The input is the script text and image files, and the output is a video file with narration. Google Text-to-Speech (gTTS) is used for speech synthesis.

[0855] Step 5:

[0856] Video distribution

[0857] The server uploads the generated video file to the distribution platform and sends a notification to the user's device. The input is the video file and user information, and the output is the video uploaded to the distribution platform and a notification message to the user (e.g., "A new fitness video is available to watch"). This allows the user to be aware of the existence of the video and watch it.

[0858] Step 6:

[0859] User Viewing

[0860] The user receives a notification sent to their device and watches the video. The input is the notification message and the device, and the output is the video being played and the user can watch the content. This allows the user to enjoy personalized content.

[0861] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0862] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0863] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0864] [Third embodiment]

[0865] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0866] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[0867] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0868] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0869] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0870] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0871] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0872] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0873] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0874] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0875] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0876] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0877] The present invention provides a system for automatically generating and distributing video content optimized for each user based on the user's past behavioral data and registration information. Hereinafter, a specific embodiment of the present invention will be described.

[0878] Retrieving User Information

[0879] The server obtains the user's past behavioral data and registration information, including the articles the user has viewed on the online platform, the portfolio stocks registered, search history, and categories of interest.

[0880] Script generation

[0881] Based on the acquired user information, the server uses generative AI to generate a script tailored to each user, specifically generating news, market trends, product descriptions, and other information related to the user's areas of interest in natural language.

[0882] example:

[0883] If User A is interested in "technology-related stock information," the server generates scripts about "today's technology stock market trends" and "earnings reports for specific companies."

[0884] Image material generation and acquisition

[0885] Based on the generated script, the server collects or generates the necessary image materials, such as graphs and charts showing market trends, photos related to news articles, etc. It can also use generative AI to generate new image materials as needed.

[0886] example:

[0887] If the script for user A is "about the rise in stock prices of ABC Company," the logo of ABC Company and a graph showing fluctuations in stock prices will be collected as image materials.

[0888] Video generation

[0889] The server combines the generated script with the collected image materials to generate a video. A text-to-speech program converts the script into audio and synchronizes it with the image materials.

[0890] example:

[0891] A script is generated saying "ABC Company's stock price has risen by 5%" and a video with narration is generated along with the corresponding stock price graph.

[0892] Video distribution

[0893] The server uploads the generated video to a designated platform (e.g., a social media app) and distributes it to users. The notification function notifies users of the availability of new content on their devices.

[0894] example:

[0895] User A receives a notification on his device saying, "A new video about technology stocks has been added." User A receives the notification and can watch the video.

[0896] User Viewing

[0897] Users receive notifications and can watch videos on their devices. Content tailored to their interests and preferences is available, resulting in high levels of satisfaction.

[0898] example:

[0899] User A opens the LINE app and watches a video streamed from the server. The video includes information on "technology stock market trends" and "details on the rise in ABC Company's stock price."

[0900] Through this series of processes, the system effectively provides content that users are interested in, increasing user satisfaction. In addition, by automatically generating and delivering video content customized for each user, the way information is received becomes more personalized.

[0901] The processing flow will be explained below.

[0902] Step 1:

[0903] The server retrieves the user's past behavioral data and registration information, which includes collecting data such as the articles the user has viewed, the portfolio holdings they have registered, their search history, and categories of interest.

[0904] Step 2:

[0905] The server analyzes the acquired data to identify the user's areas of interest, including the categories the user frequently browses and extracting related keywords.

[0906] Step 3:

[0907] The server uses generative AI to generate a script tailored to each user, which generates news, market trends, product descriptions, and other information in natural language based on the user's areas of interest.

[0908] Step 4:

[0909] Based on the generated script, the server collects or generates the necessary image materials, such as graphs and charts showing market trends and photos related to the news article, and, if necessary, uses generative AI to create new image materials.

[0910] Step 5:

[0911] The server combines the image material with the generated script to create a video, using a text-to-speech program to convert the script into audio and synchronize it with the image material.

[0912] Step 6:

[0913] The server uploads the generated video file to the specified platform and prepares it for distribution. As soon as the video is published on the distribution platform, a notification is sent to the user's device.

[0914] Step 7:

[0915] Users receive a notification sent to their device to watch the video. Clicking on the notification opens the application and plays the generated video. Users can review the video content and find out the information they are interested in.

[0916] Example 1

[0917] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0918] Conventional information distribution systems have struggled to automatically generate and efficiently distribute personalized content for each user. As a result, it has been difficult to provide information tailored to the user's interests, often resulting in low user satisfaction. Furthermore, manual content creation and distribution is time-consuming and costly, so efficient operation is required.

[0919] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0920] In this invention, the server includes means for acquiring a user's area of ​​interest from the user's past behavioral data and registration information, means for generating a script suitable for each user using a generative AI based on the acquired user's area of ​​interest, means for collecting or generating related image materials based on the generated script, means for combining the script and image materials and generating audio and video using a text-to-speech program, means for distributing the generated video to a specific platform and notifying the user's device, and means for viewing the notified content on the user's device. This makes it possible to automatically generate and distribute video content personalized for each user, thereby achieving high user satisfaction.

[0921] "User past behavioral data" refers to data such as the actions a user takes on an online platform, browsing history, search history, and purchase history.

[0922] "Registration Information" means the personal information and interest category information provided by a User when registering for the Service.

[0923] "User interest areas" are the range of categories or topics in which a user is interested, identified based on the user's behavioral data and registration information.

[0924] "Generative AI" is an artificial intelligence model that generates natural language text, images, and other content based on data and prompts provided.

[0925] A "script" is text containing narration and content instructions for a video, generated based on information related to the user's area of ​​interest.

[0926] "Image material" refers to visual content used in a video, such as photographs, graphs, charts, logos, etc.

[0927] A "read-aloud program" is software that converts text into synthetic speech and plays it back as narration.

[0928] "Video" means dynamic visual content consisting of multiple frames and may include audio and text.

[0929] "Platform" is a general term for websites and applications that allow users to upload generated videos and make them viewable.

[0930] "Notifications" are messages that inform users when newly generated videos or content is available.

[0931] "Device" means a device on which a user receives notifications and views content, including a smartphone, tablet, or PC.

[0932] The present invention provides an information distribution system that automatically generates and distributes video content optimized for each user. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS The following describes in detail an embodiment of the present invention.

[0933] Retrieving User Information

[0934] The server retrieves the user's past behavioral data and registration information from the database, including browsing history, registered portfolio stocks, search history, category information of interest, etc. When the server retrieves this data, it uses the user ID to efficiently collect related data.

[0935] As a concrete example, the server issues an SQL query to retrieve data as follows:

[0936] SELECT FROM user_behavior WHERE user_id = 'userA';

[0937] Script generation

[0938] The server uses generative AI (e.g., GPT-4) to generate a script tailored to each user based on the acquired user information, and sends prompts to generate news, market trends, and product descriptions related to the user's areas of interest in natural language.

[0939] As a concrete example, if user A is interested in technology stock information, the server might use a prompt like this:

[0940] User A is interested in technology stocks. Create a news script for this user.

[0941] In response to this prompt, the generative AI outputs a script that reads, "Today's technology market trends: ABC Company's stock price rose 5%."

[0942] Image material generation and acquisition

[0943] The server collects or generates the necessary image materials based on the generated script, such as photos, graphs, charts, etc., using external data sources or generative AI (e.g., DALL-E).

[0944] You can use the API to obtain image materials. For example, issue the following API request to obtain image materials.

[0945] GET / stock_images?company=ABC

[0946] Video generation

[0947] The server combines the generated script with image materials and generates a video using automatic video generation software (e.g., Adobe Premiere Pro API), and also converts the script into audio using a speech synthesis API (e.g., Google Text-to-Speech) and embeds it in the video.

[0948] As a concrete example, an API request is issued as follows to generate a video.

[0949] POST / generate_video

[0950] {

[0951] "script": "Today's technology market trends: ABC Company's stock price rose 5%.",

[0952] "images": ["abc_logo.png", "stock_chart.png"],

[0953] "voice": "synthesized voice file"

[0954] }

[0955] Video distribution

[0956] The server uploads the generated video to a specific platform (e.g., YouTube) and notifies the user's device. The server then uses the platform's API to upload the video and obtain its URL.

[0957] As a concrete example, you can upload a video by issuing an API request as follows:

[0958] POST / upload_video

[0959] {

[0960] "video_file": "generated_video.mp4",

[0961] "platform": "YouTube"

[0962] }

[0963] After uploading, a notification message will be sent to the user's device.

[0964] Notification: "New technology stock info video added. Watch it here: [URL]"

[0965] User Viewing

[0966] The user receives a notification on their device, opens an application (e.g., a messaging app) to watch the video, and can tap the notification to go directly to the viewing page.

[0967] For example, when a user taps on a notification on their device, a messaging app opens with the URL of the specified video, which the user can click to watch the video.

[0968] Through these steps, the present invention realizes the automatic generation and distribution of video content optimized for each user, and builds an information distribution system that provides high levels of user satisfaction.

[0969] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0970] Step 1: Get user information

[0971] The server retrieves the user's past behavioral data and registration information from the database. As input, the user ID is provided to the server, and the server executes a query against the database based on this user ID. Specifically, the server issues an SQL query as follows:

[0972] SELECT FROM user_behavior WHERE user_id = 'userA';

[0973] This query outputs data related to the user's areas of interest (browsing history, registered portfolio stocks, search history, interest categories, etc.).

[0974] Step 2: Generate the script

[0975] The server uses the acquired user information as input and sends a prompt to the generative AI (e.g., GPT-4) to generate a script. Specifically, the server sends the following prompt to the generative AI:

[0976] User A is interested in technology stocks. Create a news script for this user.

[0977] Based on this prompt, the generative AI outputs text related to the user's area of ​​interest (e.g., "Today's technology market trends: ABC Company's stock price rose 5%.").

[0978] Step 3: Creating and acquiring image materials

[0979] The server uses the generated script as input to collect or generate the necessary image materials. Specifically, the server issues an API request to an external data source.

[0980] GET / stock_images?company=ABC

[0981] This request will output relevant image material (e.g., ABC company logo, stock price graph).

[0982] Step 4: Generate the video

[0983] The server generates a video using automatic video generation software (e.g., Adobe Premiere Pro API) using the generated script and collected image materials as input. Specifically, the server issues the following API request:

[0984] POST / generate_video

[0985] {

[0986] "script": "Today's technology market trends: ABC Company's stock price rose 5%.",

[0987] "images": ["abc_logo.png", "stock_chart.png"],

[0988] "voice": "synthesized voice file"

[0989] }

[0990] This request causes the generated video to be output.

[0991] Step 5: Publish your video

[0992] The server takes the generated video as input, uploads it to a specific platform (e.g., YouTube), and notifies the user's device. Specifically, the server issues the following API request:

[0993] POST / upload_video

[0994] {

[0995] "video_file": "generated_video.mp4",

[0996] "platform": "YouTube"

[0997] }

[0998] This action will output a video URL from the platform, which the server will then use to send a notification to the user's device.

[0999] Notification: "New technology stock info video added. Watch it here: [URL]"

[1000] Step 6: Watch the video

[1001] The user receives a notification on their device and opens an application (e.g., a messaging app) to watch the video. Specifically, the user taps the notification, which opens the application and displays the video URL. Clicking on this URL starts watching the video.

[1002] This series of steps realizes a system that can automatically generate and efficiently deliver video content optimized for each user.

[1003] (Application example 1)

[1004] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1005] With conventional content distribution systems, it was difficult to automatically generate and distribute video content optimized for user interests. Furthermore, content individualization was insufficient, limiting improvements in user satisfaction. Furthermore, notifications of generated video content were not provided in a timely manner, resulting in users missing viewing opportunities.

[1006] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1007] In this invention, the server includes means for acquiring a user's area of ​​interest from the user's past behavioral data and registration information, means for generating a script appropriate for each user using a generative AI based on the acquired user's area of ​​interest, means for collecting or generating related image materials based on the generated script, means for combining the script and image materials and adding narration using a text-to-speech program to generate a video, means for distributing the generated video to a platform and notifying the user's device, and means for notifying the user that the generated video is available. This makes it possible to automatically generate and effectively distribute video content optimized for user interests.

[1008] "User past behavioral data" refers to records of a user's browsing history, search history, click history, etc. on online platforms.

[1009] "Registration Information" refers to personal information, areas of interest, settings information, etc. provided by a user when registering for the service.

[1010] "Interest areas" refer to categories, themes, or topics that a user is particularly interested in.

[1011] "Generative AI" refers to an artificial intelligence model that automatically generates new content and information based on acquired data.

[1012] "Script" refers to a screenplay written in narration or text format created by generative AI.

[1013] "Image material" refers to visual materials that complement the generated script, including photographs, illustrations, graphs, etc.

[1014] A "reading program" refers to software that converts a text script into audio and reads it aloud as narration.

[1015] "Video" refers to visual and audio content created by combining a generated script with image material and narration.

[1016] "Platform" refers to online services and applications for delivering generated videos to users.

[1017] "Device" refers to the electronic device, such as a smartphone or tablet, that a User uses to receive and watch videos.

[1018] "Notification" means a message or alert that notifies the user that the generated video is available for viewing.

[1019] The present invention provides a system for automatically generating and distributing individually optimized video content based on a user's past behavioral data and registration information. Hereinafter, an embodiment of the present invention will be described in detail.

[1020] First, the server obtains the user's past behavioral data and registration information, including the articles the user has viewed on the online platform, search history, categories of interest, etc. Based on this information, the server identifies the user's areas of interest.

[1021] The server then uses a generative AI model (e.g., GPT-4) to generate a script based on the identified areas of interest. This script is tailored to the user's interests and may include specific topics such as "latest technology stock market trends." Examples of prompts include:

[1022] "Users are interested in: Technology stock information

[1023] Generation Objective: Generate a script about the latest technology stock market trends.

[1024] Based on the generated script, the server collects or generates relevant image materials, such as photos related to the news article or graphs and charts showing market trends. Generative AI can also be used to generate new image materials if necessary.

[1025] The server then combines the script with the collected image material to generate a video, adding narration using a text-to-speech program that reads the script aloud. The resulting video combines specific visual content and narration that correspond to the user's interests.

[1026] The generated video is uploaded to the specified platform (e.g., a social media app), and the server notifies the user's device that the generated video is available, allowing the user to immediately know that new video content is available for viewing.

[1027] Users receive a notification on their device and can watch the generated video. The video contains content optimized for the user's interests, resulting in a high level of satisfaction.

[1028] The specific hardware and software used to implement the above process include servers, user devices (smartphones, tablets, etc.), generative AI models using Python (e.g., GPT-4), image processing libraries (e.g., Pillow), and video editing libraries (e.g., MoviePy). This enables the automatic generation and distribution of individually optimized video content.

[1029] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1030] Step 1:

[1031] The server retrieves the user's past behavioral data and registration information. This information includes the articles the user has viewed on the online platform, their search history, and categories of interest. The user ID is used as input. The retrieved data is used to identify the user's areas of interest. Specifically, the server retrieves user information from the database using a REST API.

[1032] Step 2:

[1033] The server generates a script using a generative AI model (e.g., GPT-4) based on the acquired user's areas of interest. In this process, the server uses the user's areas of interest as input and provides a prompt sentence to the natural language generation AI. Specifically, the server inputs the areas of interest in text format into the AI ​​model and obtains the generated script text. The output is a script customized for each user.

[1034] Step 3:

[1035] The server collects or generates related image materials based on the generated script. The generated script text is used as input. Specifically, it collects images through a web API or generates new images using generative AI. The output is image materials that correspond to the script content.

[1036] Step 4:

[1037] The server combines the script with the collected image material to generate a video. The script text and image material are used as input. Specifically, it converts the script content into audio using a text-to-speech program, synchronizes it with the image material, and generates a video using a video editing library (e.g., MoviePy). The output is a video file with narration.

[1038] Step 5:

[1039] The server distributes the generated video to the specified platform. The generated video file and information about the distribution platform are used as input. The specific operation is to upload the video file to the platform using the REST API. The output is the status of distribution completion.

[1040] Step 6:

[1041] The server notifies the user's device that the generated video is available. The inputs are the delivery completion status and the user's contact information. The specific operation is to send a notification to the user's smartphone using a push notification service. The output is the notification sending status.

[1042] Step 7:

[1043] The user receives the notification on their device and watches the generated video. The notification content and the URL of the video distribution destination are used as input. The specific operation is to open a video viewing application through the smartphone's notification system and play the video. The output is log data of the video playback status.

[1044] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1045] The present invention provides an information distribution system that combines a user's past behavioral data and registered information with an emotion engine that recognizes the user's emotions. Hereinafter, embodiments of the present invention will be described in detail.

[1046] Obtaining user information and sentiment

[1047] The server retrieves the user's past behavioral data and registration information, and further recognizes the user's emotional state using an emotion engine, including the articles the user has viewed, the portfolio stocks they have registered, their search history, categories of interest, and emotional data through facial recognition and voice analysis.

[1048] example:

[1049] Let's say User B is interested in "fitness-related information" and is feeling "good" today.

[1050] Script generation

[1051] The server uses generative AI to generate a script suited to each user based on the acquired user information and emotional data. By incorporating the emotional data, a more personalized script is created.

[1052] example:

[1053] For user B, a script about "effective fitness training methods" is generated in a positive tone that reflects the emotional data of "energetic."

[1054] Image material generation and acquisition

[1055] The server collects or generates the necessary image materials based on the generated script. These materials include visuals and graphs that demonstrate the training method. If necessary, new image materials can be generated using generative AI.

[1056] example:

[1057] Based on User B's script, images showing the steps of the training method and visuals to increase motivation are collected.

[1058] Video generation

[1059] The server then combines the generated script with the collected image materials to generate a video. It uses a text-to-speech program to convert the script into audio and synchronize it with the image materials. It also adjusts the tone and content of the narration based on the user's emotions.

[1060] example:

[1061] A video is created in an energetic tone that explains to User B a training method that is highly effective for fitness.

[1062] Video distribution

[1063] The server uploads the generated video file to the specified platform and prepares it for distribution. As soon as the video is published on the distribution platform, a notification is sent to the user's device.

[1064] example:

[1065] User B receives a notification on their device saying "New fitness video available."

[1066] User Viewing

[1067] Users receive a notification sent to their device to watch the video. Clicking on the notification opens the application and plays the generated video. Users can review the video content and find out the information they are interested in.

[1068] example:

[1069] User B clicks on the notification and watches a fitness video to get training information that suits their energetic mood.

[1070] Through this process, the system can effectively and personalizedly provide content that users are interested in, increasing user satisfaction. In addition, by combining it with an emotion engine, video content can be adapted to the user's current emotional state, resulting in deeper engagement.

[1071] The processing flow will be explained below.

[1072] Step 1:

[1073] The server retrieves the user's past behavioral data and registration information, which includes collecting data such as the articles the user has viewed, the portfolio holdings they have registered, their search history, and categories of interest.

[1074] Step 2:

[1075] The server uses an emotion engine to recognize the user's current emotional state, which includes analyzing the user's facial expressions, tone of voice, and input.

[1076] Step 3:

[1077] The server combines the acquired behavioral data and registration information with the emotional data to identify the user's areas of interest, including categories of interest, keywords, and the user's emotions.

[1078] Step 4:

[1079] The server uses generative AI to generate a script based on the identified areas of interest and emotional data, with content and tone that reflects the user's current emotional state.

[1080] Step 5:

[1081] The server then collects or generates relevant image material based on the generated script, including graphs and charts showing market trends, photos related to news articles, and even uses generative AI to create new image material that matches the emotion.

[1082] Step 6:

[1083] The server combines image material with the generated script to create a video, using a text-to-speech program to convert the script into audio and adjust the tone of the narration based on emotion.

[1084] Step 7:

[1085] The server uploads the generated video file to the specified platform and prepares it for distribution. As soon as the video is published on the distribution platform, a notification is sent to the user's device.

[1086] Step 8:

[1087] Users receive a notification sent to their device to watch the video. Clicking on the notification opens the application and plays the generated video. Users can review the video content and find out the information they are interested in.

[1088] Example 2

[1089] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1090] Conventional information delivery systems have performed personalization by taking into account a user's past behavioral data and registration information, but have not been able to provide content that reflects the user's emotional state. As a result, while they can provide information that matches a user's interests and concerns, it is difficult to provide optimal content that reflects the user's emotional state at the time, which limits the ability to improve user satisfaction and engagement. Another problem is that the time and effort required to generate videos hinders the rapid provision of information.

[1091] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for acquiring a user's past behavioral data and registration information, means for recognizing the user's emotional state using the acquired data, means for generating a script suitable for each user using a generative AI model based on the acquired user information and emotional data, means for collecting or generating related image materials based on the generated script, means for combining the script and image materials and adding narration using a text-to-speech program to generate a video, and means for uploading the generated video to a distribution platform and notifying the user's terminal. This enables rapid generation and distribution of personalized video content based on a user's past behavioral data, registration information, and emotional state.

[1092] "User's past behavioral data" refers to information such as the user's previous web browsing history, search history, browsing history, and registered portfolio stocks.

[1093] "Registration Information" refers to data such as personal attributes, interests, and preferences that a user provides to the system.

[1094] "Emotional state" is data that indicates a user's current emotions and mood, and is obtained using methods such as facial recognition and voice analysis.

[1095] "Means of acquisition" refers to the methods and technologies used to collect users' past behavioral data and registration information from databases, etc.

[1096] "Means for recognizing emotional state" refers to methods or technologies for determining a user's current emotions using facial recognition technology or voice analysis technology.

[1097] A "generative AI model" refers to a machine learning model that uses artificial intelligence to perform tasks such as natural language generation and image generation.

[1098] "Script" refers to the outline or script of content created for each user using a generative AI model.

[1099] "Related image material" refers to the visual content required based on the generated script, including collected images and images newly created by the generative AI.

[1100] "Read-out program" refers to a program for converting text information into speech.

[1101] "Means for generating video" refers to methods and technologies for creating video content by synthesizing a script, image materials, and narration audio.

[1102] A "distribution platform" refers to a web service or application that allows generated video content to be published and delivered to users.

[1103] "Means of notification" refers to the methods and technologies used to notify a user's device that new content is available for viewing.

[1104] The present invention is an information distribution system that combines a user's past behavioral data and registered information with an emotion engine that recognizes the user's emotional state. Details of this system are described below.

[1105] Obtaining user information and sentiment

[1106] The server uses a database management system (e.g., MySQL) to retrieve the user's past behavioral data and registration information. This information includes the articles the user has viewed, their search history, and registered portfolio stocks. The server also uses an emotion engine (e.g., Emotion API) to recognize the user's emotional state. This emotion data includes information obtained through facial recognition and voice analysis.

[1107] Examples:

[1108] If User B has viewed "fitness-related information" multiple times in the past, the server will acquire this historical data. Also, if the emotion engine recognizes User B as "healthy," the server will acquire this emotion data.

[1109] Script generation

[1110] The server generates a personalized script using a generative AI model (e.g., GPT-4) based on the acquired user information and emotion data. It creates a prompt sentence to input into the generative AI model, and the script is generated based on the content of that sentence.

[1111] Specific prompt examples:

[1112] "User B has viewed fitness-related articles multiple times in the past and is currently feeling energetic."

[1113] Examples:

[1114] For User B, a script about "effective fitness training methods" with a positive tone is generated, reflecting the emotional data of "energetic."

[1115] Image material generation and acquisition

[1116] The server collects or generates the necessary image materials based on the generated script. Specifically, it searches and collects the necessary visual materials using an image collection platform (e.g., Unsplash API), and in some cases generates new image materials using a generative AI model (e.g., DALL-E).

[1117] Examples:

[1118] Based on User B's script, images showing training methods and visuals to increase motivation are collected.

[1119] Video generation

[1120] The server combines the generated script with the collected image material to generate a video. It uses narration synthesis software (e.g., Amazon Polly) to convert the script into audio, and video editing software (e.g., Adobe Premiere Pro) to synchronize the audio. It also adjusts the tone and content of the narration based on the user's emotions.

[1121] Examples:

[1122] Videos are created that explain effective fitness training methods in an energetic tone.

[1123] Video distribution

[1124] The server uploads the resulting video file to a distribution platform (e.g., YouTube API) and sends a notification to the user's device. As soon as the video is published, the user's device is notified that a new video is available to watch.

[1125] Examples:

[1126] User B receives a notification on their device saying "New fitness video available."

[1127] User Viewing

[1128] Users receive a notification sent to their device to watch the video. Clicking on the notification launches the video streaming application and plays the generated video, allowing users to immediately consume content of interest.

[1129] Examples:

[1130] User B clicks on the notification and watches a fitness video to get training information that suits their energetic mood.

[1131] This system aims to improve user satisfaction and engagement by quickly generating and delivering personalized content based on users' interest data and emotional state.

[1132] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1133] Step 1: Get user information

[1134] Input: User's ID

[1135] Output: User's past behavior data and registration information

[1136] The server uses a database management system (e.g., MySQL) to retrieve past behavioral data and registration information using the user's ID as a key. The data includes browsing history, search history, registered portfolio stocks, etc.

[1137] Specific behavior:

[1138] sql

[1139] SELECT FROM user_data WHERE user_id = 'B';

[1140] By executing this query, user B's past behavioral data and registration information will be retrieved.

[1141] Step 2: Obtaining emotion data

[1142] Input: User's ID

[1143] Output: User's emotional state data

[1144] The server uses an emotion engine (e.g., Emotion API) to recognize the user's emotional state through facial recognition and voice analysis.

[1145] Specific behavior:

[1146] python

[1147] emotion_data = emotion_api.recognize_emotion(user_id='B')

[1148] This API call obtains User B's current emotional state (e.g., "energetic").

[1149] Step 3: Generate a prompt statement

[1150] Input: User's past behavior data, registration information, emotional state data

[1151] Output: prompt statement

[1152] The server creates prompt sentences to input into the generative AI model based on user information and emotional data.

[1153] Specific behavior:

[1154] python

[1155] prompt = f "User B has viewed fitness-related articles multiple times in the past and is currently feeling energetic."

[1156] This generates a personalized prompt for User B.

[1157] Step 4: Generate the script

[1158] Input: prompt statement

[1159] Output: Script

[1160] The server uses a generative AI model (e.g., GPT-4) to generate a script based on the prompt.

[1161] Specific behavior:

[1162] python

[1163] script = generate_ai_model.generate(prompt)

[1164] Running this generative AI model will generate a fitness-related script that is appropriate for User B.

[1165] Step 5: Identifying image material

[1166] Input: Script

[1167] Output: A list of required image materials

[1168] The server identifies the necessary image material from the generated script.

[1169] Specific behavior:

[1170] python

[1171] required_images = extract_images_from_script(script)

[1172] This identifies a list of image materials required for the script.

[1173] Step 6: Collecting and creating image materials

[1174] Input: List of image materials

[1175] Output: Image material

[1176] The server collects the necessary visual materials using an image collection platform (e.g., Unsplash API) and, in some cases, generates new image materials using a generative AI model (e.g., DALL-E).

[1177] Specific behavior:

[1178] python

[1179] images = unsplash_api.search_images(query=required_images)

[1180] generated_images = dalle_api.generate(query=required_images)

[1181] This allows the necessary image materials to be collected or generated.

[1182] Step 7: Generate narration

[1183] Input: Script

[1184] Output: Narration audio

[1185] The server uses narration synthesis software (e.g., Amazon Polly) to convert the contents of the script into voice.

[1186] Specific behavior:

[1187] python

[1188] narration = polly_synthesize_speech(text=script)

[1189] This generates a narration voice based on the script.

[1190] Step 8: Edit your video

[1191] Input: script, image materials, narration audio

[1192] Output: Finished video

[1193] The server uses video editing software (e.g., Adobe Premiere Pro API) to generate a video that synchronizes the script, image materials, and narration audio.

[1194] Specific behavior:

[1195] python

[1196] video = create_video(narration=narration, images=images, script=script)

[1197] This process generates a video optimized for user B.

[1198] Step 9: Publish your video

[1199] Input: Finished video

[1200] Output: Video streaming link

[1201] The server uploads the generated video file to a distribution platform (e.g., YouTube API) and obtains a distribution link for the video.

[1202] Specific behavior:

[1203] python

[1204] upload_response = youtube_api.upload_video(file_path=video_path)

[1205] This will generate a distribution link for the video.

[1206] Step 10: Sending notifications

[1207] Input: Distribution link, user information

[1208] Output: Notification sent to the user's device

[1209] The server notifies the user's device that a new video is available for viewing.

[1210] Specific behavior:

[1211] python

[1212] send_notification(user_id='B', message='New fitness video available')

[1213] This will send a notification to User B's device.

[1214] Step 11: Watch the video

[1215] Input: Notification

[1216] Output: Video playback

[1217] The user receives a notification sent to their device and can watch the video. Clicking on the notification launches the video streaming application on their device and plays the generated video.

[1218] Specific behavior:

[1219] Click the notification on your device and open the video streaming application.

[1220] Play the new video to find out more.

[1221] This allows user B to view personalized video content.

[1222] (Application example 2)

[1223] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1224] Conventional information distribution systems were able to provide content based on users' past behavioral data and registration information, but it was difficult to provide personalized content that adapted to the user's emotional state. It is necessary to further improve user satisfaction and increase engagement by providing information based on the user's current emotions.

[1225] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1226] In this invention, the server includes: means for acquiring a user's areas of interest from the user's past behavioral data and registration information; means for generating a script appropriate for each user using a generative AI model based on the acquired user's areas of interest; means for collecting or generating related image materials based on the generated script; means for combining the script and image materials and adding narration using a text-to-speech program to generate a video; means for distributing the generated video to a platform and notifying the user's device; means for acquiring user emotional data through facial recognition and voice analysis; and means for adjusting the tone and content of the script based on the acquired emotional data, thereby enabling information distribution that is adapted to the user's current emotional state.

[1227] "User's past behavioral data and registration information" refers to data such as the user's previous actions and choices, browsing history, and personal information and areas of interest provided by the user when registering.

[1228] "Interest areas" are data that indicate specific categories, topics, or fields in which a user is interested.

[1229] A "generative AI model" is an artificial intelligence mechanism or program that generates specific content based on a user's areas of interest and emotional data.

[1230] A "script" is text data generated by a generative AI model to explain the content of an audio narration or video.

[1231] "Image materials" are images, graphs, and visual data that visually support the contents of the script.

[1232] A "reading program" is software that converts the contents of a script into audio and adds it to a video as narration.

[1233] A "platform" is an internet service or application for distributing video content.

[1234] "Facial recognition" is an image processing technology used to analyze emotions from a user's face.

[1235] "Voice analysis" is an acoustic processing technology used to analyze emotions and intentions from a user's speech.

[1236] "Emotional data" is data that represents a user's emotional state, obtained through facial recognition or voice analysis.

[1237] The present invention is an information distribution system that provides individually personalized information based on a user's past behavioral data, registration information, and emotional data. Specific embodiments will be described below.

[1238] Obtaining user information and sentiment

[1239] The server acquires the user's past behavioral data and registration information, and analyzes the user's emotional data using an emotion recognition engine. The specific software used for this is OpenCV and an emotion recognition model (e.g., emotion_model.onnx). The server acquires the user's emotions through facial recognition, and analyzes emotions from the user's speech through voice analysis.

[1240] example:

[1241] Let's say User B is interested in "fitness-related information" and is feeling "good" today.

[1242] Script generation

[1243] The server uses a generative AI model based on the acquired user information and emotional data to generate a script appropriate for the user. By incorporating emotional data, a more personalized script can be created. One example of the AI ​​model used is GPT-3.

[1244] Example prompt for a generative AI model:

[1245] "Generate a cheerful tone script about effective fitness training methods."

[1246] Image material generation and acquisition

[1247] The server collects or generates the necessary image materials based on the generated script. These materials include visuals that demonstrate training methods and visuals to increase motivation. If necessary, new image materials may be generated using generative AI.

[1248] example:

[1249] Based on User B's script, images showing the steps of the training method and visuals to increase motivation are collected.

[1250] Video generation

[1251] The server then combines the generated script with the collected image materials to generate a video. A text-to-speech program converts the script into audio and synchronizes it with the image materials. This program uses tools such as Google Text-to-Speech (gTTS). The server also adjusts the tone and content of the narration based on the user's emotions.

[1252] example:

[1253] A video is created in an energetic tone that explains to User B a training method that is highly effective for fitness.

[1254] Video distribution

[1255] The server uploads the generated video file to the distribution platform and sends a notification to the user's device using the distribution platform's API.

[1256] example:

[1257] User B receives a notification on their device saying "New fitness video available."

[1258] User Viewing

[1259] Users receive a notification sent to their device and can watch the video. Clicking on the notification opens the application and plays the generated video. This allows users to efficiently obtain information that matches their emotional state.

[1260] example:

[1261] User B clicks on the notification and watches a fitness video to get training information that suits their energetic mood.

[1262] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1263] Step 1:

[1264] Obtaining user information and sentiment

[1265] The server obtains the user's past behavioral data and registration information from the API, and collects emotional data using facial recognition and voice analysis. For example, the user's browsing history, search history, and registration information are input, and an emotion recognition engine (OpenCV and emotion recognition model) outputs emotional tags such as "cheerful" or "sad" from facial images and voice. This integrates the user's behavioral data and emotional data.

[1266] Step 2:

[1267] Script generation

[1268] The server generates a script using a generative AI model (e.g., GPT-3) based on the acquired user information and emotional data. The input is the user's area of ​​interest (e.g., fitness) and emotional data (e.g., "energetic"), and the script text generated based on the prompt ("Generate a cheerful tone script about effective fitness training methods.") is output. This generates text that matches the user's emotions.

[1269] Step 3:

[1270] Image material generation and acquisition

[1271] The server collects or generates related image materials based on the generated script. The input is the script text (e.g., "Explanation of effective fitness training methods in a lively tone"), and the output is a corresponding image file (e.g., a diagram showing training steps). This may involve using an image collection service or image generation AI.

[1272] Step 4:

[1273] Video generation

[1274] The server combines the generated script with image materials to generate a video. It then uses a text-to-speech program to convert the script into audio and synchronize it with the image materials. The input is the script text and image files, and the output is a video file with narration. Google Text-to-Speech (gTTS) is used for speech synthesis.

[1275] Step 5:

[1276] Video distribution

[1277] The server uploads the generated video file to the distribution platform and sends a notification to the user's device. The input is the video file and user information, and the output is the video uploaded to the distribution platform and a notification message to the user (e.g., "A new fitness video is available to watch"). This allows the user to be aware of the existence of the video and watch it.

[1278] Step 6:

[1279] User Viewing

[1280] The user receives a notification sent to their device and watches the video. The input is the notification message and the device, and the output is the video being played and the user can watch the content. This allows the user to enjoy personalized content.

[1281] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1282] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1283] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1284] [Fourth embodiment]

[1285] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1286] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1287] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1288] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1289] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1290] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1291] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1292] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1293] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1294] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1295] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1296] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1297] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1298] The present invention provides a system for automatically generating and distributing video content optimized for each user based on the user's past behavioral data and registration information. Hereinafter, a specific embodiment of the present invention will be described.

[1299] Retrieving User Information

[1300] The server obtains the user's past behavioral data and registration information, including the articles the user has viewed on the online platform, the portfolio stocks registered, search history, and categories of interest.

[1301] Script generation

[1302] Based on the acquired user information, the server uses generative AI to generate a script tailored to each user, specifically generating news, market trends, product descriptions, and other information related to the user's areas of interest in natural language.

[1303] example:

[1304] If User A is interested in "technology-related stock information," the server generates scripts about "today's technology stock market trends" and "earnings reports for specific companies."

[1305] Image material generation and acquisition

[1306] Based on the generated script, the server collects or generates the necessary image materials, such as graphs and charts showing market trends, photos related to news articles, etc. It can also use generative AI to generate new image materials as needed.

[1307] example:

[1308] If the script for user A is "about the rise in stock prices of ABC Company," the logo of ABC Company and a graph showing fluctuations in stock prices will be collected as image materials.

[1309] Video generation

[1310] The server combines the generated script with the collected image materials to generate a video. A text-to-speech program converts the script into audio and synchronizes it with the image materials.

[1311] example:

[1312] A script is generated saying "ABC Company's stock price has risen by 5%" and a video with narration is generated along with the corresponding stock price graph.

[1313] Video distribution

[1314] The server uploads the generated video to a designated platform (e.g., a social media app) and distributes it to users. The notification function notifies users of the availability of new content on their devices.

[1315] example:

[1316] User A receives a notification on his device saying, "A new video about technology stocks has been added." User A receives the notification and can watch the video.

[1317] User Viewing

[1318] Users receive notifications and can watch videos on their devices. Content tailored to their interests and preferences is available, resulting in high levels of satisfaction.

[1319] example:

[1320] User A opens the LINE app and watches a video streamed from the server. The video includes information on "technology stock market trends" and "details on the rise in ABC Company's stock price."

[1321] Through this series of processes, the system effectively provides content that users are interested in, increasing user satisfaction. In addition, by automatically generating and delivering video content customized for each user, the way information is received becomes more personalized.

[1322] The processing flow will be explained below.

[1323] Step 1:

[1324] The server retrieves the user's past behavioral data and registration information, which includes collecting data such as the articles the user has viewed, the portfolio holdings they have registered, their search history, and categories of interest.

[1325] Step 2:

[1326] The server analyzes the acquired data to identify the user's areas of interest, including the categories the user frequently browses and extracting related keywords.

[1327] Step 3:

[1328] The server uses generative AI to generate a script tailored to each user, which generates news, market trends, product descriptions, and other information in natural language based on the user's areas of interest.

[1329] Step 4:

[1330] Based on the generated script, the server collects or generates the necessary image materials, such as graphs and charts showing market trends and photos related to the news article, and, if necessary, uses generative AI to create new image materials.

[1331] Step 5:

[1332] The server combines the image material with the generated script to create a video, using a text-to-speech program to convert the script into audio and synchronize it with the image material.

[1333] Step 6:

[1334] The server uploads the generated video file to the specified platform and prepares it for distribution. As soon as the video is published on the distribution platform, a notification is sent to the user's device.

[1335] Step 7:

[1336] Users receive a notification sent to their device to watch the video. Clicking on the notification opens the application and plays the generated video. Users can review the video content and find out the information they are interested in.

[1337] Example 1

[1338] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1339] Conventional information distribution systems have struggled to automatically generate and efficiently distribute personalized content for each user. As a result, it has been difficult to provide information tailored to the user's interests, often resulting in low user satisfaction. Furthermore, manual content creation and distribution is time-consuming and costly, so efficient operation is required.

[1340] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1341] In this invention, the server includes means for acquiring a user's area of ​​interest from the user's past behavioral data and registration information, means for generating a script suitable for each user using a generative AI based on the acquired user's area of ​​interest, means for collecting or generating related image materials based on the generated script, means for combining the script and image materials and generating audio and video using a text-to-speech program, means for distributing the generated video to a specific platform and notifying the user's device, and means for viewing the notified content on the user's device. This makes it possible to automatically generate and distribute video content personalized for each user, thereby achieving high user satisfaction.

[1342] "User past behavioral data" refers to data such as the actions a user takes on an online platform, browsing history, search history, and purchase history.

[1343] "Registration Information" means the personal information and interest category information provided by a User when registering for the Service.

[1344] "User interest areas" are the range of categories or topics in which a user is interested, identified based on the user's behavioral data and registration information.

[1345] "Generative AI" is an artificial intelligence model that generates natural language text, images, and other content based on data and prompts provided.

[1346] A "script" is text containing narration and content instructions for a video, generated based on information related to the user's area of ​​interest.

[1347] "Image material" refers to visual content used in a video, such as photographs, graphs, charts, logos, etc.

[1348] A "read-aloud program" is software that converts text into synthetic speech and plays it back as narration.

[1349] "Video" means dynamic visual content consisting of multiple frames and may include audio and text.

[1350] "Platform" is a general term for websites and applications that allow users to upload generated videos and make them viewable.

[1351] "Notifications" are messages that inform users when newly generated videos or content is available.

[1352] "Device" means a device on which a user receives notifications and views content, including a smartphone, tablet, or PC.

[1353] The present invention provides an information distribution system that automatically generates and distributes video content optimized for each user. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS The following describes in detail an embodiment of the present invention.

[1354] Retrieving User Information

[1355] The server retrieves the user's past behavioral data and registration information from the database, including browsing history, registered portfolio stocks, search history, category information of interest, etc. When the server retrieves this data, it uses the user ID to efficiently collect related data.

[1356] As a concrete example, the server issues an SQL query to retrieve data as follows:

[1357] SELECT FROM user_behavior WHERE user_id = 'userA';

[1358] Script generation

[1359] The server uses generative AI (e.g., GPT-4) to generate a script tailored to each user based on the acquired user information, and sends prompts to generate news, market trends, and product descriptions related to the user's areas of interest in natural language.

[1360] As a concrete example, if user A is interested in technology stock information, the server might use a prompt like this:

[1361] User A is interested in technology stocks. Create a news script for this user.

[1362] In response to this prompt, the generative AI outputs a script that reads, "Today's technology market trends: ABC Company's stock price rose 5%."

[1363] Image material generation and acquisition

[1364] The server collects or generates the necessary image materials based on the generated script, such as photos, graphs, charts, etc., using external data sources or generative AI (e.g., DALL-E).

[1365] You can use the API to obtain image materials. For example, issue the following API request to obtain image materials.

[1366] GET / stock_images?company=ABC

[1367] Video generation

[1368] The server combines the generated script with image materials and generates a video using automatic video generation software (e.g., Adobe Premiere Pro API), and also converts the script into audio using a speech synthesis API (e.g., Google Text-to-Speech) and embeds it in the video.

[1369] As a concrete example, an API request is issued as follows to generate a video.

[1370] POST / generate_video

[1371] {

[1372] "script": "Today's technology market trends: ABC Company's stock price rose 5%.",

[1373] "images": ["abc_logo.png", "stock_chart.png"],

[1374] "voice": "synthesized voice file"

[1375] }

[1376] Video distribution

[1377] The server uploads the generated video to a specific platform (e.g., YouTube) and notifies the user's device. The server then uses the platform's API to upload the video and obtain its URL.

[1378] As a concrete example, you can upload a video by issuing an API request as follows:

[1379] POST / upload_video

[1380] {

[1381] "video_file": "generated_video.mp4",

[1382] "platform": "YouTube"

[1383] }

[1384] After uploading, a notification message will be sent to the user's device.

[1385] Notification: "New technology stock info video added. Watch it here: [URL]"

[1386] User Viewing

[1387] The user receives a notification on their device, opens an application (e.g., a messaging app) to watch the video, and can tap the notification to go directly to the viewing page.

[1388] For example, when a user taps on a notification on their device, a messaging app opens with the URL of the specified video, which the user can click to watch the video.

[1389] Through these steps, the present invention realizes the automatic generation and distribution of video content optimized for each user, and builds an information distribution system that provides high levels of user satisfaction.

[1390] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1391] Step 1: Get user information

[1392] The server retrieves the user's past behavioral data and registration information from the database. As input, the user ID is provided to the server, and the server executes a query against the database based on this user ID. Specifically, the server issues an SQL query as follows:

[1393] SELECT FROM user_behavior WHERE user_id = 'userA';

[1394] This query outputs data related to the user's areas of interest (browsing history, registered portfolio stocks, search history, interest categories, etc.).

[1395] Step 2: Generate the script

[1396] The server uses the acquired user information as input and sends a prompt to the generative AI (e.g., GPT-4) to generate a script. Specifically, the server sends the following prompt to the generative AI:

[1397] User A is interested in technology stocks. Create a news script for this user.

[1398] Based on this prompt, the generative AI outputs text related to the user's area of ​​interest (e.g., "Today's technology market trends: ABC Company's stock price rose 5%.").

[1399] Step 3: Creating and acquiring image materials

[1400] The server uses the generated script as input to collect or generate the necessary image materials. Specifically, the server issues an API request to an external data source.

[1401] GET / stock_images?company=ABC

[1402] This request will output relevant image material (e.g., ABC company logo, stock price graph).

[1403] Step 4: Generate the video

[1404] The server generates a video using automatic video generation software (e.g., Adobe Premiere Pro API) using the generated script and collected image materials as input. Specifically, the server issues the following API request:

[1405] POST / generate_video

[1406] {

[1407] "script": "Today's technology market trends: ABC Company's stock price rose 5%.",

[1408] "images": ["abc_logo.png", "stock_chart.png"],

[1409] "voice": "synthesized voice file"

[1410] }

[1411] This request causes the generated video to be output.

[1412] Step 5: Publish your video

[1413] The server takes the generated video as input, uploads it to a specific platform (e.g., YouTube), and notifies the user's device. Specifically, the server issues the following API request:

[1414] POST / upload_video

[1415] {

[1416] "video_file": "generated_video.mp4",

[1417] "platform": "YouTube"

[1418] }

[1419] This action will output a video URL from the platform, which the server will then use to send a notification to the user's device.

[1420] Notification: "New technology stock info video added. Watch it here: [URL]"

[1421] Step 6: Watch the video

[1422] The user receives a notification on their device and opens an application (e.g., a messaging app) to watch the video. Specifically, the user taps the notification, which opens the application and displays the video URL. Clicking on this URL starts watching the video.

[1423] This series of steps realizes a system that can automatically generate and efficiently deliver video content optimized for each user.

[1424] (Application example 1)

[1425] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1426] With conventional content distribution systems, it was difficult to automatically generate and distribute video content optimized for user interests. Furthermore, content individualization was insufficient, limiting improvements in user satisfaction. Furthermore, notifications of generated video content were not provided in a timely manner, resulting in users missing viewing opportunities.

[1427] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1428] In this invention, the server includes means for acquiring a user's area of ​​interest from the user's past behavioral data and registration information, means for generating a script appropriate for each user using a generative AI based on the acquired user's area of ​​interest, means for collecting or generating related image materials based on the generated script, means for combining the script and image materials and adding narration using a text-to-speech program to generate a video, means for distributing the generated video to a platform and notifying the user's device, and means for notifying the user that the generated video is available. This makes it possible to automatically generate and effectively distribute video content optimized for user interests.

[1429] "User past behavioral data" refers to records of a user's browsing history, search history, click history, etc. on online platforms.

[1430] "Registration Information" refers to personal information, areas of interest, settings information, etc. provided by a user when registering for the service.

[1431] "Interest areas" refer to categories, themes, or topics that a user is particularly interested in.

[1432] "Generative AI" refers to an artificial intelligence model that automatically generates new content and information based on acquired data.

[1433] "Script" refers to a screenplay written in narration or text format created by generative AI.

[1434] "Image material" refers to visual materials that complement the generated script, including photographs, illustrations, graphs, etc.

[1435] A "reading program" refers to software that converts a text script into audio and reads it aloud as narration.

[1436] "Video" refers to visual and audio content created by combining a generated script with image material and narration.

[1437] "Platform" refers to online services and applications for delivering generated videos to users.

[1438] "Device" refers to the electronic device, such as a smartphone or tablet, that a User uses to receive and watch videos.

[1439] "Notification" means a message or alert that notifies the user that the generated video is available for viewing.

[1440] The present invention provides a system for automatically generating and distributing individually optimized video content based on a user's past behavioral data and registration information. Hereinafter, an embodiment of the present invention will be described in detail.

[1441] First, the server obtains the user's past behavioral data and registration information, including the articles the user has viewed on the online platform, search history, categories of interest, etc. Based on this information, the server identifies the user's areas of interest.

[1442] The server then uses a generative AI model (e.g., GPT-4) to generate a script based on the identified areas of interest. This script is tailored to the user's interests and may include specific topics such as "latest technology stock market trends." Examples of prompts include:

[1443] "Users are interested in: Technology stock information

[1444] Generation Objective: Generate a script about the latest technology stock market trends.

[1445] Based on the generated script, the server collects or generates relevant image materials, such as photos related to the news article or graphs and charts showing market trends. Generative AI can also be used to generate new image materials if necessary.

[1446] The server then combines the script with the collected image material to generate a video, adding narration using a text-to-speech program that reads the script aloud. The resulting video combines specific visual content and narration that correspond to the user's interests.

[1447] The generated video is uploaded to the specified platform (e.g., a social media app), and the server notifies the user's device that the generated video is available, allowing the user to immediately know that new video content is available for viewing.

[1448] Users receive a notification on their device and can watch the generated video. The video contains content optimized for the user's interests, resulting in a high level of satisfaction.

[1449] The specific hardware and software used to implement the above process include servers, user devices (smartphones, tablets, etc.), generative AI models using Python (e.g., GPT-4), image processing libraries (e.g., Pillow), and video editing libraries (e.g., MoviePy). This enables the automatic generation and distribution of individually optimized video content.

[1450] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1451] Step 1:

[1452] The server retrieves the user's past behavioral data and registration information. This information includes the articles the user has viewed on the online platform, their search history, and categories of interest. The user ID is used as input. The retrieved data is used to identify the user's areas of interest. Specifically, the server retrieves user information from the database using a REST API.

[1453] Step 2:

[1454] The server generates a script using a generative AI model (e.g., GPT-4) based on the acquired user's areas of interest. In this process, the server uses the user's areas of interest as input and provides a prompt sentence to the natural language generation AI. Specifically, the server inputs the areas of interest in text format into the AI ​​model and obtains the generated script text. The output is a script customized for each user.

[1455] Step 3:

[1456] The server collects or generates related image materials based on the generated script. The generated script text is used as input. Specifically, it collects images through a web API or generates new images using generative AI. The output is image materials that correspond to the script content.

[1457] Step 4:

[1458] The server combines the script with the collected image material to generate a video. The script text and image material are used as input. Specifically, it converts the script content into audio using a text-to-speech program, synchronizes it with the image material, and generates a video using a video editing library (e.g., MoviePy). The output is a video file with narration.

[1459] Step 5:

[1460] The server distributes the generated video to the specified platform. The generated video file and information about the distribution platform are used as input. The specific operation is to upload the video file to the platform using the REST API. The output is the status of distribution completion.

[1461] Step 6:

[1462] The server notifies the user's device that the generated video is available. The inputs are the delivery completion status and the user's contact information. The specific operation is to send a notification to the user's smartphone using a push notification service. The output is the notification sending status.

[1463] Step 7:

[1464] The user receives the notification on their device and watches the generated video. The notification content and the URL of the video distribution destination are used as input. The specific operation is to open a video viewing application through the smartphone's notification system and play the video. The output is log data of the video playback status.

[1465] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1466] The present invention provides an information distribution system that combines a user's past behavioral data and registered information with an emotion engine that recognizes the user's emotions. Hereinafter, embodiments of the present invention will be described in detail.

[1467] Obtaining user information and sentiment

[1468] The server retrieves the user's past behavioral data and registration information, and further recognizes the user's emotional state using an emotion engine, including the articles the user has viewed, the portfolio stocks they have registered, their search history, categories of interest, and emotional data through facial recognition and voice analysis.

[1469] example:

[1470] Let's say User B is interested in "fitness-related information" and is feeling "good" today.

[1471] Script generation

[1472] The server uses generative AI to generate a script suited to each user based on the acquired user information and emotional data. By incorporating the emotional data, a more personalized script is created.

[1473] example:

[1474] For user B, a script about "effective fitness training methods" is generated in a positive tone that reflects the emotional data of "energetic."

[1475] Image material generation and acquisition

[1476] The server collects or generates the necessary image materials based on the generated script. These materials include visuals and graphs that demonstrate the training method. If necessary, new image materials can be generated using generative AI.

[1477] example:

[1478] Based on User B's script, images showing the steps of the training method and visuals to increase motivation are collected.

[1479] Video generation

[1480] The server then combines the generated script with the collected image materials to generate a video. It uses a text-to-speech program to convert the script into audio and synchronize it with the image materials. It also adjusts the tone and content of the narration based on the user's emotions.

[1481] example:

[1482] A video is created in an energetic tone that explains to User B a training method that is highly effective for fitness.

[1483] Video distribution

[1484] The server uploads the generated video file to the specified platform and prepares it for distribution. As soon as the video is published on the distribution platform, a notification is sent to the user's device.

[1485] example:

[1486] User B receives a notification on their device saying "New fitness video available."

[1487] User Viewing

[1488] Users receive a notification sent to their device to watch the video. Clicking on the notification opens the application and plays the generated video. Users can review the video content and find out the information they are interested in.

[1489] example:

[1490] User B clicks on the notification and watches a fitness video to get training information that suits their energetic mood.

[1491] Through this process, the system can effectively and personalizedly provide content that users are interested in, increasing user satisfaction. In addition, by combining it with an emotion engine, video content can be adapted to the user's current emotional state, resulting in deeper engagement.

[1492] The processing flow will be explained below.

[1493] Step 1:

[1494] The server retrieves the user's past behavioral data and registration information, which includes collecting data such as the articles the user has viewed, the portfolio holdings they have registered, their search history, and categories of interest.

[1495] Step 2:

[1496] The server uses an emotion engine to recognize the user's current emotional state, which includes analyzing the user's facial expressions, tone of voice, and input.

[1497] Step 3:

[1498] The server combines the acquired behavioral data and registration information with the emotional data to identify the user's areas of interest, including categories of interest, keywords, and the user's emotions.

[1499] Step 4:

[1500] The server uses generative AI to generate a script based on the identified areas of interest and emotional data, with content and tone that reflects the user's current emotional state.

[1501] Step 5:

[1502] The server then collects or generates relevant image material based on the generated script, including graphs and charts showing market trends, photos related to news articles, and even uses generative AI to create new image material that matches the emotion.

[1503] Step 6:

[1504] The server combines image material with the generated script to create a video, using a text-to-speech program to convert the script into audio and adjust the tone of the narration based on emotion.

[1505] Step 7:

[1506] The server uploads the generated video file to the specified platform and prepares it for distribution. As soon as the video is published on the distribution platform, a notification is sent to the user's device.

[1507] Step 8:

[1508] Users receive a notification sent to their device to watch the video. Clicking on the notification opens the application and plays the generated video. Users can review the video content and find out the information they are interested in.

[1509] Example 2

[1510] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1511] Conventional information delivery systems have performed personalization by taking into account a user's past behavioral data and registration information, but have not been able to provide content that reflects the user's emotional state. As a result, while they can provide information that matches a user's interests and concerns, it is difficult to provide optimal content that reflects the user's emotional state at the time, which limits the ability to improve user satisfaction and engagement. Another problem is that the time and effort required to generate videos hinders the rapid provision of information.

[1512] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for acquiring a user's past behavioral data and registration information, means for recognizing the user's emotional state using the acquired data, means for generating a script suitable for each user using a generative AI model based on the acquired user information and emotional data, means for collecting or generating related image materials based on the generated script, means for combining the script and image materials and adding narration using a text-to-speech program to generate a video, and means for uploading the generated video to a distribution platform and notifying the user's terminal. This enables rapid generation and distribution of personalized video content based on a user's past behavioral data, registration information, and emotional state.

[1513] "User's past behavioral data" refers to information such as the user's previous web browsing history, search history, browsing history, and registered portfolio stocks.

[1514] "Registration Information" refers to data such as personal attributes, interests, and preferences that a user provides to the system.

[1515] "Emotional state" is data that indicates a user's current emotions and mood, and is obtained using methods such as facial recognition and voice analysis.

[1516] "Means of acquisition" refers to the methods and technologies used to collect users' past behavioral data and registration information from databases, etc.

[1517] "Means for recognizing emotional state" refers to methods or technologies for determining a user's current emotions using facial recognition technology or voice analysis technology.

[1518] A "generative AI model" refers to a machine learning model that uses artificial intelligence to perform tasks such as natural language generation and image generation.

[1519] "Script" refers to the outline or script of content created for each user using a generative AI model.

[1520] "Related image material" refers to the visual content required based on the generated script, including collected images and images newly created by the generative AI.

[1521] "Read-out program" refers to a program for converting text information into speech.

[1522] "Means for generating video" refers to methods and technologies for creating video content by synthesizing a script, image materials, and narration audio.

[1523] A "distribution platform" refers to a web service or application that allows generated video content to be published and delivered to users.

[1524] "Means of notification" refers to the methods and technologies used to notify a user's device that new content is available for viewing.

[1525] The present invention is an information distribution system that combines a user's past behavioral data and registered information with an emotion engine that recognizes the user's emotional state. Details of this system are described below.

[1526] Obtaining user information and sentiment

[1527] The server uses a database management system (e.g., MySQL) to retrieve the user's past behavioral data and registration information. This information includes the articles the user has viewed, their search history, and registered portfolio stocks. The server also uses an emotion engine (e.g., Emotion API) to recognize the user's emotional state. This emotion data includes information obtained through facial recognition and voice analysis.

[1528] Examples:

[1529] If User B has viewed "fitness-related information" multiple times in the past, the server will acquire this historical data. Also, if the emotion engine recognizes User B as "healthy," the server will acquire this emotion data.

[1530] Script generation

[1531] The server generates a personalized script using a generative AI model (e.g., GPT-4) based on the acquired user information and emotion data. It creates a prompt sentence to input into the generative AI model, and the script is generated based on the content of that sentence.

[1532] Specific prompt examples:

[1533] "User B has viewed fitness-related articles multiple times in the past and is currently feeling energetic."

[1534] Examples:

[1535] For User B, a script about "effective fitness training methods" with a positive tone is generated, reflecting the emotional data of "energetic."

[1536] Image material generation and acquisition

[1537] The server collects or generates the necessary image materials based on the generated script. Specifically, it searches and collects the necessary visual materials using an image collection platform (e.g., Unsplash API), and in some cases generates new image materials using a generative AI model (e.g., DALL-E).

[1538] Examples:

[1539] Based on User B's script, images showing training methods and visuals to increase motivation are collected.

[1540] Video generation

[1541] The server combines the generated script with the collected image material to generate a video. It uses narration synthesis software (e.g., Amazon Polly) to convert the script into audio, and video editing software (e.g., Adobe Premiere Pro) to synchronize the audio. It also adjusts the tone and content of the narration based on the user's emotions.

[1542] Examples:

[1543] Videos are created that explain effective fitness training methods in an energetic tone.

[1544] Video distribution

[1545] The server uploads the resulting video file to a distribution platform (e.g., YouTube API) and sends a notification to the user's device. As soon as the video is published, the user's device is notified that a new video is available to watch.

[1546] Examples:

[1547] User B receives a notification on their device saying "New fitness video available."

[1548] User Viewing

[1549] Users receive a notification sent to their device to watch the video. Clicking on the notification launches the video streaming application and plays the generated video, allowing users to immediately consume content of interest.

[1550] Examples:

[1551] User B clicks on the notification and watches a fitness video to get training information that suits their energetic mood.

[1552] This system aims to improve user satisfaction and engagement by quickly generating and delivering personalized content based on users' interest data and emotional state.

[1553] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1554] Step 1: Get user information

[1555] Input: User's ID

[1556] Output: User's past behavior data and registration information

[1557] The server uses a database management system (e.g., MySQL) to retrieve past behavioral data and registration information using the user's ID as a key. The data includes browsing history, search history, registered portfolio stocks, etc.

[1558] Specific behavior:

[1559] sql

[1560] SELECT FROM user_data WHERE user_id = 'B';

[1561] By executing this query, user B's past behavioral data and registration information will be retrieved.

[1562] Step 2: Obtaining emotion data

[1563] Input: User's ID

[1564] Output: User's emotional state data

[1565] The server uses an emotion engine (e.g., Emotion API) to recognize the user's emotional state through facial recognition and voice analysis.

[1566] Specific behavior:

[1567] python

[1568] emotion_data = emotion_api.recognize_emotion(user_id='B')

[1569] This API call obtains User B's current emotional state (e.g., "energetic").

[1570] Step 3: Generate a prompt statement

[1571] Input: User's past behavior data, registration information, emotional state data

[1572] Output: prompt statement

[1573] The server creates prompt sentences to input into the generative AI model based on user information and emotional data.

[1574] Specific behavior:

[1575] python

[1576] prompt = f "User B has viewed fitness-related articles multiple times in the past and is currently feeling energetic."

[1577] This generates a personalized prompt for User B.

[1578] Step 4: Generate the script

[1579] Input: prompt statement

[1580] Output: Script

[1581] The server uses a generative AI model (e.g., GPT-4) to generate a script based on the prompt.

[1582] Specific behavior:

[1583] python

[1584] script = generate_ai_model.generate(prompt)

[1585] Running this generative AI model will generate a fitness-related script that is appropriate for User B.

[1586] Step 5: Identifying image material

[1587] Input: Script

[1588] Output: A list of required image materials

[1589] The server identifies the necessary image material from the generated script.

[1590] Specific behavior:

[1591] python

[1592] required_images = extract_images_from_script(script)

[1593] This identifies a list of image materials required for the script.

[1594] Step 6: Collecting and creating image materials

[1595] Input: List of image materials

[1596] Output: Image material

[1597] The server collects the necessary visual materials using an image collection platform (e.g., Unsplash API) and, in some cases, generates new image materials using a generative AI model (e.g., DALL-E).

[1598] Specific behavior:

[1599] python

[1600] images = unsplash_api.search_images(query=required_images)

[1601] generated_images = dalle_api.generate(query=required_images)

[1602] This allows the necessary image materials to be collected or generated.

[1603] Step 7: Generate narration

[1604] Input: Script

[1605] Output: Narration audio

[1606] The server uses narration synthesis software (e.g., Amazon Polly) to convert the contents of the script into voice.

[1607] Specific behavior:

[1608] python

[1609] narration = polly_synthesize_speech(text=script)

[1610] This generates a narration voice based on the script.

[1611] Step 8: Edit your video

[1612] Input: script, image materials, narration audio

[1613] Output: Finished video

[1614] The server uses video editing software (e.g., Adobe Premiere Pro API) to generate a video that synchronizes the script, image materials, and narration audio.

[1615] Specific behavior:

[1616] python

[1617] video = create_video(narration=narration, images=images, script=script)

[1618] This process generates a video optimized for user B.

[1619] Step 9: Publish your video

[1620] Input: Finished video

[1621] Output: Video streaming link

[1622] The server uploads the generated video file to a distribution platform (e.g., YouTube API) and obtains a distribution link for the video.

[1623] Specific behavior:

[1624] python

[1625] upload_response = youtube_api.upload_video(file_path=video_path)

[1626] This will generate a distribution link for the video.

[1627] Step 10: Sending notifications

[1628] Input: Distribution link, user information

[1629] Output: Notification sent to the user's device

[1630] The server notifies the user's device that a new video is available for viewing.

[1631] Specific behavior:

[1632] python

[1633] send_notification(user_id='B', message='New fitness video available')

[1634] This will send a notification to User B's device.

[1635] Step 11: Watch the video

[1636] Input: Notification

[1637] Output: Video playback

[1638] The user receives a notification sent to their device and can watch the video. Clicking on the notification launches the video streaming application on their device and plays the generated video.

[1639] Specific behavior:

[1640] Click the notification on your device and open the video streaming application.

[1641] Play the new video to find out more.

[1642] This allows user B to view personalized video content.

[1643] (Application example 2)

[1644] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1645] Conventional information distribution systems were able to provide content based on users' past behavioral data and registration information, but it was difficult to provide personalized content that adapted to the user's emotional state. It is necessary to further improve user satisfaction and increase engagement by providing information based on the user's current emotions.

[1646] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1647] In this invention, the server includes: means for acquiring a user's areas of interest from the user's past behavioral data and registration information; means for generating a script appropriate for each user using a generative AI model based on the acquired user's areas of interest; means for collecting or generating related image materials based on the generated script; means for combining the script and image materials and adding narration using a text-to-speech program to generate a video; means for distributing the generated video to a platform and notifying the user's device; means for acquiring user emotional data through facial recognition and voice analysis; and means for adjusting the tone and content of the script based on the acquired emotional data, thereby enabling information distribution that is adapted to the user's current emotional state.

[1648] "User's past behavioral data and registration information" refers to data such as the user's previous actions and choices, browsing history, and personal information and areas of interest provided by the user when registering.

[1649] "Interest areas" are data that indicate specific categories, topics, or fields in which a user is interested.

[1650] A "generative AI model" is an artificial intelligence mechanism or program that generates specific content based on a user's areas of interest and emotional data.

[1651] A "script" is text data generated by a generative AI model to explain the content of an audio narration or video.

[1652] "Image materials" are images, graphs, and visual data that visually support the contents of the script.

[1653] A "reading program" is software that converts the contents of a script into audio and adds it to a video as narration.

[1654] A "platform" is an internet service or application for distributing video content.

[1655] "Facial recognition" is an image processing technology used to analyze emotions from a user's face.

[1656] "Voice analysis" is an acoustic processing technology used to analyze emotions and intentions from a user's speech.

[1657] "Emotional data" is data that represents a user's emotional state, obtained through facial recognition or voice analysis.

[1658] The present invention is an information distribution system that provides individually personalized information based on a user's past behavioral data, registration information, and emotional data. Specific embodiments will be described below.

[1659] Obtaining user information and sentiment

[1660] The server acquires the user's past behavioral data and registration information, and analyzes the user's emotional data using an emotion recognition engine. The specific software used for this is OpenCV and an emotion recognition model (e.g., emotion_model.onnx). The server acquires the user's emotions through facial recognition, and analyzes emotions from the user's speech through voice analysis.

[1661] example:

[1662] Let's say User B is interested in "fitness-related information" and is feeling "good" today.

[1663] Script generation

[1664] The server uses a generative AI model based on the acquired user information and emotional data to generate a script appropriate for the user. By incorporating emotional data, a more personalized script can be created. One example of the AI ​​model used is GPT-3.

[1665] Example prompt for a generative AI model:

[1666] "Generate a cheerful tone script about effective fitness training methods."

[1667] Image material generation and acquisition

[1668] The server collects or generates the necessary image materials based on the generated script. These materials include visuals that demonstrate training methods and visuals to increase motivation. If necessary, new image materials may be generated using generative AI.

[1669] example:

[1670] Based on User B's script, images showing the steps of the training method and visuals to increase motivation are collected.

[1671] Video generation

[1672] The server then combines the generated script with the collected image materials to generate a video. A text-to-speech program converts the script into audio and synchronizes it with the image materials. This program uses tools such as Google Text-to-Speech (gTTS). The server also adjusts the tone and content of the narration based on the user's emotions.

[1673] example:

[1674] A video is created in an energetic tone that explains to User B a training method that is highly effective for fitness.

[1675] Video distribution

[1676] The server uploads the generated video file to the distribution platform and sends a notification to the user's device using the distribution platform's API.

[1677] example:

[1678] User B receives a notification on their device saying "New fitness video available."

[1679] User Viewing

[1680] Users receive a notification sent to their device and can watch the video. Clicking on the notification opens the application and plays the generated video. This allows users to efficiently obtain information that matches their emotional state.

[1681] example:

[1682] User B clicks on the notification and watches a fitness video to get training information that suits their energetic mood.

[1683] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1684] Step 1:

[1685] Obtaining user information and sentiment

[1686] The server obtains the user's past behavioral data and registration information from the API, and collects emotional data using facial recognition and voice analysis. For example, the user's browsing history, search history, and registration information are input, and an emotion recognition engine (OpenCV and emotion recognition model) outputs emotional tags such as "cheerful" or "sad" from facial images and voice. This integrates the user's behavioral data and emotional data.

[1687] Step 2:

[1688] Script generation

[1689] The server generates a script using a generative AI model (e.g., GPT-3) based on the acquired user information and emotional data. The input is the user's area of ​​interest (e.g., fitness) and emotional data (e.g., "energetic"), and the script text generated based on the prompt ("Generate a cheerful tone script about effective fitness training methods.") is output. This generates text that matches the user's emotions.

[1690] Step 3:

[1691] Image material generation and acquisition

[1692] The server collects or generates related image materials based on the generated script. The input is the script text (e.g., "Explanation of effective fitness training methods in a lively tone"), and the output is a corresponding image file (e.g., a diagram showing training steps). This may involve using an image collection service or image generation AI.

[1693] Step 4:

[1694] Video generation

[1695] The server combines the generated script with image materials to generate a video. It then uses a text-to-speech program to convert the script into audio and synchronize it with the image materials. The input is the script text and image files, and the output is a video file with narration. Google Text-to-Speech (gTTS) is used for speech synthesis.

[1696] Step 5:

[1697] Video distribution

[1698] The server uploads the generated video file to the distribution platform and sends a notification to the user's device. The input is the video file and user information, and the output is the video uploaded to the distribution platform and a notification message to the user (e.g., "A new fitness video is available to watch"). This allows the user to be aware of the existence of the video and watch it.

[1699] Step 6:

[1700] User Viewing

[1701] The user receives a notification sent to their device and watches the video. The input is the notification message and the device, and the output is the video being played and the user can watch the content. This allows the user to enjoy personalized content.

[1702] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1703] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1704] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1705] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1706] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1707] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1708] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1709] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1710] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1711] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1712] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1713] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1714] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1715] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1716] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1717] The hardware resource for executing a specific process can be any of the following types of processors: A processor, for example, is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1718] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1719] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1720] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1721] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1722] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1723] The following is further disclosed regarding the above embodiment.

[1724] (Claim 1)

[1725] A means for acquiring the user's area of ​​interest from the user's past behavioral data and registration information;

[1726] A means for generating a script suitable for each user using a generative AI based on the acquired user's area of ​​interest;

[1727] means for collecting or generating related image material based on the generated script;

[1728] A method of combining a script with image material, adding narration using a reading program, and generating a video;

[1729] A means of distributing the generated video to the platform and notifying the user's device,

[1730] An information distribution system including:

[1731] (Claim 2)

[1732] 2. The information distribution system according to claim 1, further comprising means for a user to generate a script describing newly added items in a certain category.

[1733] (Claim 3)

[1734] 2. The information distribution system according to claim 1, further comprising means for generating local news and weather information for the user's local area as a script.

[1735] "Example 1"

[1736] (Claim 1)

[1737] A means for acquiring the user's area of ​​interest from the user's past behavioral data and registration information;

[1738] A means for generating a script suitable for each user using a generative AI based on the acquired user's area of ​​interest;

[1739] means for collecting or generating related image material based on the generated script;

[1740] A means for combining a script and image material, generating audio and video using a reading program;

[1741] A means to distribute the generated video to a specific platform and notify the user's device,

[1742] A means for viewing the notified content on the user's device;

[1743] A system including:

[1744] (Claim 2)

[1745] 10. The system of claim 1, further comprising means for a user to generate a script that includes newly added items in a category.

[1746] (Claim 3)

[1747] 10. The system of claim 1, further comprising means for generating scripted local news and weather information for a user's local area.

[1748] "Application Example 1"

[1749] (Claim 1)

[1750] A means for acquiring the user's area of ​​interest from the user's past behavioral data and registration information;

[1751] A means for generating a script suitable for each user using a generative AI based on the acquired user's area of ​​interest;

[1752] means for collecting or generating related image material based on the generated script;

[1753] A method of combining a script with image material, adding narration using a reading program, and generating a video;

[1754] A means of distributing the generated video to the platform and notifying the user's device,

[1755] a means for notifying the user when the generated video is available;

[1756] A system including:

[1757] (Claim 2)

[1758] 10. The system of claim 1, further comprising means for a user to generate a script describing newly added items in a category.

[1759] (Claim 3)

[1760] 10. The system of claim 1, further comprising means for generating scripted local news and weather information for a user's local area.

[1761] "Example 2: Combining Emotion Engines"

[1762] (Claim 1)

[1763] A means for obtaining user past behavior data and registration information;

[1764] means for recognizing an emotional state of a user using the acquired data;

[1765] A means for generating a script suitable for each user using a generative AI model based on the acquired user information and emotion data;

[1766] means for collecting or generating related image material based on the generated script;

[1767] A method of combining a script with image material, adding narration using a reading program, and generating a video;

[1768] A means to upload the generated video to a distribution platform and notify the user's device,

[1769] A system including:

[1770] (Claim 2)

[1771] 10. The system of claim 1, further comprising means for a user to generate a script describing newly added items in a category.

[1772] (Claim 3)

[1773] 10. The system of claim 1, further comprising means for generating scripted local news and weather information for a user's local area.

[1774] "Application example 2 when combining emotion engines"

[1775] (Claim 1)

[1776] A means for acquiring the user's area of ​​interest from the user's past behavioral data and registration information;

[1777] A means for generating a script suitable for each user using a generation AI model based on the acquired user's area of ​​interest;

[1778] means for collecting or generating related image material based on the generated script;

[1779] A method of combining a script with image material, adding narration using a reading program, and generating a video;

[1780] A means of distributing the generated video to the platform and notifying the user's device,

[1781] A means of obtaining user emotional data through facial recognition and voice analysis,

[1782] A means to adjust the tone and content of the script based on the acquired emotional data,

[1783] A system including:

[1784] (Claim 2)

[1785] 10. The system of claim 1, further comprising means for a user to generate a script describing newly added items in a category.

[1786] (Claim 3)

[1787] 10. The system of claim 1, further comprising means for generating scripted local news and weather information for a user's local area. [Explanation of symbols]

[1788] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. A means for acquiring the user's area of ​​interest from the user's past behavioral data and registration information; A means for generating a script suitable for each user using a generative AI based on the acquired user's area of ​​interest; means for collecting or generating related image material based on the generated script; A method of combining a script with image material, adding narration using a reading program, and generating a video; A means of distributing the generated video to the platform and notifying the user's device, An information distribution system including:

2. 2. The information distribution system according to claim 1, further comprising means for a user to generate a script describing an item newly added to a certain category.

3. 2. The information distribution system according to claim 1, further comprising means for generating local news and weather information for the user's local area as a script.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A