System
The system addresses the challenge of inefficient content delivery by automating personalized video content generation and distribution based on user interests and local information, enhancing user satisfaction through tailored and regularly updated content.
Patent Information
- Application Number
- JP2024125370
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-31
- Publication Date
- 2026-02-13
AI Technical Summary
Users spend significant time searching for appropriate content and face a lack of personalized content based on their interests and local information, leading to low satisfaction, and manually creating and distributing content is labor-intensive and inefficient for large user bases.
A system that collects user information, generates personalized video content using generative AI models, integrates audio and images, and distributes it to user devices, allowing for automated content creation and delivery tailored to individual interests and geographical locations.
Efficiently delivers personalized video content optimized for each user, reducing the effort required for content creation and improving user satisfaction by leveraging user feedback for continuous improvement.
Smart Images

Figure 2026023435000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] Many users spend a lot of time searching for appropriate content and have to expend time searching for the right source. Furthermore, a lack of personalized content based on user interests and local information leads to low user satisfaction. Furthermore, manually creating and distributing content is labor-intensive and difficult to efficiently serve a large user base. This invention aims to solve these problems by providing a system that automatically generates and regularly distributes short video content optimized for each user. [Means for solving the problem]
[0005] The present invention provides a system that includes a means for collecting user information, a means for generating a script for each user using a generative AI model, a means for generating audio based on the script using a text-to-speech program, a means for generating images based on the script using an image generation AI, a means for generating a video by combining the generated audio and images, and a means for delivering the generated video to a user's device. This makes it possible to automatically generate and efficiently deliver personalized video content based on the user's interests and regional information. This system provides content optimized for each user, reduces the effort required for content creation and delivery, and improves the overall user experience.
[0006] "User Information" refers to various data about a user, including the user's interests and geographical location.
[0007] A "generative AI model" is artificial intelligence software or algorithm that automatically generates text or scripts based on user information.
[0008] A "script" is a text or script that forms the basis of video content, and includes instructions for voice reading and image generation.
[0009] A "text-to-speech program" is software or algorithm that generates speech based on a script.
[0010] "Image generation AI" refers to artificial intelligence software or algorithms that automatically generate images based on the content of a script.
[0011] "Video" refers to multimedia content created by integrating generated audio and images.
[0012] "User device" refers to a digital device owned by a user (e.g., smartphone, tablet, PC, etc.).
[0013] "Distribution" refers to the act of sending the generated video content to the user's device.
[0014] "Platform" refers to the services and applications (e.g., LINE, websites) through which video content is distributed.
[0015] "Interests" are the topics or areas in which a user is particularly interested.
[0016] "Local information" refers to information about the area where the user lives or the area in which they are interested. [Brief explanation of the drawings]
[0017] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0018] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0019] First, the terms used in the following description will be explained.
[0020] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0021] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0022] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0023] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0024] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0025] [First embodiment]
[0026] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0027] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0028] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0029] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0030] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0031] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0032] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0033] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0034] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0035] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0036] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0037] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0038] The system of the present invention automatically generates short video content optimized for each user and distributes it periodically. This system generates content based on the user's interests and local information and distributes it to the user's device, thereby improving user satisfaction.
[0039] Server Operation
[0040] 1. Collection of User Information
[0041] The server collects information about the user's interests and region and stores it in a database. For example, if the user is interested in "technology," the server will retrieve and store that information.
[0042] 2. Script generation using generative AI models
[0043] The server retrieves user information from a database and uses a generative AI model to generate a script based on the user's interests, for example, "Today's Hot Tech News."
[0044] 3. Integration of text-to-speech programs and image generation AI
[0045] Based on the generated script, the server uses a text-to-speech program to generate audio and uses image generation AI to generate images that match the script. The generated audio and images are then integrated to generate a video.
[0046] 4. Video distribution
[0047] The server then delivers the generated video to the user's device via the platform that the user normally uses.
[0048] Device behavior
[0049] 1. User Authentication
[0050] The user terminal goes through an authentication process with the server and provides user information, which allows the server to obtain detailed information about the user.
[0051] 2. Receiving and saving videos
[0052] The user device receives the video delivered from the server and stores it locally, allowing the video to be played offline.
[0053] 3. Play the video
[0054] The user's device plays the stored video to the user, allowing the user to watch the latest information on their area of interest.
[0055] User Actions
[0056] 1. Setting your interests
[0057] Users can set their interests through their devices and send them to the server. For example, they can set interests such as "technology" or "sports."
[0058] 2. Watch the video and give feedback
[0059] After watching a video, users can send feedback about the content to the server via their device, which is reflected in the next content generation.
[0060] Specific examples
[0061] For example, if user A specifies that he is interested in "technology," the server will do the following:
[0062] 1. Collection of User Information
[0063] The server collects user A's interest data and stores it in a database.
[0064] 2. Script generation using generative AI models
[0065] The server retrieves user A's information about "technology" from the database and generates a script using the generative AI model. For example, a script might be generated that reads, "Today's hot technology news: A new AI technology has been announced."
[0066] 3. Integration of text-to-speech programs and image generation AI
[0067] The server inputs the script into a speech program to generate audio.
[0068] Based on the generated script, image generation AI is used to generate related images.
[0069] Integrates audio and images to generate videos.
[0070] 4. Video distribution
[0071] The video generated by the server is distributed to User A's device. For example, distribution can be done using the SNS platform that User A normally uses.
[0072] 5. Receiving and saving videos
[0073] User A's device receives the video and saves it locally.
[0074] 6. Video playback
[0075] User A's device plays the saved video and User A watches it.
[0076] 7. Providing Feedback
[0077] After viewing, User A enters feedback and sends it to the server via their device. For example, feedback such as "The technical content was easy to understand."
[0078] This system enables efficient delivery of content customized for each user, thereby improving user satisfaction.
[0079] The processing flow will be explained below.
[0080] Step 1:
[0081] The server collects user interests and geographical information and stores it in a database. When a user inputs their interests and geographical information, such as "technology" or "Tokyo," through their device, the server receives this information and records it in the user's profile.
[0082] Step 2:
[0083] The server periodically retrieves user information from the database. For example, every morning the system automatically retrieves the interests and geographical locations of all registered users.
[0084] Step 3:
[0085] The server then calls up a generative AI model based on user information to generate a script tailored to each individual user. For example, if a user is interested in "technology news," a script containing a summary of the previous day's new technology news will be generated.
[0086] Step 4:
[0087] The server then passes the generated script to a text-to-speech program to generate audio. For example, the server uses the Google Text-to-Speech API to convert the script into audio data.
[0088] Step 5:
[0089] The server then passes the generated script to an image generation AI to generate images for the video, for example, using an image generation model such as DALL-E to generate images related to the news content.
[0090] Step 6:
[0091] The server then combines the generated audio and images to create a video. The program then creates a slideshow-style video, switching between images as needed against the audio file in the background.
[0092] Step 7:
[0093] The server then distributes the generated video to the platform specified by the user, for example, uploading it to LINE VOOM and obtaining the URL.
[0094] Step 8:
[0095] The server notifies the user's device of the video URL to which the video is being delivered. The user's device receives this notification and saves the video URL locally.
[0096] Step 9:
[0097] The device will use the saved video URL to download the video and store it locally, allowing users to watch the video even without an internet connection.
[0098] Step 10:
[0099] The device plays the stored video for the user to view, and the user operates the device to start the video and view the provided information.
[0100] Step 11:
[0101] After watching a video, users can provide their impressions and feedback via their device, for example, by entering comments such as "The content was easy to understand."
[0102] Step 12:
[0103] The terminal transmits the feedback received from the user to the server, so that the server can utilize the feedback information for generating the next content.
[0104] In this way, the system automatically generates video content customized for each user and distributes it regularly.
[0105] Example 1
[0106] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0107] Conventional short video content distribution systems have difficulty providing personalized content based on user interests and local information. Furthermore, they lack a feedback function to improve the user experience, making it difficult to efficiently deliver individually optimized content.
[0108] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0109] In this invention, the server includes means for collecting user information, means for generating a script for each user using a generative AI model, means for generating audio based on the script using a text-to-speech program, means for generating images based on the script using an image generation AI, means for generating a video by combining the generated audio and images, means for delivering the generated video to the user's terminal, means for receiving and storing the generated video in the user's terminal, means for playing the generated video on the user's terminal, and means for the user to provide feedback after viewing. This makes it possible to efficiently generate and deliver short video content customized for each user and improve the quality of the content by utilizing the feedback.
[0110] "User Information" is data related to an individual user, such as the user's interests and geographical location.
[0111] A "generative AI model" is an artificial intelligence algorithm that automatically generates a script based on user input data.
[0112] A "script" is text generated by a generative AI model that constitutes the content of video content.
[0113] A "reading program" is software that converts text data into speech.
[0114] "Image generation AI" is an artificial intelligence algorithm that automatically generates images based on text and other input data.
[0115] "Video generation" is the process of combining the generated audio and images into a single video file.
[0116] "User device" refers to an electronic device owned by the user, such as a smartphone, tablet, or PC.
[0117] "Video distribution" is the process of sending video data from a server to a user's device.
[0118] "Feedback" refers to reactions such as opinions and impressions that users provide regarding the content they have viewed.
[0119] The system of the present invention automatically generates and periodically distributes short video content optimized for each user. The main components of the system include a server, a terminal, and a user.
[0120] Server Operation
[0121] The server first collects the user's interests and local area information through an API and stores this information in a database. If the user indicates that they are interested in "technology," this information is retrieved and stored by the server. The collected information is then used as prompts to generate scripts using a generative AI model. For example, a prompt such as "Tell me the latest technology-related news" is input into the generative AI model.
[0122] The generative AI model uses a natural language generation model such as GPT-3. Based on this, the server generates a specific script, such as "Today's hot tech news: New AI technology has been announced." Next, a text-to-speech program (such as Google's Text-to-Speech API) generates audio based on this script. At the same time, an image generation AI such as DALL-E or Stable Diffusion generates appropriate images based on the script.
[0123] The generated audio and video are then combined to create a single video. This is done using video editing software (such as FFmpeg). The resulting video can then be distributed via the user's preferred social media platform or a dedicated app. For example, the server could upload the video via the YouTube API and send a link to the user's device.
[0124] Device behavior
[0125] The user device first goes through an authentication process with the server and provides user information. If this authentication is successful, the device can receive the video delivered from the server and store it locally. This allows the user to play the video even offline. When the user presses the "play" button, the stored video is played using video player software.
[0126] User Actions
[0127] Users set their interests through their device. This is done by selecting an area such as "technology" or "sports" from the settings screen within the device's application. Once the selection is complete, the information is immediately sent to the server. After watching the video, users enter their opinions and thoughts in a feedback form and send it to the server. For example, by providing feedback such as "the content of the video was easy to understand," this will be reflected in the next content generation.
[0128] Specific examples
[0129] For example, if User A specifies that he or she is interested in "technology," the procedure is as follows: The server collects User A's interest data in "technology" and stores it in a database. Next, a generative AI model is used to generate a script that reads, "Today's hot technology news: A new AI technology has been announced." The script is input into Google's Text-to-Speech API to generate audio, and DALL-E is used to generate images based on the script, which are then integrated to create a video. The created video is distributed to User A's device via YouTube, and the device saves the video locally. User A plays the saved video, and after watching, sends feedback to the server that "the technical content was easy to understand."
[0130] This system makes it possible to efficiently generate and distribute short video content customized for each user, and to improve the quality of the content by utilizing feedback.
[0131] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0132] Step 1:
[0133] The server collects user information. It receives interest and region information sent by the user from their device via API and stores it in a database. The input is the interest data set by the user (e.g., "technology"), and the output is that data is stored in the database.
[0134] Step 2:
[0135] A script is generated using a generative AI model. The server retrieves the user's interest data from the database and inputs it as a prompt sentence to the generative AI model. The inputs are the "user's interest data" and the "prompt sentence" (e.g., "Tell me the latest technology news"), and the output is a script generated by the generative AI model (e.g., "Today's hot technology news: A new AI technology has been announced").
[0136] Step 3:
[0137] The server inputs the script into a text-to-speech program to generate audio. Specifically, the generated script is input into Google's Text-to-Speech API to obtain audio data. The input at this time is the generated script, and the output is the audio data generated by the text-to-speech program.
[0138] Step 4:
[0139] The server generates an image based on the script. The generated script is input into an image generation AI (DALL-E or Stable Diffusion) to generate a related image. The input is the generated script, and the output is image data generated by the image generation AI.
[0140] Step 5:
[0141] The server combines the audio and image data to generate a video. The audio and image data are input into video editing software (such as FFmpeg) to generate a single video file. The input in this case is the audio and image data, and the output is the generated video file.
[0142] Step 6:
[0143] The server delivers the generated video to the user's device. The generated video file is uploaded to a social media platform (YouTube, Instagram, etc.) and the link is sent to the user's device. The input in this case is the generated video file, and the output is the video link.
[0144] Step 7:
[0145] The user device receives the video and saves it locally. The video is downloaded via the video link sent from the server and saved in local storage. The input is the video link, and the output is the saved video file.
[0146] Step 8:
[0147] The user device plays the stored video. When the user presses the "play" button, the stored video is played using video player software. The input is the stored video file, and the output is the playback of the video.
[0148] Step 9:
[0149] Users provide feedback after watching. After watching a video, users enter their opinions and thoughts in a feedback form and send the data to the server. The input is the user's feedback data, and the output is the feedback information stored on the server.
[0150] (Application example 1)
[0151] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0152] Conventional automatic short video content generation systems, when generating videos optimized for each user, mainly focused on the user's interests and local information, and therefore were unable to adequately stimulate purchasing motivation or promote products in physical stores. In particular, they were unable to create customized content that reflected the customer's purchasing history or real-time interactions in the store, which created challenges in improving the customer experience and promoting sales.
[0153] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0154] In this invention, the server includes means for collecting user information, means for generating a script for each user using a generative AI model, means for generating audio based on the script using a text-to-speech program, means for generating images based on the script using an image generation AI, means for generating a video by combining the generated audio and images, means for delivering the generated video to the user's device and a display in a physical store, means for performing user authentication and setting interest information, and means for collecting feedback after viewing the video. This makes it possible to provide customers visiting a store with a customized promotional video that reflects their purchasing history, interests, and real-time interactions, thereby improving the customer experience and promoting sales.
[0155] "User Information" is data about individual users, including their interests, geographic location, and purchasing history.
[0156] A "generative AI model" is an artificial intelligence model that generates scripts based on user information.
[0157] A "reading program" is software that automatically generates speech based on a generated script.
[0158] "Image generation AI" is an artificial intelligence model that generates relevant images based on a script.
[0159] "User authentication" refers to the means by which a user's identity is verified, which is the prerequisite process for collecting and personalizing user information.
[0160] "Interest information" is data that indicates a user's interests in a particular field.
[0161] "Feedback" refers to the ratings and opinions provided by users after watching a video, and is data that is reflected in the next content generation.
[0162] A "physical store" is a sales facility located in a physical location that offers goods and services.
[0163] A "terminal" is an electronic device that allows a user to receive or send information, including a smartphone, smart glasses, or a display.
[0164] A "display" is a device for displaying visual information, and includes digital signage used in stores.
[0165] This system automatically generates short video content optimized for each user, improving promotions and customer experiences in brick-and-mortar stores. By generating customized videos based on the purchasing history and interests of customers visiting brick-and-mortar stores and providing them via displays or smart glasses, more personalized promotions can be achieved.
[0166] Server Operation
[0167] The server collects user information and generates a script using a generative AI model. Based on the script, a text-to-speech program generates audio and an image generation AI generates images. These are then combined to generate a video, which is then distributed to displays in physical stores and to users' devices.
[0168] Main components
[0169] 1. Collection of User Information
[0170] The server collects user interests, location information, and purchasing history and stores them in a database.
[0171] 2. Script generation using generative AI models
[0172] The server uses a generative AI model to generate a personalized script based on collected user information.
[0173] 3. Integration of text-to-speech programs and image generation AI
[0174] The server inputs the generated script into a speech program to generate audio, and also uses image generation AI to generate related images.
[0175] 4. Video Creation and Distribution
[0176] The server integrates audio and images to generate a video, which is then distributed to displays in physical stores and to users' devices.
[0177] Device behavior
[0178] The device provides user information through an authentication process with the server, receives the video, and stores it locally for offline playback. After watching the video, users can provide feedback, which is reflected in the next content generation.
[0179] Main features
[0180] 1. User Authentication
[0181] The user terminal goes through an authentication process with the server and provides user information.
[0182] 2. Receiving and saving videos
[0183] The user device receives the video delivered from the server and stores it locally.
[0184] 3. Feedback function
[0185] After watching a video, users provide feedback, which is sent to the server and used to generate the next video.
[0186] Hardware and software used
[0187] Server: A Python server application using Flask
[0188] Generative AI models: script generation based on user interests (e.g., GPT-3)
[0189] Text-to-speech program: A text-to-speech tool (e.g., Google Text-to-Speech)
[0190] Image generation AI: Deep learning model (e.g., DALL-E)
[0191] Video generation tools: video editing libraries (e.g., moviepy)
[0192] Devices: Smartphones, smart glasses, in-store displays
[0193] Specific examples
[0194] For example, when a specific customer enters a store, a customized video based on the customer's past purchase history and interests will be displayed on the in-store display, allowing the customer to view new product information and attractive promotions that are tailored to them.
[0195] Example prompt for a generative AI model:
[0196] Customer Interest: "Cooking"
[0197] Prompt for generative AI model: "Generate a short video about the latest cooking appliances and fresh ingredients."
[0198] This system allows physical stores to provide promotional videos optimized for each individual customer, improving the customer experience and promoting sales.
[0199] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0200] Step 1:
[0201] Collection of User Information
[0202] The server collects users' interests, local area information, and purchasing history. The information provided by each user through their device and their in-store purchasing history are stored in a database. This allows the collection of data necessary to generate customized content. The input data are users' interests, local area information, and purchasing history. The output is a database of the collected user information.
[0203] Step 2:
[0204] Script generation using generative AI models
[0205] The server generates a personalized script using a generative AI model based on the collected user information. The script is generated by using data obtained from the user information database as input and providing a prompt to the generative AI model (e.g., "Customer interest: 'Cooking'. Prompt to the generative AI model: 'Generate a short video about the latest cooking equipment and fresh ingredients.'"). The output is the generated customized script.
[0206] Step 3:
[0207] Speech generation by a reading program
[0208] The server inputs the generated script into a text-to-speech program to generate audio. The input data is the generated script, and the output is the generated audio file. The text-to-speech program uses a speech synthesis tool such as Google Text-to-Speech.
[0209] Step 4:
[0210] Image generation using image generation AI
[0211] The server inputs the generated script into the image generation AI, which then generates an image based on it. The input data is the generated script, and the output is a generated image file. The image generation AI model uses a deep learning model such as DALL-E.
[0212] Step 5:
[0213] Video generation
[0214] The server generates a video by integrating the generated audio and images. The input data are the generated audio and image files, and the output is the generated video file. To generate the video, a video editing library such as moviepy is used.
[0215] Step 6:
[0216] Video distribution
[0217] The server distributes the generated video to the user's device and the display in the physical store. The input data is the generated video file, and the output is the video distributed to the user's device and the display in the physical store. Distribution is done via the Internet.
[0218] Step 7:
[0219] User authentication and interest settings
[0220] Users set and update their authentication and interest information on the server through their devices. The input data is the user's authentication information and interest information, and the output is the authenticated user information stored on the server.
[0221] Step 8:
[0222] Providing feedback after watching the video
[0223] Users watch videos through their devices and provide feedback after viewing. The input data is feedback information from the users, and the output is stored in the server's feedback database. This provides data that will be reflected in the next content generation.
[0224] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0225] The system of the present invention automatically generates short video content optimized for each user and distributes it periodically. Furthermore, by combining it with an emotion engine, it can recognize the user's emotions and provide personalized content accordingly, thereby further improving user satisfaction.
[0226] Server Operation
[0227] 1. Collection of User Information
[0228] The server collects user interests and geographical information and stores it in a database. When a user inputs their interests and geographical information, such as "technology" or "Tokyo," through their device, the server receives this information and records it in the user's profile.
[0229] 2. Script generation using generative AI models
[0230] The server retrieves user information from a database and uses a generative AI model to generate a script based on the user's interests, for example, "Today's Hot Tech News."
[0231] 3. Integration of text-to-speech programs and image generation AI
[0232] Based on the generated script, the server uses a text-to-speech program to generate audio and uses image generation AI to generate images that match the script. The generated audio and images are then integrated to generate a video.
[0233] 4. Video distribution
[0234] The server then delivers the generated video to the user's device via the platform the user normally uses, such as a social networking platform or email.
[0235] 5. Gathering Feedback
[0236] After watching a video, users input their impressions and feedback, which are then sent to the server via their device, where they are stored in a database.
[0237] 6. Emotion Recognition by Emotion Engine
[0238] The server passes the collected feedback to the emotion engine to analyze the user's emotions. The emotion engine recognizes the emotions the user felt when giving feedback and stores them in a database.
[0239] 7. Use of Emotional Data
[0240] The server uses the emotion data recognized by the emotion engine to correct the script generation process for the next time. For example, if the user was recognized as "very satisfied" during the previous video viewing, a script with similar content and style will be generated.
[0241] Device behavior
[0242] 1. User Authentication
[0243] The user terminal goes through an authentication process with the server and provides user information, which allows the server to obtain detailed information about the user.
[0244] 2. Receiving and saving videos
[0245] The user's device receives the video streamed from the server and stores it locally, allowing the video to be played even without an internet connection.
[0246] 3. Play the video
[0247] The user device plays the stored video to the user, who then operates the device to start the video and view the provided information.
[0248] 4. Providing Feedback
[0249] After watching the video, users can enter their impressions and feedback through the device. For example, they can enter feedback such as "The content was easy to understand."
[0250] User Actions
[0251] 1. Setting your interests
[0252] Users can set their interests through their devices and send them to the server. For example, they can set interests such as "technology" or "sports."
[0253] 2. Watch the video and give feedback
[0254] After watching a video, users can send feedback about the content to the server via their device, which is reflected in the next content generation.
[0255] Specific examples
[0256] For example, if user A specifies that he is interested in "technology," the server will do the following:
[0257] 1. Collection of User Information
[0258] The server collects user A's interest data and stores it in a database.
[0259] 2. Script generation using generative AI models
[0260] The server retrieves user A's information about "technology" from the database and generates a script using the generative AI model. For example, a script might be generated that reads, "Today's hot technology news: A new AI technology has been announced."
[0261] 3. Integration of text-to-speech programs and image generation AI
[0262] The server inputs the script into a speech program to generate audio.
[0263] Based on the generated script, image generation AI is used to generate related images.
[0264] Integrates audio and images to generate videos.
[0265] 4. Video distribution
[0266] The video generated by the server is distributed to User A's device. For example, distribution can be done using the SNS platform that User A normally uses.
[0267] 5. Gathering Feedback
[0268] User A enters feedback after watching the video. For example, he / she enters a comment including emotion such as "Very satisfied."
[0269] 6. Emotion Recognition by Emotion Engine
[0270] The server passes the feedback to the emotion engine to recognize User A's emotion.
[0271] 7. Use of Emotional Data
[0272] The server then uses the recognized emotion data to correct the script generation process for the next time, so that the next script will be generated in the same style.
[0273] This system makes it possible to further personalize and efficiently deliver content customized for each user based on their emotions, thereby further improving user satisfaction.
[0274] The processing flow will be explained below.
[0275] Step 1:
[0276] The server collects user interests and geographical location information and stores it in a database. For example, if a user enters interests and geographical location such as "technology" or "Tokyo," the server receives this and records it in the user profile.
[0277] Step 2:
[0278] The server periodically retrieves user information from the database. For example, every morning the server can be configured to automatically retrieve user interests and location information.
[0279] Step 3:
[0280] The server then calls up a generative AI model based on user information to generate a script tailored to each individual user. For example, if a user is interested in "technology news," a script containing a summary of the previous day's new technology news will be generated.
[0281] Step 4:
[0282] The server then passes the generated script to a text-to-speech program to generate audio. For example, the server uses the Google Text-to-Speech API to convert the script into audio data.
[0283] Step 5:
[0284] The server then passes the generated script to an image generation AI to generate images for the video, for example, using an image generation model such as DALL-E to generate images related to the news content.
[0285] Step 6:
[0286] The server combines the generated audio and images to create a video, and the program generates a slideshow-style video that displays the images while switching between them against the audio file in the background.
[0287] Step 7:
[0288] The server then distributes the generated video to the platform specified by the user, for example, uploading it to LINE VOOM and obtaining the URL.
[0289] Step 8:
[0290] The server notifies the user's device of the video URL to which the video is being delivered. The user's device receives this notification and saves the video URL locally.
[0291] Step 9:
[0292] The device will use the saved video URL to download the video and store it locally, allowing users to watch the video even without an internet connection.
[0293] Step 10:
[0294] The device plays the stored video for the user to view, and the user operates the device to start the video and view the provided information.
[0295] Step 11:
[0296] After watching a video, users can provide their impressions and feedback via their device, for example, by entering comments such as "The content was easy to understand."
[0297] Step 12:
[0298] The terminal transmits the feedback received from the user to the server, so that the server can utilize the feedback information for generating the next content.
[0299] Step 13:
[0300] The server passes the feedback to an emotion engine to recognize the user's emotion. The emotion engine identifies the user's emotion (e.g., "very satisfied") from the feedback content using, for example, text analysis technology.
[0301] Step 14:
[0302] The server stores the emotion data recognized by the emotion engine in a database, for example, adding emotion data such as "very satisfied" to the user profile.
[0303] Step 15:
[0304] The server uses the recognized emotion data to adjust the script generation process for the next time. For example, if the user was recognized as "very satisfied" during the previous video viewing, the server will generate a script with a similar content style next time.
[0305] Through this process, the system can provide more personalized content based on the user's emotions, thereby improving the quality of the user experience and further increasing satisfaction.
[0306] Example 2
[0307] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0308] In modern content delivery systems, it is important to generate and deliver personalized content optimized for each user. However, conventional systems have difficulty efficiently collecting user sentiment and feedback and reflecting it in the next content generation process. This can easily lead to a decrease in user satisfaction, and more effective personalization is required.
[0309] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0310] In this invention, the server includes means for collecting user information, means for generating a script for each user using a generative AI model, means for generating audio based on the script using a text-to-speech program, means for generating images based on the script using an image generation program, means for generating a video by combining the generated audio and images, means for delivering the generated video to the user's device, means for collecting feedback from the user, means for passing the collected feedback to a sentiment analysis engine to recognize the user's emotions, and means for correcting the next script generation process using the recognized emotion data. This enables the generation and delivery of content optimized for each user, making it possible to provide personalized content that reflects the user's emotions and feedback and provides a high level of satisfaction.
[0311] "User information" refers to data necessary for personalizing content, such as user interests and regional information.
[0312] A "generative AI model" is a model that uses artificial intelligence technology to automatically generate scripts based on user information.
[0313] A "reading program" is a program for generating voice based on a generated script.
[0314] An "image generation program" is a program for generating images based on a generated script.
[0315] "Video" is media content created by combining audio and video.
[0316] "Feedback" refers to the impressions and evaluations provided by users after viewing content.
[0317] An "emotion analysis engine" is an engine that analyzes feedback collected from users and recognizes their emotions.
[0318] "Emotional Data" is data that indicates the emotional state of a user as recognized by an emotion analysis engine.
[0319] The system of the present invention automatically generates short video content optimized for each user and distributes it periodically. The system is mainly composed of three components: a server, a terminal, and a user. The specific operation of each component is explained below.
[0320] Server Operation
[0321] The server first collects user information, including user interests and local area information. For example, if a user inputs their interests such as "technology" or "Tokyo" through their device, the server receives this information and records it in a database.
[0322] The server then uses a generative AI model to generate a script based on the user's information. A natural language generation model such as GPT-4 can be used as the generative AI model. A specific example of a prompt is, "User interest: technology. Please elaborate on the latest technology news in your next script." This prompt causes the AI model to generate a script with appropriate content.
[0323] Based on the generated script, the server uses a text-to-speech program (e.g., Google Text-to-Speech) to generate audio. At the same time, it uses an image generation program (e.g., DALL-E) to generate images related to the script. These audio and images are then combined to generate video content.
[0324] The generated video is then distributed to the user's device by the server via the user's usual social media platform or email. For example, it may be posted on Twitter or Instagram.
[0325] After the user watches the video, the server collects feedback from the user, including their thoughts and feelings, and stores the collected feedback in a database.
[0326] The collected feedback is passed to a sentiment analysis engine (e.g., IBM Watson Emotion Analysis) to recognize the user's emotions. This emotional data is used to adjust the script generation process for the next time. For example, if the feedback for the previous video was "very satisfied," the next time a script will be generated, it will have a similar content and style.
[0327] Device behavior
[0328] The user terminal provides user information through an authentication process with the server, which allows the server to obtain detailed information about the user.
[0329] The user's device receives the video delivered from the server and stores it locally, allowing the video to be played even in an offline environment. The user can operate the device to watch the video and enjoy the provided content.
[0330] After watching the video, the user's device provides an interface to collect feedback from the user, who can then input their impressions and ratings and send them to the server via their device.
[0331] User Actions
[0332] Users can set their interests through their devices, for example by selecting specific interests such as "technology" or "sports," thereby informing the server of their preferences.
[0333] After watching the video, users can input their feedback through their device, which will be reflected in the next content generation, thereby contributing to increasing user satisfaction.
[0334] Specific examples
[0335] For example, if User A specifies that he or she is interested in "technology," the server operates as follows: First, the server collects User A's interest data and stores it in a database. Next, the generative AI model generates a script with content related to "technology." The AI model generates a specific script using the prompt, "User interest: technology. For the next script you create, please detail the latest technology news."
[0336] Based on the script, a reading program generates audio and an image generation program generates related images. The generated audio and images are integrated to create a video, which is then sent to User A's device. After watching the video, User A enters feedback and sends it to the server. The server passes this feedback to an emotion analysis engine, and the recognized emotional data is reflected in the next script generation.
[0337] This makes it possible to generate and deliver personalized content optimized for each user, thereby increasing user satisfaction.
[0338] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0339] System program processing flow
[0340] Step 1: Collect user information
[0341] The server receives information about the user's interests and local area from the device. When the user inputs information such as "technology" or "Tokyo" through the device, the server records this information in a database.
[0342] Input: User interests and location information
[0343] Output: User profile information stored in the database
[0344] Specific operation: The user selects "technology" on the device settings screen, and the information is sent to the server and stored in a database.
[0345] Step 2: Generative AI model generates script
[0346] The server sends prompts to the generative AI model based on the collected user information to generate a script based on the user's interests. For example, "User interest: technology. For your next script, please detail the latest technology news."
[0347] Input: User information retrieved from the database, prompt text
[0348] Output: Generated script
[0349] How it works: The server inputs the prompt "User interest: Technology" into a generative AI model (e.g., GPT-4), and the model generates a related news script.
[0350] Step 3: Integrating the Text-to-Speech Program and the Image Generator
[0351] The server generates audio based on the generated script using a reading program and generates associated images using an image generation program.
[0352] Input: Generated script
[0353] Output: Generated audio and image files
[0354] Specific operation: The server inputs the script into a speech program (e.g., Google Text-to-Speech) to generate an audio file. In parallel, the server inputs the script into an image generation program (e.g., DALL-E) to generate related images.
[0355] Step 4: Audio and video integration
[0356] The server combines the generated audio and images to create a video.
[0357] Input: Generated audio and image files
[0358] Output: Finished video file
[0359] Specific operation: The server places the audio and image files on the timeline and exports them as a single video file.
[0360] Step 5: Publish your video
[0361] The server then distributes the generated video to the user's device via the user's usual social media platform or email.
[0362] Input: Finished video file
[0363] Output: Video delivered to the user's device
[0364] Specific operation: The video file generated by the server is posted to Twitter or Instagram, or sent by email.
[0365] Step 6: Gather feedback
[0366] After watching a video, users can enter their feedback and send it to the server via their device, which then stores it in a database.
[0367] Input: User feedback
[0368] Output: Feedback stored in a database
[0369] Specific operation: The user enters feedback such as "very satisfied" on the device, and the server receives it and records it in the database.
[0370] Step 7: Emotion Recognition with the Emotion Engine
[0371] The server passes the collected feedback to an emotion engine to analyze the user's emotions.
[0372] Input: User feedback data
[0373] Output: Recognized emotion data
[0374] Specific operation: The server passes the feedback "very satisfied" to an emotion engine (e.g., IBM Watson Emotion Analysis), which then recognizes the emotion "satisfied" and stores it in a database.
[0375] Step 8: Use emotion data
[0376] The server uses the emotion data obtained from the emotion engine to correct the next script generation process.
[0377] Input: Recognized emotion data
[0378] Output: Corrected script generation process
[0379] Specific operation: The server sets the conditions for the next script generation based on the "satisfaction" information and reflects this when sending a new prompt sentence to the generation AI model.
[0380] (Application example 2)
[0381] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0382] In today's world, content distribution services must provide optimal content for each individual user. However, with conventional technologies, content generated based on user interests and local information does not adequately reflect user emotions or real-time feedback. This makes it difficult to improve user satisfaction, and there is a need to further increase engagement for each user.
[0383] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for collecting user information, means for generating a script for each user using a generation AI model, means for generating audio based on the script using a text-to-speech program, means for generating images based on the script using an image generation AI, means for generating a video by combining the generated audio and images, means for delivering the generated video to the user's device, means for recognizing the user's emotion using an emotion recognition engine, and means for reflecting the recognized emotion data in the next content generation. This makes it possible to deliver personalized short video content based not only on the user's interests and regional information but also on their emotions.
[0384] "Means for collecting user information" refers to a mechanism for collecting data including user interests and regional information and storing it on a server.
[0385] A "generative AI model" is a system that includes an algorithm that automatically generates the optimal script based on each user's interests and regional information.
[0386] A "means for generating a script" is a process that uses a generative AI model to create a script appropriate for a particular user.
[0387] A "reading program" is software that automatically generates speech based on a generated script.
[0388] "Image generation AI" is an artificial intelligence system that automatically creates images that match the content of a generated script.
[0389] "Means for generating video" refers to technology that integrates generated audio and images to create a single video.
[0390] "Means for delivering video" refers to the infrastructure for transmitting the generated video to the user's device.
[0391] An "emotion recognition engine" is a system that includes an algorithm for analyzing and recognizing emotions from a user's facial expressions, voice, etc.
[0392] "Means for reflecting emotional data in the next content generation" refers to a process of using the recognized emotional data of the user to provide appropriate feedback to the script or content to be generated next.
[0393] This invention is a system that automatically generates short video content optimized for each user and distributes it periodically. Furthermore, by combining it with an emotion engine, it recognizes the user's emotions and provides personalized content accordingly. The various means required for this and their specific processing steps are described below.
[0394] Collection of User Information
[0395] The server first collects the user's interests and area information. The user enters their interests and area of residence through their device, and the information is sent to the server. This information is stored in a database and used for subsequent processing.
[0396] Script generation
[0397] The server uses the collected user information to generate a script using a generative AI model. This generative AI model inputs the user's interests and local information as prompts and outputs a short script based on that data. For example, a prompt might be, "The user's interest is 'technology.' Please generate a script to deliver today's latest technology news to him."
[0398] Sound and image generation
[0399] The server inputs the generated script into a text-to-speech program to generate audio. It then uses an image generation AI to generate images based on the script. Specifically, the text-to-speech program and image generation AI use a TextToSpeech engine and an ImageGenerator.
[0400] Video Creation and Delivery
[0401] The server combines the generated audio and video to generate a video, which is then distributed to the user's device. The video is generated using a video editing software library and distributed via social media platforms or email.
[0402] Emotion recognition and feedback collection
[0403] After the user watches the video, the device uses an emotion recognition engine to analyze the user's emotions. This engine recognizes emotions in real time through facial expressions and voice analysis, and sends the data to the server. For example, it collects emotional feedback such as "very satisfied."
[0404] Reflection in next script generation
[0405] The server reflects the collected emotional data in the next content generation, correcting the script generation process based on the previous feedback and generating content that is in line with the user's preferences and emotions.
[0406] Thus, the present invention can significantly improve user satisfaction by efficiently generating and delivering personalized short video content that takes user emotions into consideration.
[0407] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0408] Step 1: Collect user information
[0409] Users input their interests and local area information through their devices and send that information to the server. The server stores the received data in a database. This input information includes the user's content preferences and location information. For example, a user might input information that they are interested in "technology" and live in "Tokyo."
[0410] Step 2: Generate the script
[0411] The server generates a script using a generative AI model based on the user information collected in step 1. In this process, the server inputs a prompt such as "The user's interest is 'technology'. Please generate a script that conveys today's latest technology news to him" into the generative AI model, and obtains the optimal script as the output.
[0412] Step 3: Generate audio
[0413] The server inputs the script generated in step 2 into a speech program (such as a TextToSpeech engine) to automatically generate speech. Based on the input, acoustic analysis and synthesis are performed, and the output is a generated audio file.
[0414] Step 4: Generate images
[0415] The server inputs the generated script into an image generation AI (such as ImageGenerator) to generate an image related to the script. The image generation AI analyzes the contents of the script and generates and outputs an appropriate image. For example, it generates an image related to an article about the latest AI technology.
[0416] Step 5: Generate the video
[0417] The server generates a video by combining the audio generated in step 3 with the images generated in step 4. This process uses a video editing software library to combine the audio and images in a sequence and output a single video file containing technology news.
[0418] Step 6: Publish your video
[0419] The server distributes the generated video to the user's device. Distribution is done via the user's usual social media platform or email. Based on the distribution destination information, the server sends the video file to the corresponding platform. The user then watches the video using their device.
[0420] Step 7: Recognize emotions
[0421] While the user is watching a video, the device uses a built-in emotion recognition engine to analyze the user's emotions in real time. This process involves using a camera and microphone to collect the user's facial expressions and tone of voice, and then inputting this data into the emotion recognition algorithm. The output is the user's emotional data.
[0422] Step 8: Gather feedback
[0423] After viewing, users input their feedback and emotional data into their device and send it to the server. The server stores this in a database. The collected feedback is reflected in the next content generation process. For example, emotional feedback such as "very satisfied" can be input.
[0424] Step 9: Reflecting the changes in the next script generation
[0425] The server uses the emotion data collected in step 8 to correct the next content generation process. This ensures that the next script generated is in line with the user's preferences and emotions. This makes it possible to continuously improve user satisfaction.
[0426] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0427] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0428] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0429] [Second embodiment]
[0430] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0431] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0432] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0433] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0434] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0435] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0436] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0437] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0438] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0439] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0440] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0441] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0442] The system of the present invention automatically generates short video content optimized for each user and distributes it periodically. This system generates content based on the user's interests and local information and distributes it to the user's device, thereby improving user satisfaction.
[0443] Server Operation
[0444] 1. Collection of User Information
[0445] The server collects information about the user's interests and region and stores it in a database. For example, if the user is interested in "technology," the server will retrieve and store that information.
[0446] 2. Script generation using generative AI models
[0447] The server retrieves user information from a database and uses a generative AI model to generate a script based on the user's interests, for example, "Today's Hot Tech News."
[0448] 3. Integration of text-to-speech programs and image generation AI
[0449] Based on the generated script, the server uses a text-to-speech program to generate audio and uses image generation AI to generate images that match the script. The generated audio and images are then integrated to generate a video.
[0450] 4. Video distribution
[0451] The server then delivers the generated video to the user's device via the platform that the user normally uses.
[0452] Device behavior
[0453] 1. User Authentication
[0454] The user terminal goes through an authentication process with the server and provides user information, which allows the server to obtain detailed information about the user.
[0455] 2. Receiving and saving videos
[0456] The user device receives the video delivered from the server and stores it locally, allowing the video to be played offline.
[0457] 3. Play the video
[0458] The user's device plays the stored video to the user, allowing the user to watch the latest information on their area of interest.
[0459] User Actions
[0460] 1. Setting your interests
[0461] Users can set their interests through their devices and send them to the server. For example, they can set interests such as "technology" or "sports."
[0462] 2. Watch the video and give feedback
[0463] After watching a video, users can send feedback about the content to the server via their device, which is reflected in the next content generation.
[0464] Specific examples
[0465] For example, if user A specifies that he is interested in "technology," the server will do the following:
[0466] 1. Collection of User Information
[0467] The server collects user A's interest data and stores it in a database.
[0468] 2. Script generation using generative AI models
[0469] The server retrieves user A's information about "technology" from the database and generates a script using the generative AI model. For example, a script might be generated that reads, "Today's hot technology news: A new AI technology has been announced."
[0470] 3. Integration of text-to-speech programs and image generation AI
[0471] The server inputs the script into a speech program to generate audio.
[0472] Based on the generated script, image generation AI is used to generate related images.
[0473] Integrates audio and images to generate videos.
[0474] 4. Video distribution
[0475] The video generated by the server is distributed to User A's device. For example, distribution can be done using the SNS platform that User A normally uses.
[0476] 5. Receiving and saving videos
[0477] User A's device receives the video and saves it locally.
[0478] 6. Video playback
[0479] User A's device plays the saved video and User A watches it.
[0480] 7. Providing Feedback
[0481] After viewing, User A enters feedback and sends it to the server via their device. For example, feedback such as "The technical content was easy to understand."
[0482] This system enables efficient delivery of content customized for each user, thereby improving user satisfaction.
[0483] The processing flow will be explained below.
[0484] Step 1:
[0485] The server collects user interests and geographical information and stores it in a database. When a user inputs their interests and geographical information, such as "technology" or "Tokyo," through their device, the server receives this information and records it in the user's profile.
[0486] Step 2:
[0487] The server periodically retrieves user information from the database. For example, every morning the system automatically retrieves the interests and geographical locations of all registered users.
[0488] Step 3:
[0489] The server then calls up a generative AI model based on user information to generate a script tailored to each individual user. For example, if a user is interested in "technology news," a script containing a summary of the previous day's new technology news will be generated.
[0490] Step 4:
[0491] The server then passes the generated script to a text-to-speech program to generate audio. For example, the server uses the Google Text-to-Speech API to convert the script into audio data.
[0492] Step 5:
[0493] The server then passes the generated script to an image generation AI to generate images for the video, for example, using an image generation model such as DALL-E to generate images related to the news content.
[0494] Step 6:
[0495] The server then combines the generated audio and images to create a video. The program then creates a slideshow-style video, switching between images as needed against the audio file in the background.
[0496] Step 7:
[0497] The server then distributes the generated video to the platform specified by the user, for example, uploading it to LINE VOOM and obtaining the URL.
[0498] Step 8:
[0499] The server notifies the user's device of the video URL to which the video is being delivered. The user's device receives this notification and saves the video URL locally.
[0500] Step 9:
[0501] The device will use the saved video URL to download the video and store it locally, allowing users to watch the video even without an internet connection.
[0502] Step 10:
[0503] The device plays the stored video for the user to view, and the user operates the device to start the video and view the provided information.
[0504] Step 11:
[0505] After watching a video, users can provide their impressions and feedback via their device, for example, by entering comments such as "The content was easy to understand."
[0506] Step 12:
[0507] The terminal transmits the feedback received from the user to the server, so that the server can utilize the feedback information for generating the next content.
[0508] In this way, the system automatically generates video content customized for each user and distributes it regularly.
[0509] Example 1
[0510] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0511] Conventional short video content distribution systems have difficulty providing personalized content based on user interests and local information. Furthermore, they lack a feedback function to improve the user experience, making it difficult to efficiently deliver individually optimized content.
[0512] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0513] In this invention, the server includes means for collecting user information, means for generating a script for each user using a generative AI model, means for generating audio based on the script using a text-to-speech program, means for generating images based on the script using an image generation AI, means for generating a video by combining the generated audio and images, means for delivering the generated video to the user's terminal, means for receiving and storing the generated video in the user's terminal, means for playing the generated video on the user's terminal, and means for the user to provide feedback after viewing. This makes it possible to efficiently generate and deliver short video content customized for each user and improve the quality of the content by utilizing the feedback.
[0514] "User Information" is data related to an individual user, such as the user's interests and geographical location.
[0515] A "generative AI model" is an artificial intelligence algorithm that automatically generates a script based on user input data.
[0516] A "script" is text generated by a generative AI model that constitutes the content of video content.
[0517] A "reading program" is software that converts text data into speech.
[0518] "Image generation AI" is an artificial intelligence algorithm that automatically generates images based on text and other input data.
[0519] "Video generation" is the process of combining the generated audio and images into a single video file.
[0520] "User device" refers to an electronic device owned by the user, such as a smartphone, tablet, or PC.
[0521] "Video distribution" is the process of sending video data from a server to a user's device.
[0522] "Feedback" refers to reactions such as opinions and impressions that users provide regarding the content they have viewed.
[0523] The system of the present invention automatically generates and periodically distributes short video content optimized for each user. The main components of the system include a server, a terminal, and a user.
[0524] Server Operation
[0525] The server first collects the user's interests and local area information through an API and stores this information in a database. If the user indicates that they are interested in "technology," this information is retrieved and stored by the server. The collected information is then used as prompts to generate scripts using a generative AI model. For example, a prompt such as "Tell me the latest technology-related news" is input into the generative AI model.
[0526] The generative AI model uses a natural language generation model such as GPT-3. Based on this, the server generates a specific script, such as "Today's hot tech news: New AI technology has been announced." Next, a text-to-speech program (such as Google's Text-to-Speech API) generates audio based on this script. At the same time, an image generation AI such as DALL-E or Stable Diffusion generates appropriate images based on the script.
[0527] The generated audio and video are then combined to create a single video. This is done using video editing software (such as FFmpeg). The resulting video can then be distributed via the user's preferred social media platform or a dedicated app. For example, the server could upload the video via the YouTube API and send a link to the user's device.
[0528] Device behavior
[0529] The user device first goes through an authentication process with the server and provides user information. If this authentication is successful, the device can receive the video delivered from the server and store it locally. This allows the user to play the video even offline. When the user presses the "play" button, the stored video is played using video player software.
[0530] User Actions
[0531] Users set their interests through their device. This is done by selecting an area such as "technology" or "sports" from the settings screen within the device's application. Once the selection is complete, the information is immediately sent to the server. After watching the video, users enter their opinions and thoughts in a feedback form and send it to the server. For example, by providing feedback such as "the content of the video was easy to understand," this will be reflected in the next content generation.
[0532] Specific examples
[0533] For example, if User A specifies that he or she is interested in "technology," the procedure is as follows: The server collects User A's interest data in "technology" and stores it in a database. Next, a generative AI model is used to generate a script that reads, "Today's hot technology news: A new AI technology has been announced." The script is input into Google's Text-to-Speech API to generate audio, and DALL-E is used to generate images based on the script, which are then integrated to create a video. The created video is distributed to User A's device via YouTube, and the device saves the video locally. User A plays the saved video, and after watching, sends feedback to the server that "the technical content was easy to understand."
[0534] This system makes it possible to efficiently generate and distribute short video content customized for each user, and to improve the quality of the content by utilizing feedback.
[0535] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0536] Step 1:
[0537] The server collects user information. It receives interest and region information sent by the user from their device via API and stores it in a database. The input is the interest data set by the user (e.g., "technology"), and the output is that data is stored in the database.
[0538] Step 2:
[0539] A script is generated using a generative AI model. The server retrieves the user's interest data from the database and inputs it as a prompt sentence to the generative AI model. The inputs are the "user's interest data" and the "prompt sentence" (e.g., "Tell me the latest technology news"), and the output is a script generated by the generative AI model (e.g., "Today's hot technology news: A new AI technology has been announced").
[0540] Step 3:
[0541] The server inputs the script into a text-to-speech program to generate audio. Specifically, the generated script is input into Google's Text-to-Speech API to obtain audio data. The input at this time is the generated script, and the output is the audio data generated by the text-to-speech program.
[0542] Step 4:
[0543] The server generates an image based on the script. The generated script is input into an image generation AI (DALL-E or Stable Diffusion) to generate a related image. The input is the generated script, and the output is image data generated by the image generation AI.
[0544] Step 5:
[0545] The server combines the audio and image data to generate a video. The audio and image data are input into video editing software (such as FFmpeg) to generate a single video file. The input in this case is the audio and image data, and the output is the generated video file.
[0546] Step 6:
[0547] The server delivers the generated video to the user's device. The generated video file is uploaded to a social media platform (YouTube, Instagram, etc.) and the link is sent to the user's device. The input in this case is the generated video file, and the output is the video link.
[0548] Step 7:
[0549] The user device receives the video and saves it locally. The video is downloaded via the video link sent from the server and saved in local storage. The input is the video link, and the output is the saved video file.
[0550] Step 8:
[0551] The user device plays the stored video. When the user presses the "play" button, the stored video is played using video player software. The input is the stored video file, and the output is the playback of the video.
[0552] Step 9:
[0553] Users provide feedback after watching. After watching a video, users enter their opinions and thoughts in a feedback form and send the data to the server. The input is the user's feedback data, and the output is the feedback information stored on the server.
[0554] (Application example 1)
[0555] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0556] Conventional automatic short video content generation systems, when generating videos optimized for each user, mainly focused on the user's interests and local information, and therefore were unable to adequately stimulate purchasing motivation or promote products in physical stores. In particular, they were unable to create customized content that reflected the customer's purchasing history or real-time interactions in the store, which created challenges in improving the customer experience and promoting sales.
[0557] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0558] In this invention, the server includes means for collecting user information, means for generating a script for each user using a generative AI model, means for generating audio based on the script using a text-to-speech program, means for generating images based on the script using an image generation AI, means for generating a video by combining the generated audio and images, means for delivering the generated video to the user's device and a display in a physical store, means for performing user authentication and setting interest information, and means for collecting feedback after viewing the video. This makes it possible to provide customers visiting a store with a customized promotional video that reflects their purchasing history, interests, and real-time interactions, thereby improving the customer experience and promoting sales.
[0559] "User Information" is data about individual users, including their interests, geographic location, and purchasing history.
[0560] A "generative AI model" is an artificial intelligence model that generates scripts based on user information.
[0561] A "reading program" is software that automatically generates speech based on a generated script.
[0562] "Image generation AI" is an artificial intelligence model that generates relevant images based on a script.
[0563] "User authentication" refers to the means by which a user's identity is verified, which is the prerequisite process for collecting and personalizing user information.
[0564] "Interest information" is data that indicates a user's interests in a particular field.
[0565] "Feedback" refers to the ratings and opinions provided by users after watching a video, and is data that is reflected in the next content generation.
[0566] A "physical store" is a sales facility located in a physical location that offers goods and services.
[0567] A "terminal" is an electronic device that allows a user to receive or send information, including a smartphone, smart glasses, or a display.
[0568] A "display" is a device for displaying visual information, and includes digital signage used in stores.
[0569] This system automatically generates short video content optimized for each user, improving promotions and customer experiences in brick-and-mortar stores. By generating customized videos based on the purchasing history and interests of customers visiting brick-and-mortar stores and providing them via displays or smart glasses, more personalized promotions can be achieved.
[0570] Server Operation
[0571] The server collects user information and generates a script using a generative AI model. Based on the script, a text-to-speech program generates audio and an image generation AI generates images. These are then combined to generate a video, which is then distributed to displays in physical stores and to users' devices.
[0572] Main components
[0573] 1. Collection of User Information
[0574] The server collects user interests, location information, and purchasing history and stores them in a database.
[0575] 2. Script generation using generative AI models
[0576] The server uses a generative AI model to generate a personalized script based on collected user information.
[0577] 3. Integration of text-to-speech programs and image generation AI
[0578] The server inputs the generated script into a speech program to generate audio, and also uses image generation AI to generate related images.
[0579] 4. Video Creation and Distribution
[0580] The server integrates audio and images to generate a video, which is then distributed to displays in physical stores and to users' devices.
[0581] Device behavior
[0582] The device provides user information through an authentication process with the server, receives the video, and stores it locally for offline playback. After watching the video, users can provide feedback, which is reflected in the next content generation.
[0583] Main features
[0584] 1. User Authentication
[0585] The user terminal goes through an authentication process with the server and provides user information.
[0586] 2. Receiving and saving videos
[0587] The user device receives the video delivered from the server and stores it locally.
[0588] 3. Feedback function
[0589] After watching a video, users provide feedback, which is sent to the server and used to generate the next video.
[0590] Hardware and software used
[0591] Server: A Python server application using Flask
[0592] Generative AI models: script generation based on user interests (e.g., GPT-3)
[0593] Text-to-speech program: A text-to-speech tool (e.g., Google Text-to-Speech)
[0594] Image generation AI: Deep learning model (e.g., DALL-E)
[0595] Video generation tools: video editing libraries (e.g., moviepy)
[0596] Devices: Smartphones, smart glasses, in-store displays
[0597] Specific examples
[0598] For example, when a specific customer enters a store, a customized video based on the customer's past purchase history and interests will be displayed on the in-store display, allowing the customer to view new product information and attractive promotions that are tailored to them.
[0599] Example prompt for a generative AI model:
[0600] Customer Interest: "Cooking"
[0601] Prompt for generative AI model: "Generate a short video about the latest cooking appliances and fresh ingredients."
[0602] This system allows physical stores to provide promotional videos optimized for each individual customer, improving the customer experience and promoting sales.
[0603] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0604] Step 1:
[0605] Collection of User Information
[0606] The server collects users' interests, local area information, and purchasing history. The information provided by each user through their device and their in-store purchasing history are stored in a database. This allows the collection of data necessary to generate customized content. The input data are users' interests, local area information, and purchasing history. The output is a database of the collected user information.
[0607] Step 2:
[0608] Script generation using generative AI models
[0609] The server generates a personalized script using a generative AI model based on the collected user information. The script is generated by using data obtained from the user information database as input and providing a prompt to the generative AI model (e.g., "Customer interest: 'Cooking'. Prompt to the generative AI model: 'Generate a short video about the latest cooking equipment and fresh ingredients.'"). The output is the generated customized script.
[0610] Step 3:
[0611] Speech generation by a reading program
[0612] The server inputs the generated script into a text-to-speech program to generate audio. The input data is the generated script, and the output is the generated audio file. The text-to-speech program uses a speech synthesis tool such as Google Text-to-Speech.
[0613] Step 4:
[0614] Image generation using image generation AI
[0615] The server inputs the generated script into the image generation AI, which then generates an image based on it. The input data is the generated script, and the output is a generated image file. The image generation AI model uses a deep learning model such as DALL-E.
[0616] Step 5:
[0617] Video generation
[0618] The server generates a video by integrating the generated audio and images. The input data are the generated audio and image files, and the output is the generated video file. To generate the video, a video editing library such as moviepy is used.
[0619] Step 6:
[0620] Video distribution
[0621] The server distributes the generated video to the user's device and the display in the physical store. The input data is the generated video file, and the output is the video distributed to the user's device and the display in the physical store. Distribution is done via the Internet.
[0622] Step 7:
[0623] User authentication and interest settings
[0624] Users set and update their authentication and interest information on the server through their devices. The input data is the user's authentication information and interest information, and the output is the authenticated user information stored on the server.
[0625] Step 8:
[0626] Providing feedback after watching the video
[0627] Users watch videos through their devices and provide feedback after viewing. The input data is feedback information from the users, and the output is stored in the server's feedback database. This provides data that will be reflected in the next content generation.
[0628] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0629] The system of the present invention automatically generates short video content optimized for each user and distributes it periodically. Furthermore, by combining it with an emotion engine, it can recognize the user's emotions and provide personalized content accordingly, thereby further improving user satisfaction.
[0630] Server Operation
[0631] 1. Collection of User Information
[0632] The server collects user interests and geographical information and stores it in a database. When a user inputs their interests and geographical information, such as "technology" or "Tokyo," through their device, the server receives this information and records it in the user's profile.
[0633] 2. Script generation using generative AI models
[0634] The server retrieves user information from a database and uses a generative AI model to generate a script based on the user's interests, for example, "Today's Hot Tech News."
[0635] 3. Integration of text-to-speech programs and image generation AI
[0636] Based on the generated script, the server uses a text-to-speech program to generate audio and uses image generation AI to generate images that match the script. The generated audio and images are then integrated to generate a video.
[0637] 4. Video distribution
[0638] The server then delivers the generated video to the user's device via the platform the user normally uses, such as a social networking platform or email.
[0639] 5. Gathering Feedback
[0640] After watching a video, users input their impressions and feedback, which are then sent to the server via their device, where they are stored in a database.
[0641] 6. Emotion Recognition by Emotion Engine
[0642] The server passes the collected feedback to the emotion engine to analyze the user's emotions. The emotion engine recognizes the emotions the user felt when giving feedback and stores them in a database.
[0643] 7. Use of Emotional Data
[0644] The server uses the emotion data recognized by the emotion engine to correct the script generation process for the next time. For example, if the user was recognized as "very satisfied" during the previous video viewing, a script with similar content and style will be generated.
[0645] Device behavior
[0646] 1. User Authentication
[0647] The user terminal goes through an authentication process with the server and provides user information, which allows the server to obtain detailed information about the user.
[0648] 2. Receiving and saving videos
[0649] The user's device receives the video streamed from the server and stores it locally, allowing the video to be played even without an internet connection.
[0650] 3. Play the video
[0651] The user device plays the stored video to the user, who then operates the device to start the video and view the provided information.
[0652] 4. Providing Feedback
[0653] After watching the video, users can enter their impressions and feedback through the device. For example, they can enter feedback such as "The content was easy to understand."
[0654] User Actions
[0655] 1. Setting your interests
[0656] Users can set their interests through their devices and send them to the server. For example, they can set interests such as "technology" or "sports."
[0657] 2. Watch the video and give feedback
[0658] After watching a video, users can send feedback about the content to the server via their device, which is reflected in the next content generation.
[0659] Specific examples
[0660] For example, if user A specifies that he is interested in "technology," the server will do the following:
[0661] 1. Collection of User Information
[0662] The server collects user A's interest data and stores it in a database.
[0663] 2. Script generation using generative AI models
[0664] The server retrieves user A's information about "technology" from the database and generates a script using the generative AI model. For example, a script might be generated that reads, "Today's hot technology news: A new AI technology has been announced."
[0665] 3. Integration of text-to-speech programs and image generation AI
[0666] The server inputs the script into a speech program to generate audio.
[0667] Based on the generated script, image generation AI is used to generate related images.
[0668] Integrates audio and images to generate videos.
[0669] 4. Video distribution
[0670] The video generated by the server is distributed to User A's device. For example, distribution can be done using the SNS platform that User A normally uses.
[0671] 5. Gathering Feedback
[0672] User A enters feedback after watching the video. For example, he / she enters a comment including emotion such as "Very satisfied."
[0673] 6. Emotion Recognition by Emotion Engine
[0674] The server passes the feedback to the emotion engine to recognize User A's emotion.
[0675] 7. Use of Emotional Data
[0676] The server then uses the recognized emotion data to correct the script generation process for the next time, so that the next script will be generated in the same style.
[0677] This system makes it possible to further personalize and efficiently deliver content customized for each user based on their emotions, thereby further improving user satisfaction.
[0678] The processing flow will be explained below.
[0679] Step 1:
[0680] The server collects user interests and geographical location information and stores it in a database. For example, if a user enters interests and geographical location such as "technology" or "Tokyo," the server receives this and records it in the user profile.
[0681] Step 2:
[0682] The server periodically retrieves user information from the database. For example, every morning the server can be configured to automatically retrieve user interests and location information.
[0683] Step 3:
[0684] The server then calls up a generative AI model based on user information to generate a script tailored to each individual user. For example, if a user is interested in "technology news," a script containing a summary of the previous day's new technology news will be generated.
[0685] Step 4:
[0686] The server then passes the generated script to a text-to-speech program to generate audio. For example, the server uses the Google Text-to-Speech API to convert the script into audio data.
[0687] Step 5:
[0688] The server then passes the generated script to an image generation AI to generate images for the video, for example, using an image generation model such as DALL-E to generate images related to the news content.
[0689] Step 6:
[0690] The server combines the generated audio and images to create a video, and the program generates a slideshow-style video that displays the images while switching between them against the audio file in the background.
[0691] Step 7:
[0692] The server then distributes the generated video to the platform specified by the user, for example, uploading it to LINE VOOM and obtaining the URL.
[0693] Step 8:
[0694] The server notifies the user's device of the video URL to which the video is being delivered. The user's device receives this notification and saves the video URL locally.
[0695] Step 9:
[0696] The device will use the saved video URL to download the video and store it locally, allowing users to watch the video even without an internet connection.
[0697] Step 10:
[0698] The device plays the stored video for the user to view, and the user operates the device to start the video and view the provided information.
[0699] Step 11:
[0700] After watching a video, users can provide their impressions and feedback via their device, for example, by entering comments such as "The content was easy to understand."
[0701] Step 12:
[0702] The terminal transmits the feedback received from the user to the server, so that the server can utilize the feedback information for generating the next content.
[0703] Step 13:
[0704] The server passes the feedback to an emotion engine to recognize the user's emotion. The emotion engine identifies the user's emotion (e.g., "very satisfied") from the feedback content using, for example, text analysis technology.
[0705] Step 14:
[0706] The server stores the emotion data recognized by the emotion engine in a database, for example, adding emotion data such as "very satisfied" to the user profile.
[0707] Step 15:
[0708] The server uses the recognized emotion data to adjust the script generation process for the next time. For example, if the user was recognized as "very satisfied" during the previous video viewing, the server will generate a script with a similar content style next time.
[0709] Through this process, the system can provide more personalized content based on the user's emotions, thereby improving the quality of the user experience and further increasing satisfaction.
[0710] Example 2
[0711] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0712] In modern content delivery systems, it is important to generate and deliver personalized content optimized for each user. However, conventional systems have difficulty efficiently collecting user sentiment and feedback and reflecting it in the next content generation process. This can easily lead to a decrease in user satisfaction, and more effective personalization is required.
[0713] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0714] In this invention, the server includes means for collecting user information, means for generating a script for each user using a generative AI model, means for generating audio based on the script using a text-to-speech program, means for generating images based on the script using an image generation program, means for generating a video by combining the generated audio and images, means for delivering the generated video to the user's device, means for collecting feedback from the user, means for passing the collected feedback to a sentiment analysis engine to recognize the user's emotions, and means for correcting the next script generation process using the recognized emotion data. This enables the generation and delivery of content optimized for each user, making it possible to provide personalized content that reflects the user's emotions and feedback and provides a high level of satisfaction.
[0715] "User information" refers to data necessary for personalizing content, such as user interests and regional information.
[0716] A "generative AI model" is a model that uses artificial intelligence technology to automatically generate scripts based on user information.
[0717] A "reading program" is a program for generating voice based on a generated script.
[0718] An "image generation program" is a program for generating images based on a generated script.
[0719] "Video" is media content created by combining audio and video.
[0720] "Feedback" refers to the impressions and evaluations provided by users after viewing content.
[0721] An "emotion analysis engine" is an engine that analyzes feedback collected from users and recognizes their emotions.
[0722] "Emotional Data" is data that indicates the emotional state of a user as recognized by an emotion analysis engine.
[0723] The system of the present invention automatically generates short video content optimized for each user and distributes it periodically. The system is mainly composed of three components: a server, a terminal, and a user. The specific operation of each component is explained below.
[0724] Server Operation
[0725] The server first collects user information, including user interests and local area information. For example, if a user inputs their interests such as "technology" or "Tokyo" through their device, the server receives this information and records it in a database.
[0726] The server then uses a generative AI model to generate a script based on the user's information. A natural language generation model such as GPT-4 can be used as the generative AI model. A specific example of a prompt is, "User interest: technology. Please elaborate on the latest technology news in your next script." This prompt causes the AI model to generate a script with appropriate content.
[0727] Based on the generated script, the server uses a text-to-speech program (e.g., Google Text-to-Speech) to generate audio. At the same time, it uses an image generation program (e.g., DALL-E) to generate images related to the script. These audio and images are then combined to generate video content.
[0728] The generated video is then distributed to the user's device by the server via the user's usual social media platform or email. For example, it may be posted on Twitter or Instagram.
[0729] After the user watches the video, the server collects feedback from the user, including their thoughts and feelings, and stores the collected feedback in a database.
[0730] The collected feedback is passed to a sentiment analysis engine (e.g., IBM Watson Emotion Analysis) to recognize the user's emotions. This emotional data is used to adjust the script generation process for the next time. For example, if the feedback for the previous video was "very satisfied," the next time a script will be generated, it will have a similar content and style.
[0731] Device behavior
[0732] The user terminal provides user information through an authentication process with the server, which allows the server to obtain detailed information about the user.
[0733] The user's device receives the video delivered from the server and stores it locally, allowing the video to be played even in an offline environment. The user can operate the device to watch the video and enjoy the provided content.
[0734] After watching the video, the user's device provides an interface to collect feedback from the user, who can then input their impressions and ratings and send them to the server via their device.
[0735] User Actions
[0736] Users can set their interests through their devices, for example by selecting specific interests such as "technology" or "sports," thereby informing the server of their preferences.
[0737] After watching the video, users can input their feedback through their device, which will be reflected in the next content generation, thereby contributing to increasing user satisfaction.
[0738] Specific examples
[0739] For example, if User A specifies that he or she is interested in "technology," the server operates as follows: First, the server collects User A's interest data and stores it in a database. Next, the generative AI model generates a script with content related to "technology." The AI model generates a specific script using the prompt, "User interest: technology. For the next script you create, please detail the latest technology news."
[0740] Based on the script, a reading program generates audio and an image generation program generates related images. The generated audio and images are integrated to create a video, which is then sent to User A's device. After watching the video, User A enters feedback and sends it to the server. The server passes this feedback to an emotion analysis engine, and the recognized emotional data is reflected in the next script generation.
[0741] This makes it possible to generate and deliver personalized content optimized for each user, thereby increasing user satisfaction.
[0742] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0743] System program processing flow
[0744] Step 1: Collect user information
[0745] The server receives information about the user's interests and local area from the device. When the user inputs information such as "technology" or "Tokyo" through the device, the server records this information in a database.
[0746] Input: User interests and location information
[0747] Output: User profile information stored in the database
[0748] Specific operation: The user selects "technology" on the device settings screen, and the information is sent to the server and stored in a database.
[0749] Step 2: Generative AI model generates script
[0750] The server sends prompts to the generative AI model based on the collected user information to generate a script based on the user's interests. For example, "User interest: technology. For your next script, please detail the latest technology news."
[0751] Input: User information retrieved from the database, prompt text
[0752] Output: Generated script
[0753] How it works: The server inputs the prompt "User interest: Technology" into a generative AI model (e.g., GPT-4), and the model generates a related news script.
[0754] Step 3: Integrating the Text-to-Speech Program and the Image Generator
[0755] The server generates audio based on the generated script using a reading program and generates associated images using an image generation program.
[0756] Input: Generated script
[0757] Output: Generated audio and image files
[0758] Specific operation: The server inputs the script into a speech program (e.g., Google Text-to-Speech) to generate an audio file. In parallel, the server inputs the script into an image generation program (e.g., DALL-E) to generate related images.
[0759] Step 4: Audio and video integration
[0760] The server combines the generated audio and images to create a video.
[0761] Input: Generated audio and image files
[0762] Output: Finished video file
[0763] Specific operation: The server places the audio and image files on the timeline and exports them as a single video file.
[0764] Step 5: Publish your video
[0765] The server then distributes the generated video to the user's device via the user's usual social media platform or email.
[0766] Input: Finished video file
[0767] Output: Video delivered to the user's device
[0768] Specific operation: The video file generated by the server is posted to Twitter or Instagram, or sent by email.
[0769] Step 6: Gather feedback
[0770] After watching a video, users can enter their feedback and send it to the server via their device, which then stores it in a database.
[0771] Input: User feedback
[0772] Output: Feedback stored in a database
[0773] Specific operation: The user enters feedback such as "very satisfied" on the device, and the server receives it and records it in the database.
[0774] Step 7: Emotion Recognition with the Emotion Engine
[0775] The server passes the collected feedback to an emotion engine to analyze the user's emotions.
[0776] Input: User feedback data
[0777] Output: Recognized emotion data
[0778] Specific operation: The server passes the feedback "very satisfied" to an emotion engine (e.g., IBM Watson Emotion Analysis), which then recognizes the emotion "satisfied" and stores it in a database.
[0779] Step 8: Use emotion data
[0780] The server uses the emotion data obtained from the emotion engine to correct the next script generation process.
[0781] Input: Recognized emotion data
[0782] Output: Corrected script generation process
[0783] Specific operation: The server sets the conditions for the next script generation based on the "satisfaction" information and reflects this when sending a new prompt sentence to the generation AI model.
[0784] (Application example 2)
[0785] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0786] In today's world, content distribution services must provide optimal content for each individual user. However, with conventional technologies, content generated based on user interests and local information does not adequately reflect user emotions or real-time feedback. This makes it difficult to improve user satisfaction, and there is a need to further increase engagement for each user.
[0787] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for collecting user information, means for generating a script for each user using a generation AI model, means for generating audio based on the script using a text-to-speech program, means for generating images based on the script using an image generation AI, means for generating a video by combining the generated audio and images, means for delivering the generated video to the user's device, means for recognizing the user's emotion using an emotion recognition engine, and means for reflecting the recognized emotion data in the next content generation. This makes it possible to deliver personalized short video content based not only on the user's interests and regional information but also on their emotions.
[0788] "Means for collecting user information" refers to a mechanism for collecting data including user interests and regional information and storing it on a server.
[0789] A "generative AI model" is a system that includes an algorithm that automatically generates the optimal script based on each user's interests and regional information.
[0790] A "means for generating a script" is a process that uses a generative AI model to create a script appropriate for a particular user.
[0791] A "reading program" is software that automatically generates speech based on a generated script.
[0792] "Image generation AI" is an artificial intelligence system that automatically creates images that match the content of a generated script.
[0793] "Means for generating video" refers to technology that integrates generated audio and images to create a single video.
[0794] "Means for delivering video" refers to the infrastructure for transmitting the generated video to the user's device.
[0795] An "emotion recognition engine" is a system that includes an algorithm for analyzing and recognizing emotions from a user's facial expressions, voice, etc.
[0796] "Means for reflecting emotional data in the next content generation" refers to a process of using the recognized emotional data of the user to provide appropriate feedback to the script or content to be generated next.
[0797] This invention is a system that automatically generates short video content optimized for each user and distributes it periodically. Furthermore, by combining it with an emotion engine, it recognizes the user's emotions and provides personalized content accordingly. The various means required for this and their specific processing steps are described below.
[0798] Collection of User Information
[0799] The server first collects the user's interests and area information. The user enters their interests and area of residence through their device, and the information is sent to the server. This information is stored in a database and used for subsequent processing.
[0800] Script generation
[0801] The server uses the collected user information to generate a script using a generative AI model. This generative AI model inputs the user's interests and local information as prompts and outputs a short script based on that data. For example, a prompt might be, "The user's interest is 'technology.' Please generate a script to deliver today's latest technology news to him."
[0802] Sound and image generation
[0803] The server inputs the generated script into a text-to-speech program to generate audio. It then uses an image generation AI to generate images based on the script. Specifically, the text-to-speech program and image generation AI use a TextToSpeech engine and an ImageGenerator.
[0804] Video Creation and Delivery
[0805] The server combines the generated audio and video to generate a video, which is then distributed to the user's device. The video is generated using a video editing software library and distributed via social media platforms or email.
[0806] Emotion recognition and feedback collection
[0807] After the user watches the video, the device uses an emotion recognition engine to analyze the user's emotions. This engine recognizes emotions in real time through facial expressions and voice analysis, and sends the data to the server. For example, it collects emotional feedback such as "very satisfied."
[0808] Reflection in next script generation
[0809] The server reflects the collected emotional data in the next content generation, correcting the script generation process based on the previous feedback and generating content that is in line with the user's preferences and emotions.
[0810] Thus, the present invention can significantly improve user satisfaction by efficiently generating and delivering personalized short video content that takes user emotions into consideration.
[0811] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0812] Step 1: Collect user information
[0813] Users input their interests and local area information through their devices and send that information to the server. The server stores the received data in a database. This input information includes the user's content preferences and location information. For example, a user might input information that they are interested in "technology" and live in "Tokyo."
[0814] Step 2: Generate the script
[0815] The server generates a script using a generative AI model based on the user information collected in step 1. In this process, the server inputs a prompt such as "The user's interest is 'technology'. Please generate a script that conveys today's latest technology news to him" into the generative AI model, and obtains the optimal script as the output.
[0816] Step 3: Generate audio
[0817] The server inputs the script generated in step 2 into a speech program (such as a TextToSpeech engine) to automatically generate speech. Based on the input, acoustic analysis and synthesis are performed, and the output is a generated audio file.
[0818] Step 4: Generate images
[0819] The server inputs the generated script into an image generation AI (such as ImageGenerator) to generate an image related to the script. The image generation AI analyzes the contents of the script and generates and outputs an appropriate image. For example, it generates an image related to an article about the latest AI technology.
[0820] Step 5: Generate the video
[0821] The server generates a video by combining the audio generated in step 3 with the images generated in step 4. This process uses a video editing software library to combine the audio and images in a sequence and output a single video file containing technology news.
[0822] Step 6: Publish your video
[0823] The server distributes the generated video to the user's device. Distribution is done via the user's usual social media platform or email. Based on the distribution destination information, the server sends the video file to the corresponding platform. The user then watches the video using their device.
[0824] Step 7: Recognize emotions
[0825] While the user is watching a video, the device uses a built-in emotion recognition engine to analyze the user's emotions in real time. This process involves using a camera and microphone to collect the user's facial expressions and tone of voice, and then inputting this data into the emotion recognition algorithm. The output is the user's emotional data.
[0826] Step 8: Gather feedback
[0827] After viewing, users input their feedback and emotional data into their device and send it to the server. The server stores this in a database. The collected feedback is reflected in the next content generation process. For example, emotional feedback such as "very satisfied" can be input.
[0828] Step 9: Reflecting the changes in the next script generation
[0829] The server uses the emotion data collected in step 8 to correct the next content generation process. This ensures that the next script generated is in line with the user's preferences and emotions. This makes it possible to continuously improve user satisfaction.
[0830] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0831] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0832] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0833] [Third embodiment]
[0834] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0835] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[0836] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0837] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0838] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0839] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0840] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0841] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0842] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0843] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0844] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0845] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0846] The system of the present invention automatically generates short video content optimized for each user and distributes it periodically. This system generates content based on the user's interests and local information and distributes it to the user's device, thereby improving user satisfaction.
[0847] Server Operation
[0848] 1. Collection of User Information
[0849] The server collects information about the user's interests and region and stores it in a database. For example, if the user is interested in "technology," the server will retrieve and store that information.
[0850] 2. Script generation using generative AI models
[0851] The server retrieves user information from a database and uses a generative AI model to generate a script based on the user's interests, for example, "Today's Hot Tech News."
[0852] 3. Integration of text-to-speech programs and image generation AI
[0853] Based on the generated script, the server uses a text-to-speech program to generate audio and uses image generation AI to generate images that match the script. The generated audio and images are then integrated to generate a video.
[0854] 4. Video distribution
[0855] The server then delivers the generated video to the user's device via the platform that the user normally uses.
[0856] Device behavior
[0857] 1. User Authentication
[0858] The user terminal goes through an authentication process with the server and provides user information, which allows the server to obtain detailed information about the user.
[0859] 2. Receiving and saving videos
[0860] The user device receives the video delivered from the server and stores it locally, allowing the video to be played offline.
[0861] 3. Play the video
[0862] The user's device plays the stored video to the user, allowing the user to watch the latest information on their area of interest.
[0863] User Actions
[0864] 1. Setting your interests
[0865] Users can set their interests through their devices and send them to the server. For example, they can set interests such as "technology" or "sports."
[0866] 2. Watch the video and give feedback
[0867] After watching a video, users can send feedback about the content to the server via their device, which is reflected in the next content generation.
[0868] Specific examples
[0869] For example, if user A specifies that he is interested in "technology," the server will do the following:
[0870] 1. Collection of User Information
[0871] The server collects user A's interest data and stores it in a database.
[0872] 2. Script generation using generative AI models
[0873] The server retrieves user A's information about "technology" from the database and generates a script using the generative AI model. For example, a script might be generated that reads, "Today's hot technology news: A new AI technology has been announced."
[0874] 3. Integration of text-to-speech programs and image generation AI
[0875] The server inputs the script into a speech program to generate audio.
[0876] Based on the generated script, image generation AI is used to generate related images.
[0877] Integrates audio and images to generate videos.
[0878] 4. Video distribution
[0879] The video generated by the server is distributed to User A's device. For example, distribution can be done using the SNS platform that User A normally uses.
[0880] 5. Receiving and saving videos
[0881] User A's device receives the video and saves it locally.
[0882] 6. Video playback
[0883] User A's device plays the saved video and User A watches it.
[0884] 7. Providing Feedback
[0885] After viewing, User A enters feedback and sends it to the server via their device. For example, feedback such as "The technical content was easy to understand."
[0886] This system enables efficient delivery of content customized for each user, thereby improving user satisfaction.
[0887] The processing flow will be explained below.
[0888] Step 1:
[0889] The server collects user interests and geographical information and stores it in a database. When a user inputs their interests and geographical information, such as "technology" or "Tokyo," through their device, the server receives this information and records it in the user's profile.
[0890] Step 2:
[0891] The server periodically retrieves user information from the database. For example, every morning the system automatically retrieves the interests and geographical locations of all registered users.
[0892] Step 3:
[0893] The server then calls up a generative AI model based on user information to generate a script tailored to each individual user. For example, if a user is interested in "technology news," a script containing a summary of the previous day's new technology news will be generated.
[0894] Step 4:
[0895] The server then passes the generated script to a text-to-speech program to generate audio. For example, the server uses the Google Text-to-Speech API to convert the script into audio data.
[0896] Step 5:
[0897] The server then passes the generated script to an image generation AI to generate images for the video, for example, using an image generation model such as DALL-E to generate images related to the news content.
[0898] Step 6:
[0899] The server then combines the generated audio and images to create a video. The program then creates a slideshow-style video, switching between images as needed against the audio file in the background.
[0900] Step 7:
[0901] The server then distributes the generated video to the platform specified by the user, for example, uploading it to LINE VOOM and obtaining the URL.
[0902] Step 8:
[0903] The server notifies the user's device of the video URL to which the video is being delivered. The user's device receives this notification and saves the video URL locally.
[0904] Step 9:
[0905] The device will use the saved video URL to download the video and store it locally, allowing users to watch the video even without an internet connection.
[0906] Step 10:
[0907] The device plays the stored video for the user to view, and the user operates the device to start the video and view the provided information.
[0908] Step 11:
[0909] After watching a video, users can provide their impressions and feedback via their device, for example, by entering comments such as "The content was easy to understand."
[0910] Step 12:
[0911] The terminal transmits the feedback received from the user to the server, so that the server can utilize the feedback information for generating the next content.
[0912] In this way, the system automatically generates video content customized for each user and distributes it regularly.
[0913] Example 1
[0914] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0915] Conventional short video content distribution systems have difficulty providing personalized content based on user interests and local information. Furthermore, they lack a feedback function to improve the user experience, making it difficult to efficiently deliver individually optimized content.
[0916] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0917] In this invention, the server includes means for collecting user information, means for generating a script for each user using a generative AI model, means for generating audio based on the script using a text-to-speech program, means for generating images based on the script using an image generation AI, means for generating a video by combining the generated audio and images, means for delivering the generated video to the user's terminal, means for receiving and storing the generated video in the user's terminal, means for playing the generated video on the user's terminal, and means for the user to provide feedback after viewing. This makes it possible to efficiently generate and deliver short video content customized for each user and improve the quality of the content by utilizing the feedback.
[0918] "User Information" is data related to an individual user, such as the user's interests and geographical location.
[0919] A "generative AI model" is an artificial intelligence algorithm that automatically generates a script based on user input data.
[0920] A "script" is text generated by a generative AI model that constitutes the content of video content.
[0921] A "reading program" is software that converts text data into speech.
[0922] "Image generation AI" is an artificial intelligence algorithm that automatically generates images based on text and other input data.
[0923] "Video generation" is the process of combining the generated audio and images into a single video file.
[0924] "User device" refers to an electronic device owned by the user, such as a smartphone, tablet, or PC.
[0925] "Video distribution" is the process of sending video data from a server to a user's device.
[0926] "Feedback" refers to reactions such as opinions and impressions that users provide regarding the content they have viewed.
[0927] The system of the present invention automatically generates and periodically distributes short video content optimized for each user. The main components of the system include a server, a terminal, and a user.
[0928] Server Operation
[0929] The server first collects the user's interests and local area information through an API and stores this information in a database. If the user indicates that they are interested in "technology," this information is retrieved and stored by the server. The collected information is then used as prompts to generate scripts using a generative AI model. For example, a prompt such as "Tell me the latest technology-related news" is input into the generative AI model.
[0930] The generative AI model uses a natural language generation model such as GPT-3. Based on this, the server generates a specific script, such as "Today's hot tech news: New AI technology has been announced." Next, a text-to-speech program (such as Google's Text-to-Speech API) generates audio based on this script. At the same time, an image generation AI such as DALL-E or Stable Diffusion generates appropriate images based on the script.
[0931] The generated audio and video are then combined to create a single video. This is done using video editing software (such as FFmpeg). The resulting video can then be distributed via the user's preferred social media platform or a dedicated app. For example, the server could upload the video via the YouTube API and send a link to the user's device.
[0932] Device behavior
[0933] The user device first goes through an authentication process with the server and provides user information. If this authentication is successful, the device can receive the video delivered from the server and store it locally. This allows the user to play the video even offline. When the user presses the "play" button, the stored video is played using video player software.
[0934] User Actions
[0935] Users set their interests through their device. This is done by selecting an area such as "technology" or "sports" from the settings screen within the device's application. Once the selection is complete, the information is immediately sent to the server. After watching the video, users enter their opinions and thoughts in a feedback form and send it to the server. For example, by providing feedback such as "the content of the video was easy to understand," this will be reflected in the next content generation.
[0936] Specific examples
[0937] For example, if User A specifies that he or she is interested in "technology," the procedure is as follows: The server collects User A's interest data in "technology" and stores it in a database. Next, a generative AI model is used to generate a script that reads, "Today's hot technology news: A new AI technology has been announced." The script is input into Google's Text-to-Speech API to generate audio, and DALL-E is used to generate images based on the script, which are then integrated to create a video. The created video is distributed to User A's device via YouTube, and the device saves the video locally. User A plays the saved video, and after watching, sends feedback to the server that "the technical content was easy to understand."
[0938] This system makes it possible to efficiently generate and distribute short video content customized for each user, and to improve the quality of the content by utilizing feedback.
[0939] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0940] Step 1:
[0941] The server collects user information. It receives interest and region information sent by the user from their device via API and stores it in a database. The input is the interest data set by the user (e.g., "technology"), and the output is that data is stored in the database.
[0942] Step 2:
[0943] A script is generated using a generative AI model. The server retrieves the user's interest data from the database and inputs it as a prompt sentence to the generative AI model. The inputs are the "user's interest data" and the "prompt sentence" (e.g., "Tell me the latest technology news"), and the output is a script generated by the generative AI model (e.g., "Today's hot technology news: A new AI technology has been announced").
[0944] Step 3:
[0945] The server inputs the script into a text-to-speech program to generate audio. Specifically, the generated script is input into Google's Text-to-Speech API to obtain audio data. The input at this time is the generated script, and the output is the audio data generated by the text-to-speech program.
[0946] Step 4:
[0947] The server generates an image based on the script. The generated script is input into an image generation AI (DALL-E or Stable Diffusion) to generate a related image. The input is the generated script, and the output is image data generated by the image generation AI.
[0948] Step 5:
[0949] The server combines the audio and image data to generate a video. The audio and image data are input into video editing software (such as FFmpeg) to generate a single video file. The input in this case is the audio and image data, and the output is the generated video file.
[0950] Step 6:
[0951] The server delivers the generated video to the user's device. The generated video file is uploaded to a social media platform (YouTube, Instagram, etc.) and the link is sent to the user's device. The input in this case is the generated video file, and the output is the video link.
[0952] Step 7:
[0953] The user device receives the video and saves it locally. The video is downloaded via the video link sent from the server and saved in local storage. The input is the video link, and the output is the saved video file.
[0954] Step 8:
[0955] The user device plays the stored video. When the user presses the "play" button, the stored video is played using video player software. The input is the stored video file, and the output is the playback of the video.
[0956] Step 9:
[0957] Users provide feedback after watching. After watching a video, users enter their opinions and thoughts in a feedback form and send the data to the server. The input is the user's feedback data, and the output is the feedback information stored on the server.
[0958] (Application example 1)
[0959] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0960] Conventional automatic short video content generation systems, when generating videos optimized for each user, mainly focused on the user's interests and local information, and therefore were unable to adequately stimulate purchasing motivation or promote products in physical stores. In particular, they were unable to create customized content that reflected the customer's purchasing history or real-time interactions in the store, which created challenges in improving the customer experience and promoting sales.
[0961] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0962] In this invention, the server includes means for collecting user information, means for generating a script for each user using a generative AI model, means for generating audio based on the script using a text-to-speech program, means for generating images based on the script using an image generation AI, means for generating a video by combining the generated audio and images, means for delivering the generated video to the user's device and a display in a physical store, means for performing user authentication and setting interest information, and means for collecting feedback after viewing the video. This makes it possible to provide customers visiting a store with a customized promotional video that reflects their purchasing history, interests, and real-time interactions, thereby improving the customer experience and promoting sales.
[0963] "User Information" is data about individual users, including their interests, geographic location, and purchasing history.
[0964] A "generative AI model" is an artificial intelligence model that generates scripts based on user information.
[0965] A "reading program" is software that automatically generates speech based on a generated script.
[0966] "Image generation AI" is an artificial intelligence model that generates relevant images based on a script.
[0967] "User authentication" refers to the means by which a user's identity is verified, which is the prerequisite process for collecting and personalizing user information.
[0968] "Interest information" is data that indicates a user's interests in a particular field.
[0969] "Feedback" refers to the ratings and opinions provided by users after watching a video, and is data that is reflected in the next content generation.
[0970] A "physical store" is a sales facility located in a physical location that offers goods and services.
[0971] A "terminal" is an electronic device that allows a user to receive or send information, including a smartphone, smart glasses, or a display.
[0972] A "display" is a device for displaying visual information, and includes digital signage used in stores.
[0973] This system automatically generates short video content optimized for each user, improving promotions and customer experiences in brick-and-mortar stores. By generating customized videos based on the purchasing history and interests of customers visiting brick-and-mortar stores and providing them via displays or smart glasses, more personalized promotions can be achieved.
[0974] Server Operation
[0975] The server collects user information and generates a script using a generative AI model. Based on the script, a text-to-speech program generates audio and an image generation AI generates images. These are then combined to generate a video, which is then distributed to displays in physical stores and to users' devices.
[0976] Main components
[0977] 1. Collection of User Information
[0978] The server collects user interests, location information, and purchasing history and stores them in a database.
[0979] 2. Script generation using generative AI models
[0980] The server uses a generative AI model to generate a personalized script based on collected user information.
[0981] 3. Integration of text-to-speech programs and image generation AI
[0982] The server inputs the generated script into a speech program to generate audio, and also uses image generation AI to generate related images.
[0983] 4. Video Creation and Distribution
[0984] The server integrates audio and images to generate a video, which is then distributed to displays in physical stores and to users' devices.
[0985] Device behavior
[0986] The device provides user information through an authentication process with the server, receives the video, and stores it locally for offline playback. After watching the video, users can provide feedback, which is reflected in the next content generation.
[0987] Main features
[0988] 1. User Authentication
[0989] The user terminal goes through an authentication process with the server and provides user information.
[0990] 2. Receiving and saving videos
[0991] The user device receives the video delivered from the server and stores it locally.
[0992] 3. Feedback function
[0993] After watching a video, users provide feedback, which is sent to the server and used to generate the next video.
[0994] Hardware and software used
[0995] Server: A Python server application using Flask
[0996] Generative AI models: script generation based on user interests (e.g., GPT-3)
[0997] Text-to-speech program: A text-to-speech tool (e.g., Google Text-to-Speech)
[0998] Image generation AI: Deep learning model (e.g., DALL-E)
[0999] Video generation tools: video editing libraries (e.g., moviepy)
[1000] Devices: Smartphones, smart glasses, in-store displays
[1001] Specific examples
[1002] For example, when a specific customer enters a store, a customized video based on the customer's past purchase history and interests will be displayed on the in-store display, allowing the customer to view new product information and attractive promotions that are tailored to them.
[1003] Example prompt for a generative AI model:
[1004] Customer Interest: "Cooking"
[1005] Prompt for generative AI model: "Generate a short video about the latest cooking appliances and fresh ingredients."
[1006] This system allows physical stores to provide promotional videos optimized for each individual customer, improving the customer experience and promoting sales.
[1007] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1008] Step 1:
[1009] Collection of User Information
[1010] The server collects users' interests, local area information, and purchasing history. The information provided by each user through their device and their in-store purchasing history are stored in a database. This allows the collection of data necessary to generate customized content. The input data are users' interests, local area information, and purchasing history. The output is a database of the collected user information.
[1011] Step 2:
[1012] Script generation using generative AI models
[1013] The server generates a personalized script using a generative AI model based on the collected user information. The script is generated by using data obtained from the user information database as input and providing a prompt to the generative AI model (e.g., "Customer interest: 'Cooking'. Prompt to the generative AI model: 'Generate a short video about the latest cooking equipment and fresh ingredients.'"). The output is the generated customized script.
[1014] Step 3:
[1015] Speech generation by a reading program
[1016] The server inputs the generated script into a text-to-speech program to generate audio. The input data is the generated script, and the output is the generated audio file. The text-to-speech program uses a speech synthesis tool such as Google Text-to-Speech.
[1017] Step 4:
[1018] Image generation using image generation AI
[1019] The server inputs the generated script into the image generation AI, which then generates an image based on it. The input data is the generated script, and the output is a generated image file. The image generation AI model uses a deep learning model such as DALL-E.
[1020] Step 5:
[1021] Video generation
[1022] The server generates a video by integrating the generated audio and images. The input data are the generated audio and image files, and the output is the generated video file. To generate the video, a video editing library such as moviepy is used.
[1023] Step 6:
[1024] Video distribution
[1025] The server distributes the generated video to the user's device and the display in the physical store. The input data is the generated video file, and the output is the video distributed to the user's device and the display in the physical store. Distribution is done via the Internet.
[1026] Step 7:
[1027] User authentication and interest settings
[1028] Users set and update their authentication and interest information on the server through their devices. The input data is the user's authentication information and interest information, and the output is the authenticated user information stored on the server.
[1029] Step 8:
[1030] Providing feedback after watching the video
[1031] Users watch videos through their devices and provide feedback after viewing. The input data is feedback information from the users, and the output is stored in the server's feedback database. This provides data that will be reflected in the next content generation.
[1032] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1033] The system of the present invention automatically generates short video content optimized for each user and distributes it periodically. Furthermore, by combining it with an emotion engine, it can recognize the user's emotions and provide personalized content accordingly, thereby further improving user satisfaction.
[1034] Server Operation
[1035] 1. Collection of User Information
[1036] The server collects user interests and geographical information and stores it in a database. When a user inputs their interests and geographical information, such as "technology" or "Tokyo," through their device, the server receives this information and records it in the user's profile.
[1037] 2. Script generation using generative AI models
[1038] The server retrieves user information from a database and uses a generative AI model to generate a script based on the user's interests, for example, "Today's Hot Tech News."
[1039] 3. Integration of text-to-speech programs and image generation AI
[1040] Based on the generated script, the server uses a text-to-speech program to generate audio and uses image generation AI to generate images that match the script. The generated audio and images are then integrated to generate a video.
[1041] 4. Video distribution
[1042] The server then delivers the generated video to the user's device via the platform the user normally uses, such as a social networking platform or email.
[1043] 5. Gathering Feedback
[1044] After watching a video, users input their impressions and feedback, which are then sent to the server via their device, where they are stored in a database.
[1045] 6. Emotion Recognition by Emotion Engine
[1046] The server passes the collected feedback to the emotion engine to analyze the user's emotions. The emotion engine recognizes the emotions the user felt when giving feedback and stores them in a database.
[1047] 7. Use of Emotional Data
[1048] The server uses the emotion data recognized by the emotion engine to correct the script generation process for the next time. For example, if the user was recognized as "very satisfied" during the previous video viewing, a script with similar content and style will be generated.
[1049] Device behavior
[1050] 1. User Authentication
[1051] The user terminal goes through an authentication process with the server and provides user information, which allows the server to obtain detailed information about the user.
[1052] 2. Receiving and saving videos
[1053] The user's device receives the video streamed from the server and stores it locally, allowing the video to be played even without an internet connection.
[1054] 3. Play the video
[1055] The user device plays the stored video to the user, who then operates the device to start the video and view the provided information.
[1056] 4. Providing Feedback
[1057] After watching the video, users can enter their impressions and feedback through the device. For example, they can enter feedback such as "The content was easy to understand."
[1058] User Actions
[1059] 1. Setting your interests
[1060] Users can set their interests through their devices and send them to the server. For example, they can set interests such as "technology" or "sports."
[1061] 2. Watch the video and give feedback
[1062] After watching a video, users can send feedback about the content to the server via their device, which is reflected in the next content generation.
[1063] Specific examples
[1064] For example, if user A specifies that he is interested in "technology," the server will do the following:
[1065] 1. Collection of User Information
[1066] The server collects user A's interest data and stores it in a database.
[1067] 2. Script generation using generative AI models
[1068] The server retrieves user A's information about "technology" from the database and generates a script using the generative AI model. For example, a script might be generated that reads, "Today's hot technology news: A new AI technology has been announced."
[1069] 3. Integration of text-to-speech programs and image generation AI
[1070] The server inputs the script into a speech program to generate audio.
[1071] Based on the generated script, image generation AI is used to generate related images.
[1072] Integrates audio and images to generate videos.
[1073] 4. Video distribution
[1074] The video generated by the server is distributed to User A's device. For example, distribution can be done using the SNS platform that User A normally uses.
[1075] 5. Gathering Feedback
[1076] User A enters feedback after watching the video. For example, he / she enters a comment including emotion such as "Very satisfied."
[1077] 6. Emotion Recognition by Emotion Engine
[1078] The server passes the feedback to the emotion engine to recognize User A's emotion.
[1079] 7. Use of Emotional Data
[1080] The server then uses the recognized emotion data to correct the script generation process for the next time, so that the next script will be generated in the same style.
[1081] This system makes it possible to further personalize and efficiently deliver content customized for each user based on their emotions, thereby further improving user satisfaction.
[1082] The processing flow will be explained below.
[1083] Step 1:
[1084] The server collects user interests and geographical location information and stores it in a database. For example, if a user enters interests and geographical location such as "technology" or "Tokyo," the server receives this and records it in the user profile.
[1085] Step 2:
[1086] The server periodically retrieves user information from the database. For example, every morning the server can be configured to automatically retrieve user interests and location information.
[1087] Step 3:
[1088] The server then calls up a generative AI model based on user information to generate a script tailored to each individual user. For example, if a user is interested in "technology news," a script containing a summary of the previous day's new technology news will be generated.
[1089] Step 4:
[1090] The server then passes the generated script to a text-to-speech program to generate audio. For example, the server uses the Google Text-to-Speech API to convert the script into audio data.
[1091] Step 5:
[1092] The server then passes the generated script to an image generation AI to generate images for the video, for example, using an image generation model such as DALL-E to generate images related to the news content.
[1093] Step 6:
[1094] The server combines the generated audio and images to create a video, and the program generates a slideshow-style video that displays the images while switching between them against the audio file in the background.
[1095] Step 7:
[1096] The server then distributes the generated video to the platform specified by the user, for example, uploading it to LINE VOOM and obtaining the URL.
[1097] Step 8:
[1098] The server notifies the user's device of the video URL to which the video is being delivered. The user's device receives this notification and saves the video URL locally.
[1099] Step 9:
[1100] The device will use the saved video URL to download the video and store it locally, allowing users to watch the video even without an internet connection.
[1101] Step 10:
[1102] The device plays the stored video for the user to view, and the user operates the device to start the video and view the provided information.
[1103] Step 11:
[1104] After watching a video, users can provide their impressions and feedback via their device, for example, by entering comments such as "The content was easy to understand."
[1105] Step 12:
[1106] The terminal transmits the feedback received from the user to the server, so that the server can utilize the feedback information for generating the next content.
[1107] Step 13:
[1108] The server passes the feedback to an emotion engine to recognize the user's emotion. The emotion engine identifies the user's emotion (e.g., "very satisfied") from the feedback content using, for example, text analysis technology.
[1109] Step 14:
[1110] The server stores the emotion data recognized by the emotion engine in a database, for example, adding emotion data such as "very satisfied" to the user profile.
[1111] Step 15:
[1112] The server uses the recognized emotion data to adjust the script generation process for the next time. For example, if the user was recognized as "very satisfied" during the previous video viewing, the server will generate a script with a similar content style next time.
[1113] Through this process, the system can provide more personalized content based on the user's emotions, thereby improving the quality of the user experience and further increasing satisfaction.
[1114] Example 2
[1115] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1116] In modern content delivery systems, it is important to generate and deliver personalized content optimized for each user. However, conventional systems have difficulty efficiently collecting user sentiment and feedback and reflecting it in the next content generation process. This can easily lead to a decrease in user satisfaction, and more effective personalization is required.
[1117] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1118] In this invention, the server includes means for collecting user information, means for generating a script for each user using a generative AI model, means for generating audio based on the script using a text-to-speech program, means for generating images based on the script using an image generation program, means for generating a video by combining the generated audio and images, means for delivering the generated video to the user's device, means for collecting feedback from the user, means for passing the collected feedback to a sentiment analysis engine to recognize the user's emotions, and means for correcting the next script generation process using the recognized emotion data. This enables the generation and delivery of content optimized for each user, making it possible to provide personalized content that reflects the user's emotions and feedback and provides a high level of satisfaction.
[1119] "User information" refers to data necessary for personalizing content, such as user interests and regional information.
[1120] A "generative AI model" is a model that uses artificial intelligence technology to automatically generate scripts based on user information.
[1121] A "reading program" is a program for generating voice based on a generated script.
[1122] An "image generation program" is a program for generating images based on a generated script.
[1123] "Video" is media content created by combining audio and video.
[1124] "Feedback" refers to the impressions and evaluations provided by users after viewing content.
[1125] An "emotion analysis engine" is an engine that analyzes feedback collected from users and recognizes their emotions.
[1126] "Emotional Data" is data that indicates the emotional state of a user as recognized by an emotion analysis engine.
[1127] The system of the present invention automatically generates short video content optimized for each user and distributes it periodically. The system is mainly composed of three components: a server, a terminal, and a user. The specific operation of each component is explained below.
[1128] Server Operation
[1129] The server first collects user information, including user interests and local area information. For example, if a user inputs their interests such as "technology" or "Tokyo" through their device, the server receives this information and records it in a database.
[1130] The server then uses a generative AI model to generate a script based on the user's information. A natural language generation model such as GPT-4 can be used as the generative AI model. A specific example of a prompt is, "User interest: technology. Please elaborate on the latest technology news in your next script." This prompt causes the AI model to generate a script with appropriate content.
[1131] Based on the generated script, the server uses a text-to-speech program (e.g., Google Text-to-Speech) to generate audio. At the same time, it uses an image generation program (e.g., DALL-E) to generate images related to the script. These audio and images are then combined to generate video content.
[1132] The generated video is then distributed to the user's device by the server via the user's usual social media platform or email. For example, it may be posted on Twitter or Instagram.
[1133] After the user watches the video, the server collects feedback from the user, including their thoughts and feelings, and stores the collected feedback in a database.
[1134] The collected feedback is passed to a sentiment analysis engine (e.g., IBM Watson Emotion Analysis) to recognize the user's emotions. This emotional data is used to adjust the script generation process for the next time. For example, if the feedback for the previous video was "very satisfied," the next time a script will be generated, it will have a similar content and style.
[1135] Device behavior
[1136] The user terminal provides user information through an authentication process with the server, which allows the server to obtain detailed information about the user.
[1137] The user's device receives the video delivered from the server and stores it locally, allowing the video to be played even in an offline environment. The user can operate the device to watch the video and enjoy the provided content.
[1138] After watching the video, the user's device provides an interface to collect feedback from the user, who can then input their impressions and ratings and send them to the server via their device.
[1139] User Actions
[1140] Users can set their interests through their devices, for example by selecting specific interests such as "technology" or "sports," thereby informing the server of their preferences.
[1141] After watching the video, users can input their feedback through their device, which will be reflected in the next content generation, thereby contributing to increasing user satisfaction.
[1142] Specific examples
[1143] For example, if User A specifies that he or she is interested in "technology," the server operates as follows: First, the server collects User A's interest data and stores it in a database. Next, the generative AI model generates a script with content related to "technology." The AI model generates a specific script using the prompt, "User interest: technology. For the next script you create, please detail the latest technology news."
[1144] Based on the script, a reading program generates audio and an image generation program generates related images. The generated audio and images are integrated to create a video, which is then sent to User A's device. After watching the video, User A enters feedback and sends it to the server. The server passes this feedback to an emotion analysis engine, and the recognized emotional data is reflected in the next script generation.
[1145] This makes it possible to generate and deliver personalized content optimized for each user, thereby increasing user satisfaction.
[1146] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1147] System program processing flow
[1148] Step 1: Collect user information
[1149] The server receives information about the user's interests and local area from the device. When the user inputs information such as "technology" or "Tokyo" through the device, the server records this information in a database.
[1150] Input: User interests and location information
[1151] Output: User profile information stored in the database
[1152] Specific operation: The user selects "technology" on the device settings screen, and the information is sent to the server and stored in a database.
[1153] Step 2: Generative AI model generates script
[1154] The server sends prompts to the generative AI model based on the collected user information to generate a script based on the user's interests. For example, "User interest: technology. For your next script, please detail the latest technology news."
[1155] Input: User information retrieved from the database, prompt text
[1156] Output: Generated script
[1157] How it works: The server inputs the prompt "User interest: Technology" into a generative AI model (e.g., GPT-4), and the model generates a related news script.
[1158] Step 3: Integrating the Text-to-Speech Program and the Image Generator
[1159] The server generates audio based on the generated script using a reading program and generates associated images using an image generation program.
[1160] Input: Generated script
[1161] Output: Generated audio and image files
[1162] Specific operation: The server inputs the script into a speech program (e.g., Google Text-to-Speech) to generate an audio file. In parallel, the server inputs the script into an image generation program (e.g., DALL-E) to generate related images.
[1163] Step 4: Audio and video integration
[1164] The server combines the generated audio and images to create a video.
[1165] Input: Generated audio and image files
[1166] Output: Finished video file
[1167] Specific operation: The server places the audio and image files on the timeline and exports them as a single video file.
[1168] Step 5: Publish your video
[1169] The server then distributes the generated video to the user's device via the user's usual social media platform or email.
[1170] Input: Finished video file
[1171] Output: Video delivered to the user's device
[1172] Specific operation: The video file generated by the server is posted to Twitter or Instagram, or sent by email.
[1173] Step 6: Gather feedback
[1174] After watching a video, users can enter their feedback and send it to the server via their device, which then stores it in a database.
[1175] Input: User feedback
[1176] Output: Feedback stored in a database
[1177] Specific operation: The user enters feedback such as "very satisfied" on the device, and the server receives it and records it in the database.
[1178] Step 7: Emotion Recognition with the Emotion Engine
[1179] The server passes the collected feedback to an emotion engine to analyze the user's emotions.
[1180] Input: User feedback data
[1181] Output: Recognized emotion data
[1182] Specific operation: The server passes the feedback "very satisfied" to an emotion engine (e.g., IBM Watson Emotion Analysis), which then recognizes the emotion "satisfied" and stores it in a database.
[1183] Step 8: Use emotion data
[1184] The server uses the emotion data obtained from the emotion engine to correct the next script generation process.
[1185] Input: Recognized emotion data
[1186] Output: Corrected script generation process
[1187] Specific operation: The server sets the conditions for the next script generation based on the "satisfaction" information and reflects this when sending a new prompt sentence to the generation AI model.
[1188] (Application example 2)
[1189] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1190] In today's world, content distribution services must provide optimal content for each individual user. However, with conventional technologies, content generated based on user interests and local information does not adequately reflect user emotions or real-time feedback. This makes it difficult to improve user satisfaction, and there is a need to further increase engagement for each user.
[1191] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for collecting user information, means for generating a script for each user using a generation AI model, means for generating audio based on the script using a text-to-speech program, means for generating images based on the script using an image generation AI, means for generating a video by combining the generated audio and images, means for delivering the generated video to the user's device, means for recognizing the user's emotion using an emotion recognition engine, and means for reflecting the recognized emotion data in the next content generation. This makes it possible to deliver personalized short video content based not only on the user's interests and regional information but also on their emotions.
[1192] "Means for collecting user information" refers to a mechanism for collecting data including user interests and regional information and storing it on a server.
[1193] A "generative AI model" is a system that includes an algorithm that automatically generates the optimal script based on each user's interests and regional information.
[1194] A "means for generating a script" is a process that uses a generative AI model to create a script appropriate for a particular user.
[1195] A "reading program" is software that automatically generates speech based on a generated script.
[1196] "Image generation AI" is an artificial intelligence system that automatically creates images that match the content of a generated script.
[1197] "Means for generating video" refers to technology that integrates generated audio and images to create a single video.
[1198] "Means for delivering video" refers to the infrastructure for transmitting the generated video to the user's device.
[1199] An "emotion recognition engine" is a system that includes an algorithm for analyzing and recognizing emotions from a user's facial expressions, voice, etc.
[1200] "Means for reflecting emotional data in the next content generation" refers to a process of using the recognized emotional data of the user to provide appropriate feedback to the script or content to be generated next.
[1201] This invention is a system that automatically generates short video content optimized for each user and distributes it periodically. Furthermore, by combining it with an emotion engine, it recognizes the user's emotions and provides personalized content accordingly. The various means required for this and their specific processing steps are described below.
[1202] Collection of User Information
[1203] The server first collects the user's interests and area information. The user enters their interests and area of residence through their device, and the information is sent to the server. This information is stored in a database and used for subsequent processing.
[1204] Script generation
[1205] The server uses the collected user information to generate a script using a generative AI model. This generative AI model inputs the user's interests and local information as prompts and outputs a short script based on that data. For example, a prompt might be, "The user's interest is 'technology.' Please generate a script to deliver today's latest technology news to him."
[1206] Sound and image generation
[1207] The server inputs the generated script into a text-to-speech program to generate audio. It then uses an image generation AI to generate images based on the script. Specifically, the text-to-speech program and image generation AI use a TextToSpeech engine and an ImageGenerator.
[1208] Video Creation and Delivery
[1209] The server combines the generated audio and video to generate a video, which is then distributed to the user's device. The video is generated using a video editing software library and distributed via social media platforms or email.
[1210] Emotion recognition and feedback collection
[1211] After the user watches the video, the device uses an emotion recognition engine to analyze the user's emotions. This engine recognizes emotions in real time through facial expressions and voice analysis, and sends the data to the server. For example, it collects emotional feedback such as "very satisfied."
[1212] Reflection in next script generation
[1213] The server reflects the collected emotional data in the next content generation, correcting the script generation process based on the previous feedback and generating content that is in line with the user's preferences and emotions.
[1214] Thus, the present invention can significantly improve user satisfaction by efficiently generating and delivering personalized short video content that takes user emotions into consideration.
[1215] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1216] Step 1: Collect user information
[1217] Users input their interests and local area information through their devices and send that information to the server. The server stores the received data in a database. This input information includes the user's content preferences and location information. For example, a user might input information that they are interested in "technology" and live in "Tokyo."
[1218] Step 2: Generate the script
[1219] The server generates a script using a generative AI model based on the user information collected in step 1. In this process, the server inputs a prompt such as "The user's interest is 'technology'. Please generate a script that conveys today's latest technology news to him" into the generative AI model, and obtains the optimal script as the output.
[1220] Step 3: Generate audio
[1221] The server inputs the script generated in step 2 into a speech program (such as a TextToSpeech engine) to automatically generate speech. Based on the input, acoustic analysis and synthesis are performed, and the output is a generated audio file.
[1222] Step 4: Generate images
[1223] The server inputs the generated script into an image generation AI (such as ImageGenerator) to generate an image related to the script. The image generation AI analyzes the contents of the script and generates and outputs an appropriate image. For example, it generates an image related to an article about the latest AI technology.
[1224] Step 5: Generate the video
[1225] The server generates a video by combining the audio generated in step 3 with the images generated in step 4. This process uses a video editing software library to combine the audio and images in a sequence and output a single video file containing technology news.
[1226] Step 6: Publish your video
[1227] The server distributes the generated video to the user's device. Distribution is done via the user's usual social media platform or email. Based on the distribution destination information, the server sends the video file to the corresponding platform. The user then watches the video using their device.
[1228] Step 7: Recognize emotions
[1229] While the user is watching a video, the device uses a built-in emotion recognition engine to analyze the user's emotions in real time. This process involves using a camera and microphone to collect the user's facial expressions and tone of voice, and then inputting this data into the emotion recognition algorithm. The output is the user's emotional data.
[1230] Step 8: Gather feedback
[1231] After viewing, users input their feedback and emotional data into their device and send it to the server. The server stores this in a database. The collected feedback is reflected in the next content generation process. For example, emotional feedback such as "very satisfied" can be input.
[1232] Step 9: Reflecting the changes in the next script generation
[1233] The server uses the emotion data collected in step 8 to correct the next content generation process. This ensures that the next script generated is in line with the user's preferences and emotions. This makes it possible to continuously improve user satisfaction.
[1234] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1235] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1236] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1237] [Fourth embodiment]
[1238] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1239] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1240] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1241] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1242] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1243] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1244] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1245] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1246] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1247] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1248] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1249] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1250] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1251] The system of the present invention automatically generates short video content optimized for each user and distributes it periodically. This system generates content based on the user's interests and local information and distributes it to the user's device, thereby improving user satisfaction.
[1252] Server Operation
[1253] 1. Collection of User Information
[1254] The server collects information about the user's interests and region and stores it in a database. For example, if the user is interested in "technology," the server will retrieve and store that information.
[1255] 2. Script generation using generative AI models
[1256] The server retrieves user information from a database and uses a generative AI model to generate a script based on the user's interests, for example, "Today's Hot Tech News."
[1257] 3. Integration of text-to-speech programs and image generation AI
[1258] Based on the generated script, the server uses a text-to-speech program to generate audio and uses image generation AI to generate images that match the script. The generated audio and images are then integrated to generate a video.
[1259] 4. Video distribution
[1260] The server then delivers the generated video to the user's device via the platform that the user normally uses.
[1261] Device behavior
[1262] 1. User Authentication
[1263] The user terminal goes through an authentication process with the server and provides user information, which allows the server to obtain detailed information about the user.
[1264] 2. Receiving and saving videos
[1265] The user device receives the video delivered from the server and stores it locally, allowing the video to be played offline.
[1266] 3. Play the video
[1267] The user's device plays the stored video to the user, allowing the user to watch the latest information on their area of interest.
[1268] User Actions
[1269] 1. Setting your interests
[1270] Users can set their interests through their devices and send them to the server. For example, they can set interests such as "technology" or "sports."
[1271] 2. Watch the video and give feedback
[1272] After watching a video, users can send feedback about the content to the server via their device, which is reflected in the next content generation.
[1273] Specific examples
[1274] For example, if user A specifies that he is interested in "technology," the server will do the following:
[1275] 1. Collection of User Information
[1276] The server collects user A's interest data and stores it in a database.
[1277] 2. Script generation using generative AI models
[1278] The server retrieves user A's information about "technology" from the database and generates a script using the generative AI model. For example, a script might be generated that reads, "Today's hot technology news: A new AI technology has been announced."
[1279] 3. Integration of text-to-speech programs and image generation AI
[1280] The server inputs the script into a speech program to generate audio.
[1281] Based on the generated script, image generation AI is used to generate related images.
[1282] Integrates audio and images to generate videos.
[1283] 4. Video distribution
[1284] The video generated by the server is distributed to User A's device. For example, distribution can be done using the SNS platform that User A normally uses.
[1285] 5. Receiving and saving videos
[1286] User A's device receives the video and saves it locally.
[1287] 6. Video playback
[1288] User A's device plays the saved video and User A watches it.
[1289] 7. Providing Feedback
[1290] After viewing, User A enters feedback and sends it to the server via their device. For example, feedback such as "The technical content was easy to understand."
[1291] This system enables efficient delivery of content customized for each user, thereby improving user satisfaction.
[1292] The processing flow will be explained below.
[1293] Step 1:
[1294] The server collects user interests and geographical information and stores it in a database. When a user inputs their interests and geographical information, such as "technology" or "Tokyo," through their device, the server receives this information and records it in the user's profile.
[1295] Step 2:
[1296] The server periodically retrieves user information from the database. For example, every morning the system automatically retrieves the interests and geographical locations of all registered users.
[1297] Step 3:
[1298] The server then calls up a generative AI model based on user information to generate a script tailored to each individual user. For example, if a user is interested in "technology news," a script containing a summary of the previous day's new technology news will be generated.
[1299] Step 4:
[1300] The server then passes the generated script to a text-to-speech program to generate audio. For example, the server uses the Google Text-to-Speech API to convert the script into audio data.
[1301] Step 5:
[1302] The server then passes the generated script to an image generation AI to generate images for the video, for example, using an image generation model such as DALL-E to generate images related to the news content.
[1303] Step 6:
[1304] The server then combines the generated audio and images to create a video. The program then creates a slideshow-style video, switching between images as needed against the audio file in the background.
[1305] Step 7:
[1306] The server then distributes the generated video to the platform specified by the user, for example, uploading it to LINE VOOM and obtaining the URL.
[1307] Step 8:
[1308] The server notifies the user's device of the video URL to which the video is being delivered. The user's device receives this notification and saves the video URL locally.
[1309] Step 9:
[1310] The device will use the saved video URL to download the video and store it locally, allowing users to watch the video even without an internet connection.
[1311] Step 10:
[1312] The device plays the stored video for the user to view, and the user operates the device to start the video and view the provided information.
[1313] Step 11:
[1314] After watching a video, users can provide their impressions and feedback via their device, for example, by entering comments such as "The content was easy to understand."
[1315] Step 12:
[1316] The terminal transmits the feedback received from the user to the server, so that the server can utilize the feedback information for generating the next content.
[1317] In this way, the system automatically generates video content customized for each user and distributes it regularly.
[1318] Example 1
[1319] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1320] Conventional short video content distribution systems have difficulty providing personalized content based on user interests and local information. Furthermore, they lack a feedback function to improve the user experience, making it difficult to efficiently deliver individually optimized content.
[1321] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1322] In this invention, the server includes means for collecting user information, means for generating a script for each user using a generative AI model, means for generating audio based on the script using a text-to-speech program, means for generating images based on the script using an image generation AI, means for generating a video by combining the generated audio and images, means for delivering the generated video to the user's terminal, means for receiving and storing the generated video in the user's terminal, means for playing the generated video on the user's terminal, and means for the user to provide feedback after viewing. This makes it possible to efficiently generate and deliver short video content customized for each user and improve the quality of the content by utilizing the feedback.
[1323] "User Information" is data related to an individual user, such as the user's interests and geographical location.
[1324] A "generative AI model" is an artificial intelligence algorithm that automatically generates a script based on user input data.
[1325] A "script" is text generated by a generative AI model that constitutes the content of video content.
[1326] A "reading program" is software that converts text data into speech.
[1327] "Image generation AI" is an artificial intelligence algorithm that automatically generates images based on text and other input data.
[1328] "Video generation" is the process of combining the generated audio and images into a single video file.
[1329] "User device" refers to an electronic device owned by the user, such as a smartphone, tablet, or PC.
[1330] "Video distribution" is the process of sending video data from a server to a user's device.
[1331] "Feedback" refers to reactions such as opinions and impressions that users provide regarding the content they have viewed.
[1332] The system of the present invention automatically generates and periodically distributes short video content optimized for each user. The main components of the system include a server, a terminal, and a user.
[1333] Server Operation
[1334] The server first collects the user's interests and local area information through an API and stores this information in a database. If the user indicates that they are interested in "technology," this information is retrieved and stored by the server. The collected information is then used as prompts to generate scripts using a generative AI model. For example, a prompt such as "Tell me the latest technology-related news" is input into the generative AI model.
[1335] The generative AI model uses a natural language generation model such as GPT-3. Based on this, the server generates a specific script, such as "Today's hot tech news: New AI technology has been announced." Next, a text-to-speech program (such as Google's Text-to-Speech API) generates audio based on this script. At the same time, an image generation AI such as DALL-E or Stable Diffusion generates appropriate images based on the script.
[1336] The generated audio and video are then combined to create a single video. This is done using video editing software (such as FFmpeg). The resulting video can then be distributed via the user's preferred social media platform or a dedicated app. For example, the server could upload the video via the YouTube API and send a link to the user's device.
[1337] Device behavior
[1338] The user device first goes through an authentication process with the server and provides user information. If this authentication is successful, the device can receive the video delivered from the server and store it locally. This allows the user to play the video even offline. When the user presses the "play" button, the stored video is played using video player software.
[1339] User Actions
[1340] Users set their interests through their device. This is done by selecting an area such as "technology" or "sports" from the settings screen within the device's application. Once the selection is complete, the information is immediately sent to the server. After watching the video, users enter their opinions and thoughts in a feedback form and send it to the server. For example, by providing feedback such as "the content of the video was easy to understand," this will be reflected in the next content generation.
[1341] Specific examples
[1342] For example, if User A specifies that he or she is interested in "technology," the procedure is as follows: The server collects User A's interest data in "technology" and stores it in a database. Next, a generative AI model is used to generate a script that reads, "Today's hot technology news: A new AI technology has been announced." The script is input into Google's Text-to-Speech API to generate audio, and DALL-E is used to generate images based on the script, which are then integrated to create a video. The created video is distributed to User A's device via YouTube, and the device saves the video locally. User A plays the saved video, and after watching, sends feedback to the server that "the technical content was easy to understand."
[1343] This system makes it possible to efficiently generate and distribute short video content customized for each user, and to improve the quality of the content by utilizing feedback.
[1344] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1345] Step 1:
[1346] The server collects user information. It receives interest and region information sent by the user from their device via API and stores it in a database. The input is the interest data set by the user (e.g., "technology"), and the output is that data is stored in the database.
[1347] Step 2:
[1348] A script is generated using a generative AI model. The server retrieves the user's interest data from the database and inputs it as a prompt sentence to the generative AI model. The inputs are the "user's interest data" and the "prompt sentence" (e.g., "Tell me the latest technology news"), and the output is a script generated by the generative AI model (e.g., "Today's hot technology news: A new AI technology has been announced").
[1349] Step 3:
[1350] The server inputs the script into a text-to-speech program to generate audio. Specifically, the generated script is input into Google's Text-to-Speech API to obtain audio data. The input at this time is the generated script, and the output is the audio data generated by the text-to-speech program.
[1351] Step 4:
[1352] The server generates an image based on the script. The generated script is input into an image generation AI (DALL-E or Stable Diffusion) to generate a related image. The input is the generated script, and the output is image data generated by the image generation AI.
[1353] Step 5:
[1354] The server combines the audio and image data to generate a video. The audio and image data are input into video editing software (such as FFmpeg) to generate a single video file. The input in this case is the audio and image data, and the output is the generated video file.
[1355] Step 6:
[1356] The server delivers the generated video to the user's device. The generated video file is uploaded to a social media platform (YouTube, Instagram, etc.) and the link is sent to the user's device. The input in this case is the generated video file, and the output is the video link.
[1357] Step 7:
[1358] The user device receives the video and saves it locally. The video is downloaded via the video link sent from the server and saved in local storage. The input is the video link, and the output is the saved video file.
[1359] Step 8:
[1360] The user device plays the stored video. When the user presses the "play" button, the stored video is played using video player software. The input is the stored video file, and the output is the playback of the video.
[1361] Step 9:
[1362] Users provide feedback after watching. After watching a video, users enter their opinions and thoughts in a feedback form and send the data to the server. The input is the user's feedback data, and the output is the feedback information stored on the server.
[1363] (Application example 1)
[1364] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1365] Conventional automatic short video content generation systems, when generating videos optimized for each user, mainly focused on the user's interests and local information, and therefore were unable to adequately stimulate purchasing motivation or promote products in physical stores. In particular, they were unable to create customized content that reflected the customer's purchasing history or real-time interactions in the store, which created challenges in improving the customer experience and promoting sales.
[1366] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1367] In this invention, the server includes means for collecting user information, means for generating a script for each user using a generative AI model, means for generating audio based on the script using a text-to-speech program, means for generating images based on the script using an image generation AI, means for generating a video by combining the generated audio and images, means for delivering the generated video to the user's device and a display in a physical store, means for performing user authentication and setting interest information, and means for collecting feedback after viewing the video. This makes it possible to provide customers visiting a store with a customized promotional video that reflects their purchasing history, interests, and real-time interactions, thereby improving the customer experience and promoting sales.
[1368] "User Information" is data about individual users, including their interests, geographic location, and purchasing history.
[1369] A "generative AI model" is an artificial intelligence model that generates scripts based on user information.
[1370] A "reading program" is software that automatically generates speech based on a generated script.
[1371] "Image generation AI" is an artificial intelligence model that generates relevant images based on a script.
[1372] "User authentication" refers to the means by which a user's identity is verified, which is the prerequisite process for collecting and personalizing user information.
[1373] "Interest information" is data that indicates a user's interests in a particular field.
[1374] "Feedback" refers to the ratings and opinions provided by users after watching a video, and is data that is reflected in the next content generation.
[1375] A "physical store" is a sales facility located in a physical location that offers goods and services.
[1376] A "terminal" is an electronic device that allows a user to receive or send information, including a smartphone, smart glasses, or a display.
[1377] A "display" is a device for displaying visual information, and includes digital signage used in stores.
[1378] This system automatically generates short video content optimized for each user, improving promotions and customer experiences in brick-and-mortar stores. By generating customized videos based on the purchasing history and interests of customers visiting brick-and-mortar stores and providing them via displays or smart glasses, more personalized promotions can be achieved.
[1379] Server Operation
[1380] The server collects user information and generates a script using a generative AI model. Based on the script, a text-to-speech program generates audio and an image generation AI generates images. These are then combined to generate a video, which is then distributed to displays in physical stores and to users' devices.
[1381] Main components
[1382] 1. Collection of User Information
[1383] The server collects user interests, location information, and purchasing history and stores them in a database.
[1384] 2. Script generation using generative AI models
[1385] The server uses a generative AI model to generate a personalized script based on collected user information.
[1386] 3. Integration of text-to-speech programs and image generation AI
[1387] The server inputs the generated script into a speech program to generate audio, and also uses image generation AI to generate related images.
[1388] 4. Video Creation and Distribution
[1389] The server integrates audio and images to generate a video, which is then distributed to displays in physical stores and to users' devices.
[1390] Device behavior
[1391] The device provides user information through an authentication process with the server, receives the video, and stores it locally for offline playback. After watching the video, users can provide feedback, which is reflected in the next content generation.
[1392] Main features
[1393] 1. User Authentication
[1394] The user terminal goes through an authentication process with the server and provides user information.
[1395] 2. Receiving and saving videos
[1396] The user device receives the video delivered from the server and stores it locally.
[1397] 3. Feedback function
[1398] After watching a video, users provide feedback, which is sent to the server and used to generate the next video.
[1399] Hardware and software used
[1400] Server: A Python server application using Flask
[1401] Generative AI models: script generation based on user interests (e.g., GPT-3)
[1402] Text-to-speech program: A text-to-speech tool (e.g., Google Text-to-Speech)
[1403] Image generation AI: Deep learning model (e.g., DALL-E)
[1404] Video generation tools: video editing libraries (e.g., moviepy)
[1405] Devices: Smartphones, smart glasses, in-store displays
[1406] Specific examples
[1407] For example, when a specific customer enters a store, a customized video based on the customer's past purchase history and interests will be displayed on the in-store display, allowing the customer to view new product information and attractive promotions that are tailored to them.
[1408] Example prompt for a generative AI model:
[1409] Customer Interest: "Cooking"
[1410] Prompt for generative AI model: "Generate a short video about the latest cooking appliances and fresh ingredients."
[1411] This system allows physical stores to provide promotional videos optimized for each individual customer, improving the customer experience and promoting sales.
[1412] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1413] Step 1:
[1414] Collection of User Information
[1415] The server collects users' interests, local area information, and purchasing history. The information provided by each user through their device and their in-store purchasing history are stored in a database. This allows the collection of data necessary to generate customized content. The input data are users' interests, local area information, and purchasing history. The output is a database of the collected user information.
[1416] Step 2:
[1417] Script generation using generative AI models
[1418] The server generates a personalized script using a generative AI model based on the collected user information. The script is generated by using data obtained from the user information database as input and providing a prompt to the generative AI model (e.g., "Customer interest: 'Cooking'. Prompt to the generative AI model: 'Generate a short video about the latest cooking equipment and fresh ingredients.'"). The output is the generated customized script.
[1419] Step 3:
[1420] Speech generation by a reading program
[1421] The server inputs the generated script into a text-to-speech program to generate audio. The input data is the generated script, and the output is the generated audio file. The text-to-speech program uses a speech synthesis tool such as Google Text-to-Speech.
[1422] Step 4:
[1423] Image generation using image generation AI
[1424] The server inputs the generated script into the image generation AI, which then generates an image based on it. The input data is the generated script, and the output is a generated image file. The image generation AI model uses a deep learning model such as DALL-E.
[1425] Step 5:
[1426] Video generation
[1427] The server generates a video by integrating the generated audio and images. The input data are the generated audio and image files, and the output is the generated video file. To generate the video, a video editing library such as moviepy is used.
[1428] Step 6:
[1429] Video distribution
[1430] The server distributes the generated video to the user's device and the display in the physical store. The input data is the generated video file, and the output is the video distributed to the user's device and the display in the physical store. Distribution is done via the Internet.
[1431] Step 7:
[1432] User authentication and interest settings
[1433] Users set and update their authentication and interest information on the server through their devices. The input data is the user's authentication information and interest information, and the output is the authenticated user information stored on the server.
[1434] Step 8:
[1435] Providing feedback after watching the video
[1436] Users watch videos through their devices and provide feedback after viewing. The input data is feedback information from the users, and the output is stored in the server's feedback database. This provides data that will be reflected in the next content generation.
[1437] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1438] The system of the present invention automatically generates short video content optimized for each user and distributes it periodically. Furthermore, by combining it with an emotion engine, it can recognize the user's emotions and provide personalized content accordingly, thereby further improving user satisfaction.
[1439] Server Operation
[1440] 1. Collection of User Information
[1441] The server collects user interests and geographical information and stores it in a database. When a user inputs their interests and geographical information, such as "technology" or "Tokyo," through their device, the server receives this information and records it in the user's profile.
[1442] 2. Script generation using generative AI models
[1443] The server retrieves user information from a database and uses a generative AI model to generate a script based on the user's interests, for example, "Today's Hot Tech News."
[1444] 3. Integration of text-to-speech programs and image generation AI
[1445] Based on the generated script, the server uses a text-to-speech program to generate audio and uses image generation AI to generate images that match the script. The generated audio and images are then integrated to generate a video.
[1446] 4. Video distribution
[1447] The server then delivers the generated video to the user's device via the platform the user normally uses, such as a social networking platform or email.
[1448] 5. Gathering Feedback
[1449] After watching a video, users input their impressions and feedback, which are then sent to the server via their device, where they are stored in a database.
[1450] 6. Emotion Recognition by Emotion Engine
[1451] The server passes the collected feedback to the emotion engine to analyze the user's emotions. The emotion engine recognizes the emotions the user felt when giving feedback and stores them in a database.
[1452] 7. Use of Emotional Data
[1453] The server uses the emotion data recognized by the emotion engine to correct the script generation process for the next time. For example, if the user was recognized as "very satisfied" during the previous video viewing, a script with similar content and style will be generated.
[1454] Device behavior
[1455] 1. User Authentication
[1456] The user terminal goes through an authentication process with the server and provides user information, which allows the server to obtain detailed information about the user.
[1457] 2. Receiving and saving videos
[1458] The user's device receives the video streamed from the server and stores it locally, allowing the video to be played even without an internet connection.
[1459] 3. Play the video
[1460] The user device plays the stored video to the user, who then operates the device to start the video and view the provided information.
[1461] 4. Providing Feedback
[1462] After watching the video, users can enter their impressions and feedback through the device. For example, they can enter feedback such as "The content was easy to understand."
[1463] User Actions
[1464] 1. Setting your interests
[1465] Users can set their interests through their devices and send them to the server. For example, they can set interests such as "technology" or "sports."
[1466] 2. Watch the video and give feedback
[1467] After watching a video, users can send feedback about the content to the server via their device, which is reflected in the next content generation.
[1468] Specific examples
[1469] For example, if user A specifies that he is interested in "technology," the server will do the following:
[1470] 1. Collection of User Information
[1471] The server collects user A's interest data and stores it in a database.
[1472] 2. Script generation using generative AI models
[1473] The server retrieves user A's information about "technology" from the database and generates a script using the generative AI model. For example, a script might be generated that reads, "Today's hot technology news: A new AI technology has been announced."
[1474] 3. Integration of text-to-speech programs and image generation AI
[1475] The server inputs the script into a speech program to generate audio.
[1476] Based on the generated script, image generation AI is used to generate related images.
[1477] Integrates audio and images to generate videos.
[1478] 4. Video distribution
[1479] The video generated by the server is distributed to User A's device. For example, distribution can be done using the SNS platform that User A normally uses.
[1480] 5. Gathering Feedback
[1481] User A enters feedback after watching the video. For example, he / she enters a comment including emotion such as "Very satisfied."
[1482] 6. Emotion Recognition by Emotion Engine
[1483] The server passes the feedback to the emotion engine to recognize User A's emotion.
[1484] 7. Use of Emotional Data
[1485] The server then uses the recognized emotion data to correct the script generation process for the next time, so that the next script will be generated in the same style.
[1486] This system makes it possible to further personalize and efficiently deliver content customized for each user based on their emotions, thereby further improving user satisfaction.
[1487] The processing flow will be explained below.
[1488] Step 1:
[1489] The server collects user interests and geographical location information and stores it in a database. For example, if a user enters interests and geographical location such as "technology" or "Tokyo," the server receives this and records it in the user profile.
[1490] Step 2:
[1491] The server periodically retrieves user information from the database. For example, every morning the server can be configured to automatically retrieve user interests and location information.
[1492] Step 3:
[1493] The server then calls up a generative AI model based on user information to generate a script tailored to each individual user. For example, if a user is interested in "technology news," a script containing a summary of the previous day's new technology news will be generated.
[1494] Step 4:
[1495] The server then passes the generated script to a text-to-speech program to generate audio. For example, the server uses the Google Text-to-Speech API to convert the script into audio data.
[1496] Step 5:
[1497] The server then passes the generated script to an image generation AI to generate images for the video, for example, using an image generation model such as DALL-E to generate images related to the news content.
[1498] Step 6:
[1499] The server combines the generated audio and images to create a video, and the program generates a slideshow-style video that displays the images while switching between them against the audio file in the background.
[1500] Step 7:
[1501] The server then distributes the generated video to the platform specified by the user, for example, uploading it to LINE VOOM and obtaining the URL.
[1502] Step 8:
[1503] The server notifies the user's device of the video URL to which the video is being delivered. The user's device receives this notification and saves the video URL locally.
[1504] Step 9:
[1505] The device will use the saved video URL to download the video and store it locally, allowing users to watch the video even without an internet connection.
[1506] Step 10:
[1507] The device plays the stored video for the user to view, and the user operates the device to start the video and view the provided information.
[1508] Step 11:
[1509] After watching a video, users can provide their impressions and feedback via their device, for example, by entering comments such as "The content was easy to understand."
[1510] Step 12:
[1511] The terminal transmits the feedback received from the user to the server, so that the server can utilize the feedback information for generating the next content.
[1512] Step 13:
[1513] The server passes the feedback to an emotion engine to recognize the user's emotion. The emotion engine identifies the user's emotion (e.g., "very satisfied") from the feedback content using, for example, text analysis technology.
[1514] Step 14:
[1515] The server stores the emotion data recognized by the emotion engine in a database, for example, adding emotion data such as "very satisfied" to the user profile.
[1516] Step 15:
[1517] The server uses the recognized emotion data to adjust the script generation process for the next time. For example, if the user was recognized as "very satisfied" during the previous video viewing, the server will generate a script with a similar content style next time.
[1518] Through this process, the system can provide more personalized content based on the user's emotions, thereby improving the quality of the user experience and further increasing satisfaction.
[1519] Example 2
[1520] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1521] In modern content delivery systems, it is important to generate and deliver personalized content optimized for each user. However, conventional systems have difficulty efficiently collecting user sentiment and feedback and reflecting it in the next content generation process. This can easily lead to a decrease in user satisfaction, and more effective personalization is required.
[1522] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1523] In this invention, the server includes means for collecting user information, means for generating a script for each user using a generative AI model, means for generating audio based on the script using a text-to-speech program, means for generating images based on the script using an image generation program, means for generating a video by combining the generated audio and images, means for delivering the generated video to the user's device, means for collecting feedback from the user, means for passing the collected feedback to a sentiment analysis engine to recognize the user's emotions, and means for correcting the next script generation process using the recognized emotion data. This enables the generation and delivery of content optimized for each user, making it possible to provide personalized content that reflects the user's emotions and feedback and provides a high level of satisfaction.
[1524] "User information" refers to data necessary for personalizing content, such as user interests and regional information.
[1525] A "generative AI model" is a model that uses artificial intelligence technology to automatically generate scripts based on user information.
[1526] A "reading program" is a program for generating voice based on a generated script.
[1527] An "image generation program" is a program for generating images based on a generated script.
[1528] "Video" is media content created by combining audio and video.
[1529] "Feedback" refers to the impressions and evaluations provided by users after viewing content.
[1530] An "emotion analysis engine" is an engine that analyzes feedback collected from users and recognizes their emotions.
[1531] "Emotional Data" is data that indicates the emotional state of a user as recognized by an emotion analysis engine.
[1532] The system of the present invention automatically generates short video content optimized for each user and distributes it periodically. The system is mainly composed of three components: a server, a terminal, and a user. The specific operation of each component is explained below.
[1533] Server Operation
[1534] The server first collects user information, including user interests and local area information. For example, if a user inputs their interests such as "technology" or "Tokyo" through their device, the server receives this information and records it in a database.
[1535] The server then uses a generative AI model to generate a script based on the user's information. A natural language generation model such as GPT-4 can be used as the generative AI model. A specific example of a prompt is, "User interest: technology. Please elaborate on the latest technology news in your next script." This prompt causes the AI model to generate a script with appropriate content.
[1536] Based on the generated script, the server uses a text-to-speech program (e.g., Google Text-to-Speech) to generate audio. At the same time, it uses an image generation program (e.g., DALL-E) to generate images related to the script. These audio and images are then combined to generate video content.
[1537] The generated video is then distributed to the user's device by the server via the user's usual social media platform or email. For example, it may be posted on Twitter or Instagram.
[1538] After the user watches the video, the server collects feedback from the user, including their thoughts and feelings, and stores the collected feedback in a database.
[1539] The collected feedback is passed to a sentiment analysis engine (e.g., IBM Watson Emotion Analysis) to recognize the user's emotions. This emotional data is used to adjust the script generation process for the next time. For example, if the feedback for the previous video was "very satisfied," the next time a script will be generated, it will have a similar content and style.
[1540] Device behavior
[1541] The user terminal provides user information through an authentication process with the server, which allows the server to obtain detailed information about the user.
[1542] The user's device receives the video delivered from the server and stores it locally, allowing the video to be played even in an offline environment. The user can operate the device to watch the video and enjoy the provided content.
[1543] After watching the video, the user's device provides an interface to collect feedback from the user, who can then input their impressions and ratings and send them to the server via their device.
[1544] User Actions
[1545] Users can set their interests through their devices, for example by selecting specific interests such as "technology" or "sports," thereby informing the server of their preferences.
[1546] After watching the video, users can input their feedback through their device, which will be reflected in the next content generation, thereby contributing to increasing user satisfaction.
[1547] Specific examples
[1548] For example, if User A specifies that he or she is interested in "technology," the server operates as follows: First, the server collects User A's interest data and stores it in a database. Next, the generative AI model generates a script with content related to "technology." The AI model generates a specific script using the prompt, "User interest: technology. For the next script you create, please detail the latest technology news."
[1549] Based on the script, a reading program generates audio and an image generation program generates related images. The generated audio and images are integrated to create a video, which is then sent to User A's device. After watching the video, User A enters feedback and sends it to the server. The server passes this feedback to an emotion analysis engine, and the recognized emotional data is reflected in the next script generation.
[1550] This makes it possible to generate and deliver personalized content optimized for each user, thereby increasing user satisfaction.
[1551] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1552] System program processing flow
[1553] Step 1: Collect user information
[1554] The server receives information about the user's interests and local area from the device. When the user inputs information such as "technology" or "Tokyo" through the device, the server records this information in a database.
[1555] Input: User interests and location information
[1556] Output: User profile information stored in the database
[1557] Specific operation: The user selects "technology" on the device settings screen, and the information is sent to the server and stored in a database.
[1558] Step 2: Generative AI model generates script
[1559] The server sends prompts to the generative AI model based on the collected user information to generate a script based on the user's interests. For example, "User interest: technology. For your next script, please detail the latest technology news."
[1560] Input: User information retrieved from the database, prompt text
[1561] Output: Generated script
[1562] How it works: The server inputs the prompt "User interest: Technology" into a generative AI model (e.g., GPT-4), and the model generates a related news script.
[1563] Step 3: Integrating the Text-to-Speech Program and the Image Generator
[1564] The server generates audio based on the generated script using a reading program and generates associated images using an image generation program.
[1565] Input: Generated script
[1566] Output: Generated audio and image files
[1567] Specific operation: The server inputs the script into a speech program (e.g., Google Text-to-Speech) to generate an audio file. In parallel, the server inputs the script into an image generation program (e.g., DALL-E) to generate related images.
[1568] Step 4: Audio and video integration
[1569] The server combines the generated audio and images to create a video.
[1570] Input: Generated audio and image files
[1571] Output: Finished video file
[1572] Specific operation: The server places the audio and image files on the timeline and exports them as a single video file.
[1573] Step 5: Publish your video
[1574] The server then distributes the generated video to the user's device via the user's usual social media platform or email.
[1575] Input: Finished video file
[1576] Output: Video delivered to the user's device
[1577] Specific operation: The video file generated by the server is posted to Twitter or Instagram, or sent by email.
[1578] Step 6: Gather feedback
[1579] After watching a video, users can enter their feedback and send it to the server via their device, which then stores it in a database.
[1580] Input: User feedback
[1581] Output: Feedback stored in a database
[1582] Specific operation: The user enters feedback such as "very satisfied" on the device, and the server receives it and records it in the database.
[1583] Step 7: Emotion Recognition with the Emotion Engine
[1584] The server passes the collected feedback to an emotion engine to analyze the user's emotions.
[1585] Input: User feedback data
[1586] Output: Recognized emotion data
[1587] Specific operation: The server passes the feedback "very satisfied" to an emotion engine (e.g., IBM Watson Emotion Analysis), which then recognizes the emotion "satisfied" and stores it in a database.
[1588] Step 8: Use emotion data
[1589] The server uses the emotion data obtained from the emotion engine to correct the next script generation process.
[1590] Input: Recognized emotion data
[1591] Output: Corrected script generation process
[1592] Specific operation: The server sets the conditions for the next script generation based on the "satisfaction" information and reflects this when sending a new prompt sentence to the generation AI model.
[1593] (Application example 2)
[1594] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1595] In today's world, content distribution services must provide optimal content for each individual user. However, with conventional technologies, content generated based on user interests and local information does not adequately reflect user emotions or real-time feedback. This makes it difficult to improve user satisfaction, and there is a need to further increase engagement for each user.
[1596] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for collecting user information, means for generating a script for each user using a generation AI model, means for generating audio based on the script using a text-to-speech program, means for generating images based on the script using an image generation AI, means for generating a video by combining the generated audio and images, means for delivering the generated video to the user's device, means for recognizing the user's emotion using an emotion recognition engine, and means for reflecting the recognized emotion data in the next content generation. This makes it possible to deliver personalized short video content based not only on the user's interests and regional information but also on their emotions.
[1597] "Means for collecting user information" refers to a mechanism for collecting data including user interests and regional information and storing it on a server.
[1598] A "generative AI model" is a system that includes an algorithm that automatically generates the optimal script based on each user's interests and regional information.
[1599] A "means for generating a script" is a process that uses a generative AI model to create a script appropriate for a particular user.
[1600] A "reading program" is software that automatically generates speech based on a generated script.
[1601] "Image generation AI" is an artificial intelligence system that automatically creates images that match the content of a generated script.
[1602] "Means for generating video" refers to technology that integrates generated audio and images to create a single video.
[1603] "Means for delivering video" refers to the infrastructure for transmitting the generated video to the user's device.
[1604] An "emotion recognition engine" is a system that includes an algorithm for analyzing and recognizing emotions from a user's facial expressions, voice, etc.
[1605] "Means for reflecting emotional data in the next content generation" refers to a process of using the recognized emotional data of the user to provide appropriate feedback to the script or content to be generated next.
[1606] This invention is a system that automatically generates short video content optimized for each user and distributes it periodically. Furthermore, by combining it with an emotion engine, it recognizes the user's emotions and provides personalized content accordingly. The various means required for this and their specific processing steps are described below.
[1607] Collection of User Information
[1608] The server first collects the user's interests and area information. The user enters their interests and area of residence through their device, and the information is sent to the server. This information is stored in a database and used for subsequent processing.
[1609] Script generation
[1610] The server uses the collected user information to generate a script using a generative AI model. This generative AI model inputs the user's interests and local information as prompts and outputs a short script based on that data. For example, a prompt might be, "The user's interest is 'technology.' Please generate a script to deliver today's latest technology news to him."
[1611] Sound and image generation
[1612] The server inputs the generated script into a text-to-speech program to generate audio. It then uses an image generation AI to generate images based on the script. Specifically, the text-to-speech program and image generation AI use a TextToSpeech engine and an ImageGenerator.
[1613] Video Creation and Delivery
[1614] The server combines the generated audio and video to generate a video, which is then distributed to the user's device. The video is generated using a video editing software library and distributed via social media platforms or email.
[1615] Emotion recognition and feedback collection
[1616] After the user watches the video, the device uses an emotion recognition engine to analyze the user's emotions. This engine recognizes emotions in real time through facial expressions and voice analysis, and sends the data to the server. For example, it collects emotional feedback such as "very satisfied."
[1617] Reflection in next script generation
[1618] The server reflects the collected emotional data in the next content generation, correcting the script generation process based on the previous feedback and generating content that is in line with the user's preferences and emotions.
[1619] Thus, the present invention can significantly improve user satisfaction by efficiently generating and delivering personalized short video content that takes user emotions into consideration.
[1620] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1621] Step 1: Collect user information
[1622] Users input their interests and local area information through their devices and send that information to the server. The server stores the received data in a database. This input information includes the user's content preferences and location information. For example, a user might input information that they are interested in "technology" and live in "Tokyo."
[1623] Step 2: Generate the script
[1624] The server generates a script using a generative AI model based on the user information collected in step 1. In this process, the server inputs a prompt such as "The user's interest is 'technology'. Please generate a script that conveys today's latest technology news to him" into the generative AI model, and obtains the optimal script as the output.
[1625] Step 3: Generate audio
[1626] The server inputs the script generated in step 2 into a speech program (such as a TextToSpeech engine) to automatically generate speech. Based on the input, acoustic analysis and synthesis are performed, and the output is a generated audio file.
[1627] Step 4: Generate images
[1628] The server inputs the generated script into an image generation AI (such as ImageGenerator) to generate an image related to the script. The image generation AI analyzes the contents of the script and generates and outputs an appropriate image. For example, it generates an image related to an article about the latest AI technology.
[1629] Step 5: Generate the video
[1630] The server generates a video by combining the audio generated in step 3 with the images generated in step 4. This process uses a video editing software library to combine the audio and images in a sequence and output a single video file containing technology news.
[1631] Step 6: Publish your video
[1632] The server distributes the generated video to the user's device. Distribution is done via the user's usual social media platform or email. Based on the distribution destination information, the server sends the video file to the corresponding platform. The user then watches the video using their device.
[1633] Step 7: Recognize emotions
[1634] While the user is watching a video, the device uses a built-in emotion recognition engine to analyze the user's emotions in real time. This process involves using a camera and microphone to collect the user's facial expressions and tone of voice, and then inputting this data into the emotion recognition algorithm. The output is the user's emotional data.
[1635] Step 8: Gather feedback
[1636] After viewing, users input their feedback and emotional data into their device and send it to the server. The server stores this in a database. The collected feedback is reflected in the next content generation process. For example, emotional feedback such as "very satisfied" can be input.
[1637] Step 9: Reflecting the changes in the next script generation
[1638] The server uses the emotion data collected in step 8 to correct the next content generation process. This ensures that the next script generated is in line with the user's preferences and emotions. This makes it possible to continuously improve user satisfaction.
[1639] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1640] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1641] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1642] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1643] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1644] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1645] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1646] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, motorcycles, and other devices, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1647] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1648] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1649] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1650] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1651] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1652] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1653] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1654] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1655] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1656] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1657] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1658] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1659] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1660] The following is further disclosed regarding the above embodiment.
[1661] (Claim 1)
[1662] the means by which user information is collected;
[1663] A means for generating scripts for each user using a generative AI model;
[1664] means for generating speech based on the script using a text-to-speech program;
[1665] A means for generating images based on a script using image generation AI;
[1666] means for generating a video by combining the generated audio and images;
[1667] means for delivering the generated video to a user's terminal;
[1668] A system including:
[1669] (Claim 2)
[1670] 10. The system of claim 1, wherein the user information includes user interests and geographical area information.
[1671] (Claim 3)
[1672] The system of claim 1, wherein the generative AI model generates a script based on user interests and regional information.
[1673] (Claim 4)
[1674] 2. The system of claim 1, wherein the video is distributed to users periodically.
[1675] (Claim 5)
[1676] The system of claim 1, wherein the reading program and image generation AI use external services.
[1677] "Example 1"
[1678] (Claim 1)
[1679] the means by which user information is collected;
[1680] A means for generating scripts for each user using a generative AI model;
[1681] means for generating speech based on the script using a text-to-speech program;
[1682] A means for generating images based on a script using image generation AI;
[1683] means for generating a video by combining the generated audio and images;
[1684] means for delivering the generated video to a user's terminal;
[1685] means for receiving and storing the generated video on a user's terminal;
[1686] means for playing the generated video on a user's terminal;
[1687] A way for users to provide feedback after viewing;
[1688] A system including:
[1689] (Claim 2)
[1690] 10. The system of claim 1, wherein the user information includes user interests and geographical area information.
[1691] (Claim 3)
[1692] The system of claim 1, wherein the generative AI model generates a script based on user interests and regional information.
[1693] "Application Example 1"
[1694] (Claim 1)
[1695] the means by which user information is collected;
[1696] A means for generating scripts for each user using a generative AI model;
[1697] means for generating speech based on the script using a text-to-speech program;
[1698] A means for generating images based on a script using image generation AI;
[1699] means for generating a video by combining the generated audio and images;
[1700] means for delivering the generated video to a user's terminal and a display in a physical store;
[1701] a means for authenticating a user and setting user interests;
[1702] A means to collect feedback after watching the video, and
[1703] A system including:
[1704] (Claim 2)
[1705] 2. The system of claim 1, wherein the user information includes user interests, geographical location information, and purchasing history.
[1706] (Claim 3)
[1707] The system of claim 1, wherein the generative AI model generates a script based on a user's interests, regional information, and purchasing history.
[1708] "Example 2: Combining Emotion Engines"
[1709] (Claim 1)
[1710] the means by which user information is collected;
[1711] A means for generating scripts for each user using a generative AI model;
[1712] means for generating speech based on the script using a text-to-speech program;
[1713] means for generating images based on a script using an image generation program;
[1714] means for generating a video by combining the generated audio and images;
[1715] means for delivering the generated video to a user's terminal;
[1716] a means of gathering user feedback;
[1717] A means of passing the collected feedback to a sentiment analysis engine to recognize user sentiment;
[1718] means for correcting a subsequent script generation process using the recognized emotion data;
[1719] A system including:
[1720] (Claim 2)
[1721] 10. The system of claim 1, wherein the user information includes user interests and geographical area information.
[1722] (Claim 3)
[1723] The system of claim 1, wherein the generative AI model generates a script based on user interests and regional information.
[1724] "Application example 2 when combining emotion engines"
[1725] (Claim 1)
[1726] the means by which user information is collected;
[1727] A means for generating scripts for each user using a generative AI model;
[1728] means for generating speech based on the script using a text-to-speech program;
[1729] A means for generating images based on a script using image generation AI;
[1730] means for generating a video by combining the generated audio and images;
[1731] means for delivering the generated video to a user's terminal;
[1732] means for recognizing a user's emotion using an emotion recognition engine;
[1733] a means for reflecting the recognized emotion data in next content generation;
[1734] A system including:
[1735] (Claim 2)
[1736] 10. The system of claim 1, wherein the user information includes user interests and geographical area information.
[1737] (Claim 3)
[1738] The system of claim 1, wherein the generative AI model generates a script based on user interests and regional information. [Explanation of symbols]
[1739] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. the means by which user information is collected; A means for generating scripts for each user using a generative AI model; means for generating speech based on the script using a text-to-speech program; A means for generating images based on a script using image generation AI; means for generating a video by combining the generated audio and images; means for delivering the generated video to a user's terminal; A system including:
2. 10. The system of claim 1, wherein the user information includes user interests and geographical location information.
3. 2. The system of claim 1, wherein the generative AI model generates a script based on user interests and regional information.
4. 2. The system of claim 1, wherein the video is distributed to users periodically.
5. The system according to claim 1, wherein the reading program and the image generation AI use external services.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A