System
The system addresses the challenge of fixed news formats by allowing users to choose genres and customize virtual announcers, using AI to generate personalized news streams, enhancing viewer engagement and satisfaction through user feedback mechanisms.
Patent Information
- Application Number
- JP2024119078
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-24
- Publication Date
- 2026-02-05
AI Technical Summary
Traditional news broadcasts are organized with fixed announcers and set news genres, making it difficult to accommodate diverse viewer preferences and provide personalized experiences, leading to reduced engagement and satisfaction.
A system that allows users to select their favorite news genres and customize virtual announcers, using a generative AI model to generate a speech script, and provide news in a live streaming format, with mechanisms for collecting user feedback to improve the service.
Enables personalized news experiences, increasing viewer engagement and satisfaction by providing customized news content based on user preferences and continuously improving the service quality.
Smart Images

Figure 2026018017000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] Describe the "problem that the invention aims to solve" and the "means for solving the problem."
[0005] Traditional news broadcasts are organized with fixed announcers and set news genres, making it difficult to fully accommodate the diverse preferences of viewers. This makes it difficult to provide news that is in line with viewers' interests and increases engagement. Furthermore, because fixed programming makes it difficult to provide a personalized experience, some viewers may feel that they are overwhelmed or lacking information. [Means for solving the problem]
[0006] The present invention is a system that provides each viewer with a personalized news experience by providing a means for users to select their favorite news genres and a means for users to customize their virtual announcer. This system includes a means for acquiring news data and filtering it based on the user's selected genre, a means for generating a speech script from the filtered news data, a means for having a virtual announcer read the generated speech script, and a means for providing news to users in a live streaming format, thereby increasing engagement. Furthermore, by providing a means for collecting user feedback and using it to improve the service, and a means for automatically generating a speech script from news data using a generative AI model, it is possible to continuously provide more attractive content to viewers.
[0007] A "user" is someone who uses the system to select a news genre and virtual announcer and watch customized news.
[0008] "Genres" are categories that categorize news types and are specific areas of interest to viewers, such as sports, technology, or politics.
[0009] A "virtual announcer" is a computer-generated virtual announcer, a character that reads the news with a user-customized appearance and voice.
[0010] "News data" is a collection of news articles and information obtained from multiple news sources.
[0011] "Filtering" is the process of selecting necessary news from acquired news data based on the genre selected by the user.
[0012] A "reading script" is text data generated by a generative AI model based on a news article, which is then read aloud by a virtual announcer.
[0013] A "generative AI model" is an algorithm or software that uses machine learning techniques to automatically generate a reading transcript from a news article.
[0014] "Live streaming" is a technology that delivers video and audio to users in real time via the Internet.
[0015] "Feedback" refers to opinions and ratings provided by users after viewing, and is information that can be used to improve the quality of the service.
[0016] "Customization" is the act of a user changing selections and settings to suit their own preferences. [Brief explanation of the drawings]
[0017] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8]FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0018] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0019] First, the terms used in the following description will be explained.
[0020] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0021] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0022] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0023] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0024] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0025] [First embodiment]
[0026] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0027] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0028] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0029] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0030] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0031] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0032] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0033] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0034] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0035] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0036] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0037] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0038] ---
[0039] This invention is a system for providing viewers with a personalized news experience, allowing users to select a news genre and a virtual announcer and watch customized live news. This system has functions that allow users to select news genres and virtual announcers according to their preferences, and functions that automatically generate a text-to-speech script from news data and have the virtual announcer read it aloud. It also has a function for improving the service based on user feedback.
[0040] System configuration
[0041] User terminal
[0042] Users install the news distribution application on their device, create an account, and then log in. Once successfully logged in, users can select news genres and virtual announcers. The user's selection information is saved on the device and sent to the server.
[0043] server
[0044] The server has the following main functions:
[0045] 1. News data collection:
[0046] The server retrieves the latest news data from various news APIs. This data is aggregated from multiple news sources.
[0047] 2. Filtering and generating a transcript:
[0048] The collected news data is filtered based on the news genre selected by the user, and then a generative AI model is used to automatically generate a reading transcript from the filtered news articles.
[0049] 3. Creating a virtual announcer:
[0050] The virtual announcer reads the script based on the user-customized virtual announcer settings (voice, appearance, etc.), and then generates the final video material for live streaming.
[0051] 4. Live Streaming:
[0052] The generated virtual announcer will broadcast live news videos in real time over the Internet.
[0053] 5. Collecting User Feedback:
[0054] After watching the news, we collect feedback from users to help improve the service. The feedback information is stored on the server and can be used for data analysis.
[0055] Specific examples
[0056] If the user selects sports news
[0057] 1. User:
[0058] Install the news distribution app and create an account. After logging in, select the "Sports" genre on the settings screen and then set your preferred virtual announcer (e.g., a male voice with an active personality).
[0059] 2. Terminal:
[0060] The selection information is stored on the local device and transmitted to the server.
[0061] 3. Server:
[0062] Sports-related news data is collected and filtered. Then, using a generative AI model, a script is automatically generated from the filtered news data. The final video material is generated based on the script and the settings of the virtual announcer selected by the user.
[0063] 4. Terminal:
[0064] The user's device receives the streaming URL from the server and plays the sports news in real time.
[0065] 5. User:
[0066] On the way to work, users can open the app and watch a virtual announcer announce the results of last night's game. After watching, users can provide feedback within the app.
[0067] This system allows for a personalized news experience, increasing viewer engagement.
[0068] The processing flow will be explained below.
[0069] ---
[0070] Step 1:
[0071] Users install the news delivery application on their device and create an account. After creating the account, they log in by entering their email address and password on the login page.
[0072] Step 2:
[0073] The terminal sends the user's login information to the server for authentication. If authentication is successful, the server issues a session ID and returns it to the terminal. The terminal saves the session ID and maintains the logged-in state.
[0074] Step 3:
[0075] Users access the app's settings screen, select their preferred news genre (e.g., sports, technology, politics, etc.), customize their preferred virtual announcer (e.g., gender, tone of voice, appearance, etc.), and click the "Save" button.
[0076] Step 4:
[0077] The terminal transmits the user's selected news genre and virtual announcer setting information to the server, which then stores the received customization information in the user database.
[0078] Step 5:
[0079] The server retrieves the latest news data from multiple news APIs, filters the retrieved news data based on the genre selected by the user, and extracts the corresponding news data.
[0080] Step 6:
[0081] The server uses a generative AI model to automatically generate a speech transcript from the filtered news data, which is then stored in a database.
[0082] Step 7:
[0083] The server retrieves the information of the virtual announcer selected by the user and combines it with the script to generate the final video material. Specifically, it generates the reading voice using a voice synthesizer and creates the character movements using 3D modeling.
[0084] Step 8:
[0085] The server uploads the generated news video of the virtual announcer to the live streaming server and generates a stream URL, which is then sent to the device.
[0086] Step 9:
[0087] The device plays the video in real time based on the stream URL received from the server, and users can open the app at their preferred time to watch customized live news.
[0088] Step 10:
[0089] After watching the news, users can rate the announcer's performance and the content of the news in the feedback form within the app and enter the feedback information on their device, which then sends the feedback information to the server.
[0090] Step 11:
[0091] The server stores user feedback information in a database and uses it for data analysis to help improve generative AI models and services.
[0092] This completes the system's processing flow, allowing users to enjoy a personalized news experience in real time.
[0093] Example 1
[0094] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0095] Conventional news delivery systems make it difficult for viewers to select news that suits their preferences and watch it in a customized format. Furthermore, they lacked a mechanism for effectively collecting viewer feedback and utilizing it to improve services. As a result, they were unable to fully optimize the viewing experience or improve viewer satisfaction.
[0096] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0097] In this invention, the server includes a means for acquiring news data and filtering it based on a genre selected by the user, a means for generating a script from the filtered news data using a generative AI model, and a means for having a virtual announcer read the generated script. This allows users to select news from their favorite genres and watch live news broadcasts by customized virtual announcers. Furthermore, by collecting user feedback and using it to improve the service, viewer satisfaction can be increased.
[0098] 1. "A means for users to select news genres they like" refers to a function that allows users to select news categories that interest them.
[0099] 2. "Means for users to customize the virtual announcer" refers to the functionality that allows users to change settings such as the voice and appearance of the virtual announcer.
[0100] 3. "Means of acquiring news data" refers to the function of collecting the latest news information from external news data providers.
[0101] 4. "Means of filtering based on the genre selected by the user" refers to the function of extracting only information that falls into the category selected by the user from the collected news data.
[0102] 5. "Means for generating a text-to-speech transcript from filtered news data" refers to a function that uses a generative AI model to summarize filtered news information and create text that can be read aloud.
[0103] 6. "Means for having a virtual announcer read the generated script" refers to the function of converting text data into voice using speech synthesis technology and synchronizing it with the actions of the virtual announcer.
[0104] 7. "Means for providing users with generated news videos in a live streaming format" refers to the function of delivering generated audio and video to users in real time via the Internet.
[0105] 8. "Means of collecting user feedback" refers to the function of collecting ratings and comments from viewers and storing them in a database.
[0106] 9. "Means used to improve services" refers to the functions used to analyze collected feedback and improve the quality of news provision services.
[0107] 10. "Generative AI model" refers to an artificial intelligence model that generates text data using natural language processing technology.
[0108] 11. “Prompt” refers to a command or question input to a generative AI model.
[0109] These are the definitions of the important words included in this system.
[0110] The present invention is a system for providing users with a personalized news experience, allowing them to select a news genre and a virtual announcer and watch customized live news. This system has functions that allow users to select news genres and virtual announcers according to their preferences, and functions that automatically generate a text-to-speech script from news data and have the virtual announcer read it aloud. It also has a function to improve the service based on user feedback.
[0111] System configuration
[0112] User terminal
[0113] Users install the news distribution application on their device, create an account, and then log in. Once successfully logged in, users can select news genres and virtual announcers. The user's selection information is saved on the device and sent to the server. This can be done using common mobile devices such as smartphones and tablets.
[0114] server
[0115] The server has the following main functions:
[0116] 1. News data collection:
[0117] The server retrieves the latest news data from various news providers. This data is collected from multiple news sources. Specifically, "News API" and "News Data Provider Services" are used.
[0118] 2. Filtering and generating a transcript:
[0119] The collected news data is filtered based on the news genre selected by the user. Then, a generative AI model is used to automatically generate a speech transcript from the filtered news articles. The generative AI model used here is, for example, GPT-4.
[0120] 3. Creating a virtual announcer:
[0121] The virtual announcer reads the script based on the user's customized settings (voice, appearance, etc.). The final video material for live streaming is then generated. Speech synthesis software (e.g., TTS technology) is used for the speech synthesis, and 3D animation software (e.g., animation generation software) is used to generate the video.
[0122] 4. Live Streaming:
[0123] The generated news video of the virtual announcer is broadcast live. The live broadcast is carried out in real time over the Internet. A "live streaming service," for example, is used as the streaming platform.
[0124] 5. Collecting User Feedback:
[0125] After watching the news, we collect feedback from users to help improve the service. The feedback information is stored on the server and can be used for data analysis.
[0126] Specific examples
[0127] If the user selects sports news
[0128] 1. User:
[0129] Install the news distribution app and create an account. After logging in, select the "Sports" genre on the settings screen and then set your preferred virtual announcer (e.g., a male voice with an active personality).
[0130] 2. Terminal:
[0131] The selection information is stored on the local device and transmitted to the server.
[0132] 3. Server:
[0133] Sports-related news data is collected and filtered. Then, using a generative AI model, a script is automatically generated from the filtered news data. The final video material is generated based on the script and the settings of the virtual announcer selected by the user.
[0134] 4. Terminal:
[0135] The user's device receives the streaming URL from the server and plays the sports news in real time.
[0136] 5. User:
[0137] On the way to work, users can open the app and watch a virtual announcer announce the results of last night's game. After watching, users can provide feedback within the app.
[0138] Prompt Sentence Examples
[0139] "I want to watch sports news. I want the virtual announcer to be a lively male."
[0140] This system will enable viewers to enjoy a news experience that is optimized for each individual, increasing engagement.
[0141] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0142] Step 1:
[0143] A user installs a news delivery app and creates an account
[0144] Input: User's electronic device (smartphone, tablet) and internet connection
[0145] How it works: A user downloads and installs a news delivery app from the app store. They launch the app and create an account by entering their email address and password on the account creation screen that appears when the app is first launched.
[0146] Output: Account information is created and you can log in to the app. After logging in, you will be taken to a screen where you can select the news genre and virtual announcer.
[0147] Step 2:
[0148] The user selects the news genre and virtual announcer.
[0149] Input: User preference information (selected news genre, virtual announcer settings)
[0150] How it works: Within the app, users select "Sports" from a range of news genres, including "Sports," "Politics," and "Entertainment." They then customize the virtual announcer's settings screen, such as selecting a "lively male voice."
[0151] Output: The selection information is temporarily stored in the device's local storage and then sent to the server.
[0152] Step 3:
[0153] The server collects news data
[0154] Input: API requests and responses received by the server from external news providers
[0155] How it works: The server sends a request to an API endpoint such as a "news data provider service" to retrieve the latest news data in JSON format.
[0156] Output: The retrieved news data is stored in a database for further processing.
[0157] Step 4:
[0158] The server filters the news data and generates a reading script.
[0159] Input: News data to be filtered and user genre selection information
[0160] How it works: The server uses a filtering algorithm (e.g., TF-IDF) to extract only news data that falls within the user's selected "sports" genre. It then uses a generative AI model (e.g., GPT-4) to generate a summary of the filtered news data.
[0161] Output: The filtered news data and the generated speech transcripts are stored in a database.
[0162] Step 5:
[0163] The server configures the virtual announcer and generates the news video.
[0164] Input: Reading script and user's virtual announcer setting information
[0165] How it works: The server uses speech synthesis software (e.g., TTS technology) to convert the generated script into audio data, and then uses 3D animation software (e.g., animation generation software) to generate a video that coordinates the voice and movements of the virtual announcer.
[0166] Output: The created news video is stored on the server and ready to be streamed.
[0167] Step 6:
[0168] Server-generated live news video broadcast
[0169] Input: Generated news video data
[0170] How it works: The server uses a live streaming service to create a streaming URL for the generated news video, and then sends this URL to the user's device.
[0171] Output: A response containing a streaming URL is sent to the user's device.
[0172] Step 7:
[0173] User watches news
[0174] Input: Streaming URL
[0175] How it works: The user's device receives the streaming URL and uses the media player component to play the news video in real time.
[0176] Output: User watches live streaming news video.
[0177] Step 8:
[0178] Users provide feedback
[0179] Input: User feedback information (ratings, comments, etc.)
[0180] How it works: After watching a news item, users enter their rating and comments about their viewing experience on the feedback screen displayed within the app. The feedback is then sent from the app to the server.
[0181] Output: The feedback information is sent to the server and stored in a database.
[0182] The above is a detailed description of the specific processing steps of the present system. The present invention is expected to provide users with an individually customized news experience, thereby improving viewer engagement and satisfaction.
[0183] (Application example 1)
[0184] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0185] In today's world, with so much news being distributed, it is difficult for individual users to accurately obtain only the news that interests them. There are also insufficient means to customize the visual and auditory experience of news for each individual user. Furthermore, there is a lack of mechanisms for users to provide feedback on the news they have viewed and use that feedback to improve services. Effective methods to solve these issues are needed.
[0186] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0187] In this invention, the server includes means for acquiring news data and filtering it based on a genre selected by the user, means for generating a script to be read from the filtered news data, and means for having a virtual announcer read the generated script to be read. This allows each user to watch news in the genres they are interested in in real time with a customized virtual announcer. Furthermore, it is possible to stream news to a smartphone application in real time and collect feedback from users to help improve the service.
[0188] "User" refers to an individual who uses this system to receive news distribution services.
[0189] "Genre" refers to the type or category of news, such as a specific field such as sports, economics, or entertainment.
[0190] A "virtual announcer" is a computer-generated character whose role is to read the news aloud and visually.
[0191] "News Data" refers to the collection of information collected from various news sources and refers to the material used for filtering and generating the readings.
[0192] "Filtering" refers to the process of narrowing down news data based on genres selected by the user.
[0193] A "script" is a text automatically generated from filtered news data, and refers to a script that a virtual announcer will read aloud.
[0194] "Live streaming" refers to a format in which news is provided to users in real time, and refers to a method of streaming directly from a server to a user's terminal.
[0195] "Smartphone application" refers to a program that runs on a smartphone and functions as the user interface for this system.
[0196] "Real-time streaming" refers to a technology in which news videos are sent to the user's device as soon as they are generated and played back instantly.
[0197] "Feedback" refers to the opinions and evaluations provided by users regarding the news and services they have viewed, and serves as important data for improving services.
[0198] A "generative AI model" refers to an artificial intelligence algorithm that uses machine learning technology to automatically generate a reading script from news data.
[0199] The present invention is a system that allows users to customize the news genres and virtual announcers that interest them, providing a personalized news experience.
[0200] System configuration
[0201] User terminal
[0202] Users install the corresponding application on their smartphone, create an account, and log in. After logging in, users can configure their news genre and virtual announcer settings. This includes selecting a genre (e.g., sports, economics, entertainment, etc.) and customizing the announcer's appearance and voice. This configuration information is saved on the user's device and sent to the server.
[0203] server
[0204] The server has the following main functions:
[0205] 1. News data collection:
[0206] The server uses news APIs to collect the latest news data from various news sources, examples of which include NewsAPI and Google News API.
[0207] 2. Filtering and generating a transcript:
[0208] The collected news data is filtered based on the genre selected by the user, and a transcription is automatically generated from the filtered news articles using a generative AI model (e.g., GPT-3).
[0209] 3. Creating a virtual announcer:
[0210] A virtual announcer reads the script according to the announcer settings customized by the user, using tools such as Adobe Character Animator and Unity3D. The generated news video is then streamed to the user's device in real time.
[0211] 4. Gathering Feedback:
[0212] After viewing the news, we collect feedback from users and use it to improve our services, thereby continuously improving the quality of news and its delivery methods.
[0213] Process Overview
[0214] 1. User Settings:
[0215] For example, a user selects "Entertainment News" and "An announcer with a female voice and a cheerful personality." The user's selection information is sent from the smartphone terminal to the server.
[0216] 2. News gathering and filtering:
[0217] The server uses NewsAPI to collect the latest entertainment news and filters the collected data.
[0218] 3. Generate a reading transcript:
[0219] A generative AI model (GPT-3) is used to generate a read-aloud transcript from filtered news articles. For example, a sentence such as "A blockbuster film was announced at the film festival yesterday" is generated.
[0220] 4. Creating a Virtual Announcer:
[0221] Using Adobe Character Animator and Unity3D, a user-customized virtual announcer is generated, and a video is created in which the announcer reads the script.
[0222] 5. Real-time Streaming and Viewing:
[0223] The generated news videos are hosted on a server and streamed in real time to smartphones, where users open the app and listen to the news being read by an announcer.
[0224] 6. Collecting Feedback and Improving Our Services:
[0225] After watching, users provide feedback that helps improve the service.
[0226] Examples and prompts
[0227] Examples:
[0228] A user opens the app and selects a sports news item and a calm, male-voiced announcer. The server collects the latest sports news and uses a generative AI model (GPT-3) to create a script, such as "I'll tell you the results of last night's game." Using Unity3D, a virtual announcer reads the script, and the generated video is delivered to a smartphone. The user can watch the news on their way to work and provide feedback.
[0229] Prompt for the generative AI model:
[0230] News Cache: {Sports}
[0231] Announcer Voice: {Male voice, calm}
[0232] The statement read: "Here are the results of last night's match."
[0233] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0234] Step 1:
[0235] A user creates an account and logs into the smartphone application. The user selects a news genre (e.g., sports, economics, entertainment) and virtual announcer settings (e.g., voice type, gender, appearance). The selection information is stored on the user's device and sent to the server.
[0236] Input: User's genre selection information and virtual announcer settings
[0237] Output: Selections sent to the server
[0238] Specific behavior:
[0239] The user opens the app and navigates to the settings screen.
[0240] Use the drop-down menus and sliders to set your preferred news genres and anchor appearance and voice.
[0241] After the setting information is confirmed, pressing the "Save" button saves the selected information in the local storage of the device and simultaneously sends it to the server.
[0242] Step 2:
[0243] The server uses a news API to collect the latest news data for the genre selected by the user. For example, it uses NewsAPI or Google News API.
[0244] Input: News API request
[0245] Output: Collected news data
[0246] Specific behavior:
[0247] The server accesses the news API periodically or upon user request to collect the latest news data for the selected genre.
[0248] The collected news data is stored in a database on the server.
[0249] Step 3:
[0250] The collected news data is filtered based on the genre selected by the user.
[0251] Input: Collected news data, user genre preferences
[0252] Output: Filtered news data
[0253] Specific behavior:
[0254] An algorithm on the server analyzes the collected news data and extracts only articles related to the genre selected by the user.
[0255] The filtered results are stored in a database.
[0256] Step 4:
[0257] A reading script is automatically generated from filtered news data using a generative AI model (e.g., GPT-3).
[0258] Input: Filtered news data
[0259] Output: Speech manuscript
[0260] Specific behavior:
[0261] The server provides a prompt to the generative AI model, which generates a reading script based on the filtered news data.
[0262] For example, the prompt text is:
[0263] News Cache: {Sports}
[0264] Announcer Voice: {Male voice, calm}
[0265] The statement read: "Here are the results of last night's match."
[0266] The generated reading manuscript is stored in a database on the server.
[0267] Step 5:
[0268] Based on the user's customized settings, a video is generated for the virtual announcer to read the script.
[0269] Input: Reading script, user's virtual announcer settings
[0270] Output: News video with a virtual announcer
[0271] Specific behavior:
[0272] Using Adobe Character Animator and Unity3D, a virtual announcer is generated based on the announcer settings selected by the user.
[0273] The script is read aloud by an announcer, and then generated in video format.
[0274] The generated video is stored on the server.
[0275] Step 6:
[0276] The generated news video of the virtual announcer is streamed in real time and provided to users.
[0277] Input: News video with a virtual announcer
[0278] Output: Real-time streaming video played on user devices
[0279] Specific behavior:
[0280] The server hosts the generated news video and generates a streaming URL.
[0281] The user's device accesses this URL and watches the news video in real time.
[0282] Step 7:
[0283] After watching the news, we collect feedback from users to help improve our services.
[0284] Input: User feedback
[0285] Output: Data for service improvement
[0286] Specific behavior:
[0287] After users watch a news video, they can provide their ratings and opinions through an in-app feedback form.
[0288] The server analyzes the collected feedback data and uses it to improve the service.
[0289] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0290] ---
[0291] This invention is a system that achieves even higher levels of personalization by combining a personalized news delivery system with a user-selected news genre and customized virtual announcer with an emotion engine that recognizes the user's emotions. This system recognizes the user's emotions in real time and, based on that information, adjusts the virtual announcer's facial expressions and tone of voice, and recommends related news content.
[0292] System configuration
[0293] User terminal
[0294] Users install the news distribution application on their device, create an account, and then log in. Once successfully logged in, they can select and customize the news genre and virtual announcer on the settings screen. The device is also equipped with a camera and microphone, which are used to recognize the user's emotions in real time.
[0295] server
[0296] The server has the following main functions:
[0297] 1. News data collection and filtering:
[0298] The server retrieves the latest news data from various news APIs and filters it based on the genre selected by the user.
[0299] 2. Generate a reading transcript:
[0300] The server uses a generative AI model to automatically generate a reading transcript from the filtered news articles.
[0301] 3. Creating a virtual announcer:
[0302] The server retrieves the information of the virtual announcer selected by the user and combines it with the script to generate the final video material. Specifically, it generates the reading voice using a voice synthesizer and creates the character movements using 3D modeling.
[0303] 4. Emotion Recognition with Emotion Engine:
[0304] The server receives the user's camera footage and audio data sent from the device and analyzes the user's emotions in real time. The analysis results are classified into emotional states such as smile, surprise, sadness, etc.
[0305] 5. Live streaming and dynamic adjustment:
[0306] The server dynamically adjusts the virtual announcer's facial expressions and tone of voice during live broadcasts based on the analysis results of the emotion engine, and can also recommend relevant news content based on the user's emotional state.
[0307] 6. Collecting User Feedback:
[0308] After watching the news, we collect feedback from users to help improve the service. The feedback information is stored on the server and can be used for data analysis.
[0309] Specific examples
[0310] If the user selects sports news
[0311] 1. User:
[0312] Install the news distribution app, create an account, and log in. On the settings screen, select the "Sports" genre and then set your preferred virtual announcer (e.g., a male voice with an active personality).
[0313] 2. Terminal:
[0314] The selection information is stored on the local device and sent to the server, and facial expression and voice data of the user are collected in real time via the camera and microphone and sent to the server.
[0315] 3. Server:
[0316] Collect and filter sports-related news data. Use a generative AI model to automatically generate a script to read from the news data. Obtain information about the virtual announcer selected by the user and combine it with the script to generate the final video material.
[0317] 4. Emotion Engine:
[0318] At the start, the system analyzes the user's emotions and provides the results to the server, which then adaptively adjusts the virtual announcer's facial expressions and tone of voice based on the analysis results.
[0319] 5. Live Streaming:
[0320] The server uploads the generated video material to the live streaming server and provides a stream URL to the device, which then plays the video in real time.
[0321] 6. Users:
[0322] Users can open the app during their commute and watch a virtual announcer talk about the results of last night's game. While watching, the virtual announcer changes its facial expression and tone of voice depending on the user's emotions. It also provides related news based on the user's emotions.
[0323] 7. Feedback:
[0324] After watching the news, users provide feedback within the app, which is then sent to the server and used to improve the service.
[0325] This system allows users to enjoy a highly personalized, emotionally-driven news experience in real time.
[0326] The processing flow will be explained below.
[0327] ---
[0328] Step 1:
[0329] Users install the news delivery application on their device and create an account. After creating the account, they log in by entering their email address and password on the login page.
[0330] Step 2:
[0331] The terminal sends the user's login information to the server for authentication. If authentication is successful, the server issues a session ID and returns it to the terminal. The terminal saves the session ID and maintains the logged-in state.
[0332] Step 3:
[0333] Users access the app's settings screen, select their preferred news genre (e.g., sports, technology, politics, etc.), customize their preferred virtual announcer (e.g., gender, tone of voice, appearance, etc.), and click the "Save" button.
[0334] Step 4:
[0335] The terminal transmits the user's selected news genre and virtual announcer setting information to the server, which then stores the received customization information in the user database.
[0336] Step 5:
[0337] The server retrieves the latest news data from multiple news APIs, filters the retrieved news data based on the genre selected by the user, and extracts the corresponding news data.
[0338] Step 6:
[0339] The server uses a generative AI model to automatically generate a speech transcript from the filtered news data, which is then stored in a database.
[0340] Step 7:
[0341] The server retrieves the information of the virtual announcer selected by the user and combines it with the script to generate the final video material. Specifically, it generates the reading voice using a voice synthesizer and creates the character movements using 3D modeling.
[0342] Step 8:
[0343] The device uses a camera and microphone to collect the user's facial expressions and voice data in real time and transmits it to a server.
[0344] Step 9:
[0345] The server uses an emotion engine to analyze the user's facial expression and audio data sent from the device and recognize the user's emotional state. The recognition results are classified as emotions such as smile, surprise, sadness, etc.
[0346] Step 10:
[0347] Based on the analysis results of the emotion engine, the server dynamically adjusts the virtual announcer's facial expressions and tone of voice during live broadcasts, enabling a more natural and approachable news delivery that responds to the user's emotions.
[0348] Step 11:
[0349] The server uploads the generated news video of the virtual announcer to the live streaming server and sends the stream URL to the terminal.
[0350] Step 12:
[0351] The device uses the stream URL received from the server to play the video in real time, and the user opens the app to watch the virtual announcer change his or her facial expressions and tone of voice according to the user's emotions.
[0352] Step 13:
[0353] During live streaming, the server dynamically recommends relevant news content based on the user's emotional state, for example, if the user expresses surprise, it will display the latest news related to surprise.
[0354] Step 14:
[0355] After watching the news, the user inputs feedback information using a feedback form, which is then sent to the server by the terminal.
[0356] Step 15:
[0357] The server stores user feedback information in a database and uses it for data analysis to improve services and optimize generative AI models.
[0358] This concludes the specific processing flow of the system, which enables users to enjoy a personalized, emotion-sensitive news experience in real time.
[0359] Example 2
[0360] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0361] Current digital news delivery systems only provide one-way information to users, limiting personalization based on user emotions and preferences. Furthermore, passive news consumption makes it difficult to provide an experience tailored to users' interests and engagement levels. Therefore, there is a need for a news delivery system that can capture users' interest and adjust in real time to reflect their emotions.
[0362] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for selecting information of a genre that the user likes, means for the user to customize a virtual character, means for acquiring information data and filtering it based on the genre selected by the user, means for generating a text-to-speech script from the filtered information data, means for having the virtual character read the generated text-to-speech script, means for recognizing the user's emotions in real time using the camera and microphone of the terminal and transmitting the data, means for analyzing the emotions and dynamically adjusting the virtual character's facial expression and tone of voice based on the analysis results, and means for providing information to the user in a live streaming format. This enables highly personalized news delivery according to the user's emotions.
[0363] "User" refers to a general user who receives information and performs settings and customization.
[0364] A "virtual character" is a digital avatar that provides information to users and has audio and visual representations.
[0365] "Information data" refers to digital information such as text, audio, and video that is acquired based on the genre selected by the user.
[0366] "Filtering" refers to the process of selecting only relevant data from the acquired information data based on user selections.
[0367] A "reading manuscript" refers to text that is analyzed and generated by a generative AI model based on information data and is read aloud by a virtual character.
[0368] A "generative AI model" refers to an artificial intelligence model that generates and analyzes text based on large amounts of data.
[0369] "Emotion recognition" refers to a technology that uses the device's camera and microphone to analyze emotions from the user's facial expressions and voice.
[0370] "Dynamic adjustment" refers to the process of changing a virtual character's facial expression or tone of voice based on the results of a user's emotional analysis in real time.
[0371] "Live streaming" refers to a method of providing generated video material to users in real time.
[0372] "Feedback" refers to opinions and impressions provided by users after viewing.
[0373] "Customization" refers to the process by which a user sets the gender, tone of voice, character style, etc. of a virtual character according to their own preferences.
[0374] This invention is a system that allows users to select information from their favorite genres and delivers personalized information via a customized virtual character. This system achieves even higher levels of personalization by recognizing the user's emotions in real time and dynamically adjusting the virtual character's facial expressions and tone of voice based on that information.
[0375] System configuration
[0376] User
[0377] 1. The user installs a news delivery application on their device (smartphone, tablet, PC, etc.), creates an account, and logs in.
[0378] 2. After successfully logging in, you can select and customize your news genre (e.g., "Sports," "Politics," "Technology," etc.) and virtual character (gender, voice tone, character style, etc.) on the settings screen.
[0379] Terminal
[0380] 1. The device is equipped with a camera and microphone, which are used to recognize the user's emotions (facial expressions and tone of voice) in real time.
[0381] 2. The user's selected settings information (news genre, virtual character customization) is saved in a local database and sent to the server.
[0382] server
[0383] 1. News data collection and filtering: The server retrieves the latest news data from various news APIs (e.g., NewsAPI, Google News API) and filters it based on the genre selected by the user.
[0384] 2. Generate a prompt: The server uses a generative AI model (e.g., OpenAI GPT-3) to automatically generate a prompt from the filtered news articles.
[0385] Sample prompt: "Create a script based on the latest sports news article to be read aloud by a fictional character who is an active and energetic man."
[0386] 3. Virtual character generation: The server acquires the virtual character information (e.g., character model, voice tone) set by the user and combines it with the reading script to generate the final video material. Specifically, the reading voice is generated using a voice synthesizer (e.g., Amazon Polly), and the virtual character's movements are created using 3D modeling software (e.g., Blender).
[0387] 4. Emotion recognition using an emotion engine: The server receives the user's camera footage and audio data sent from the device and analyzes the user's emotions using an emotion recognition engine (e.g., Affectiva SDK). Based on the analysis results (e.g., smile, surprise, sadness), the virtual character's facial expressions and tone of voice are dynamically adjusted.
[0388] 5. Live streaming: The final video material is uploaded to the live streaming server, and a stream URL is generated and sent to the device.
[0389] Specific examples
[0390] 1. User: Install the news distribution app, create an account, and log in. Select the "Sports" genre on the settings screen and set your preferred virtual character (e.g., a male voice with an active personality).
[0391] 2. Device: The device stores the selection information locally and sends it to the server. It also collects the user's facial expressions and voice data in real time via the camera and microphone and sends them to the server.
[0392] 3. Server: Collects and filters sports-related news data. Automatically generates a script to read from the news data using a generative AI model. Information about the virtual character set by the user is acquired, and combined with the script to generate the final video material.
[0393] 4. Emotion Engine: At the start, the system analyzes the user's emotions and provides the results to the server, which then adaptively adjusts the virtual character's facial expressions and tone of voice based on the analysis results.
[0394] 5. Live streaming: The server uploads the generated video material to the live streaming server and provides a stream URL to the device, which then plays the video in real time.
[0395] 6. User: Watch the news and hear a virtual character say, "I'll tell you the results of last night's game." While watching, the virtual character changes its facial expression and tone of voice according to the user's emotions. Related news is also provided.
[0396] 7. Feedback: After watching the news, users can provide feedback within the app, which will be sent to the server and used to improve the service.
[0397] This system allows users to enjoy a highly personalized, emotionally-driven news experience in real time.
[0398] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0399] Step 1: Initial System Setup and Login
[0400] operation:
[0401] Users install the news distribution application on their device, create an account, and then log in.
[0402] input:
[0403] Your account information (email address, password, social media accounts, etc.).
[0404] output:
[0405] Session information indicating that the user logged in successfully.
[0406] Specific behavior:
[0407] A user launches the app and enters their account information on the login screen. The entered information is sent to the server, where authentication is performed. If successful, session information is returned to the device.
[0408] Step 2: Setting the news genre and virtual characters
[0409] operation:
[0410] Users select a news genre on the settings screen and customize their preferred virtual character.
[0411] input:
[0412] Customization information for the news genre and virtual character selected by the user.
[0413] output:
[0414] The setting information is stored on the device and sent to the server.
[0415] Specific behavior:
[0416] Users select a news genre (e.g., sports, politics, technology, etc.) on the app's settings screen and customize the attributes of their virtual character (e.g., gender, tone of voice, appearance, etc.). The selection information is stored in a local database and transmitted to a server over the network.
[0417] Step 3: Collecting and filtering news data
[0418] operation:
[0419] The server retrieves the latest news data from the news API and filters it based on the genre selected by the user.
[0420] input:
[0421] News data obtained from the news API, news genres selected by the user.
[0422] output:
[0423] News data filtered to user-selected genres.
[0424] Specific behavior:
[0425] The server sends a request to a news API (e.g., NewsAPI, Google News API) to retrieve the latest news data, which is then filtered by keywords and categories based on the user's selection to select only relevant news.
[0426] Step 4: Generate a transcript
[0427] operation:
[0428] The server uses a generative AI model to automatically generate a reading script from filtered news data.
[0429] input:
[0430] Filtered news data, generative AI models (e.g., OpenAI GPT-3).
[0431] output:
[0432] The generated reading transcript.
[0433] Specific behavior:
[0434] The server inputs the filtered news data into a generative AI model and generates a reading script using a prompt, such as "Please create a reading script for a lively and energetic male virtual character based on the latest sports news article." The generated script is then stored in a database.
[0435] Step 5: Generate a virtual character
[0436] operation:
[0437] The server combines the information from the virtual character with the reading script to generate the final video material.
[0438] input:
[0439] Virtual character customization information, generated reading script.
[0440] output:
[0441] A video clip of a virtual character reading the news.
[0442] Specific behavior:
[0443] The server retrieves the virtual character's customization information, generates the reading voice using a voice synthesizer (e.g., Amazon Polly), and generates the virtual character's movements and facial expressions using 3D modeling software (e.g., Blender). These are then combined to generate the final video material.
[0444] Step 6: Emotion recognition and dynamic regulation
[0445] operation:
[0446] The device's camera and microphone are used to collect the user's emotions in real time and send the data to a server.
[0447] input:
[0448] User's facial expression data, voice data.
[0449] output:
[0450] User sentiment analysis results.
[0451] Specific behavior:
[0452] The device uses a camera and microphone to collect the user's facial expressions and tone of voice in real time, and this data is sent to a server. An emotion recognition engine (e.g., Affectiva SDK) on the server analyzes this data and classifies emotional states such as smile, sadness, or surprise. Based on the analysis results, the virtual character's facial expressions and tone of voice are dynamically adjusted.
[0453] Step 7: Live Stream
[0454] operation:
[0455] The server uploads the generated video material to a live streaming server, generates a stream URL, and sends it to the terminal.
[0456] input:
[0457] The final video footage produced.
[0458] output:
[0459] The stream URL for live streaming.
[0460] Specific behavior:
[0461] The server uploads the video material to the live streaming server and generates a stream URL. This URL is sent to the device and can be accessed by the user. The device uses this URL to play the video in real time and displays a virtual character reading the news.
[0462] Step 8: Gather feedback
[0463] operation:
[0464] Users provide feedback after watching the news.
[0465] input:
[0466] User feedback information.
[0467] output:
[0468] Feedback information collected.
[0469] Specific behavior:
[0470] Users enter their opinions and thoughts into the feedback form within the app. This feedback information is collected by the device and sent to the server, which stores it in a database and uses it to improve the service in the future.
[0471] (Application example 2)
[0472] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0473] Conventional news delivery systems have limited personalization based on user emotions and provide a uniform news viewing experience, making it difficult to increase user satisfaction. Furthermore, they lack interactivity in news delivery because they are unable to reflect users' real-time reactions.
[0474] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for selecting a news genre that the user likes, means for the user to customize the virtual announcer, means for acquiring news data and filtering it based on the genre selected by the user, means for generating a script to be read from the filtered news data, means for having the virtual announcer read the generated script to be read, means for providing news to the user in a live streaming format, means for analyzing the user's emotions, and means for dynamically adjusting the virtual announcer's facial expressions and tone of voice based on the analyzed emotions. This enables a highly personalized news experience that corresponds to the user's emotions.
[0475] "Means for users to select news genres they like" refers to interfaces and functions that allow users to select news genres that interest them.
[0476] "Means for users to customize their virtual announcers" refers to interfaces and functions that allow users to set the announcer's appearance, voice, personality, etc.
[0477] "Means for acquiring news data and filtering it based on the genre selected by the user" is a function for collecting data from news sources and extracting news according to the genre selected by the user.
[0478] The "means for generating a script to be read aloud from filtered news data" is a function that automatically creates a script to be read aloud by a virtual announcer based on extracted news data.
[0479] The "means for having a virtual announcer read out the generated script" is a function for playing back the generated script in the voice of a virtual announcer.
[0480] The "means for providing news to users in a live distribution format" is a function for distributing news content to users in real time.
[0481] "Means for analyzing user emotions" refers to technology that analyzes a user's facial expressions and tone of voice in real time to recognize their emotional state.
[0482] "Means for dynamically adjusting the facial expressions and tone of voice of a virtual announcer based on analyzed emotions" refers to technology that changes the facial expressions and voice of a virtual announcer in real time based on the results of emotion analysis.
[0483] This invention is a system that achieves even higher levels of personalization by combining a news delivery system with a user-selected news genre and a customized virtual announcer with an emotion engine that recognizes the user's emotions. This system recognizes the user's emotions in real time using a smartphone, and based on that information, adjusts the virtual announcer's facial expressions and tone of voice, and recommends related news content.
[0484] System configuration
[0485] User terminal
[0486] Users install the "Emotion-Responsive News" application on their smartphones, create an account, and then log in. After logging in, they can select and customize the news genre and virtual announcer on the settings screen. The smartphones are also equipped with cameras and microphones, which are used to recognize the user's emotions in real time.
[0487] server
[0488] The server has the following main functions:
[0489] 1. News data collection and filtering
[0490] The server retrieves the latest news data from the news API and filters it based on the user's selected genre.
[0491] 2. Generation of reading script
[0492] The server uses a generative AI model to automatically generate a reading script from filtered news articles.
[0493] 3. Creating a Virtual Announcer
[0494] The server retrieves information about the virtual announcer set by the user, combines it with the script to be read, and generates the final video material using a voice synthesizer and 3D modeling.
[0495] 4. Emotion Recognition by Emotion Engine
[0496] The server receives the user's camera footage and audio data sent from the smartphone and analyzes the user's emotions in real time. The analysis results are classified into emotional states such as smile, surprise, sadness, etc.
[0497] 5. Live streaming and dynamic adjustments
[0498] The server dynamically adjusts the virtual announcer's facial expressions and tone of voice during live broadcasts based on the analysis results of the emotion engine, and can also recommend relevant news content based on the user's emotional state.
[0499] 6. Collecting User Feedback
[0500] After watching the news, feedback is collected from users to help improve the service. The feedback information is stored on the server and used for data analysis.
[0501] Specific examples
[0502] If the user selects sports news
[0503] 1. Users
[0504] Users install the "Emotionally Responsive News" app on their smartphone, create an account, and log in. On the settings screen, they select the "Sports" genre and then set their preferred virtual announcer (e.g., a man with a lively personality).
[0505] 2. Terminal
[0506] The selection information is stored on the local device and sent to the server. In addition, facial expression and voice data of the user are collected in real time via the smartphone's camera and microphone and sent to the server.
[0507] 3. Server
[0508] The server collects sports-related news data and filters it through a news API. It uses a generative AI model to automatically generate a script to read from the news data. It then obtains information about the virtual announcer selected by the user, combines it with the script, and generates the final video material using a voice synthesizer and 3D modeling.
[0509] 4. Emotion Engine
[0510] The server analyzes the video and audio transmitted from the user's smartphone to identify the user's emotional state in real time, and adaptively adjusts the virtual announcer's facial expressions and tone of voice based on the analysis results.
[0511] 5. Live Streaming
[0512] The server uploads the generated video material to a live streaming server and provides a stream URL to the device. The smartphone uses this URL to play the video in real time, and a virtual announcer announces, "I'll tell you the results of last night's game."
[0513] 6. Users
[0514] Users can open the app during their commute and watch a virtual announcer speak. While watching, the virtual announcer changes its facial expression and tone of voice depending on the user's emotions. It also provides relevant news based on the user's emotions.
[0515] Prompt Sentence Examples
[0516] Leveraging a generative AI model, we use prompts like:
[0517] Please create a news script to use in your presentation. It should be based on recent news in the sports genre and should include the following:
[0518] Last night's match results
[0519] Performances of noteworthy players
[0520] Upcoming match schedule
[0521] Please keep it easy to read and in a friendly tone."
[0522] This allows users to enjoy a highly personalized news experience that responds to their emotions in real time.
[0523] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0524] Step 1:
[0525] A user installs the "Emotion Response News" app on their smartphone, creates an account, and logs in. As input, the user provides the necessary registration information, and as output, an account is created. Specifically, information such as a username, password, and email address is entered, and the user's account data is generated based on this.
[0526] Step 2:
[0527] Users select and customize their news genre and virtual announcer on the app's settings screen. The input is the user's genre selection and announcer customization information, and the output is saved to the local device. Specifically, this includes genre selections such as "sports" and "entertainment," as well as the announcer's appearance, voice tone, and personality settings.
[0528] Step 3:
[0529] The terminal sends the setting information of the news genre and virtual announcer to the server. The user's selection information is provided as input, and the server receives this information as output and stores it in a database. Specifically, the setting information is sent to the server as JSON format data.
[0530] Step 4:
[0531] The server retrieves the latest news data from the news API and filters the data based on the genre selected by the user. As input, news data from the news API is provided, and as output, news data narrowed down to a genre is generated. Specifically, the server retrieves news data using an API key and filters the news according to the specified genre.
[0532] Step 5:
[0533] The server uses a generative AI model to automatically generate a reading script from the filtered news articles. The filtered news data is provided as input, and text data for reading is generated as output. Specifically, a prompt sentence is input to the generative AI model (such as OpenAI GPT) and the generated text is obtained.
[0534] Step 6:
[0535] The server combines the generated script with the virtual announcer information set by the user and generates the final video material using a voice synthesizer and 3D modeling. The script and announcer settings are provided as input, and video material is generated as output. Specifically, the voice synthesizer generates audio data and links it to the 3D model to create a video.
[0536] Step 7:
[0537] The device uses the smartphone's camera and microphone to collect the user's facial expressions and voice data in real time and transmits it to a server. The user's real-time video and audio are provided as input, and emotion analysis data is generated as output. Specifically, emotion analysis is performed using the Microsoft Azure Face API and Google Cloud Vision API.
[0538] Step 8:
[0539] The server analyzes the user's emotional data provided in real time and dynamically adjusts the virtual announcer's facial expressions and tone of voice. The emotion analysis results are provided as input, and the announcer's facial expressions and tone of voice are adjusted as output. Specifically, the synthesizer parameters are adjusted based on the emotional data, changing the facial expressions of the 3D model.
[0540] Step 9:
[0541] The server uploads the generated video material to the distribution server in real time and provides a stream URL to the terminal. The completed video material is provided as input, and a stream URL is generated as output. Specifically, the video upload process to the distribution server is executed.
[0542] Step 10:
[0543] The device plays the video in real time based on the stream URL and provides news to the user. The stream URL is provided as input and the video is played as output. Specifically, the video player receives the stream and plays it.
[0544] Step 11:
[0545] After watching the news, the user provides feedback, which the device sends to the server. The user's feedback is provided as input, and the server receives and stores the feedback data as output. Specifically, the user taps the feedback icon on the screen and enters their opinion or impression.
[0546] Step 12:
[0547] The server uses the collected feedback data for data analysis to help improve the service. User feedback data is provided as input, and improvement measures are formulated as output. Specifically, the feedback data is analyzed using a data analysis tool, and areas for improvement are listed.
[0548] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0549] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0550] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0551] [Second embodiment]
[0552] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0553] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0554] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0555] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0556] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0557] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0558] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0559] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0560] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0561] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0562] In the smart glasses 214, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0563] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0564] ---
[0565] This invention is a system for providing viewers with a personalized news experience, allowing users to select a news genre and a virtual announcer and watch customized live news. This system has functions that allow users to select news genres and virtual announcers according to their preferences, and functions that automatically generate a text-to-speech script from news data and have the virtual announcer read it aloud. It also has a function for improving the service based on user feedback.
[0566] System configuration
[0567] User terminal
[0568] Users install the news distribution application on their device, create an account, and then log in. Once successfully logged in, users can select news genres and virtual announcers. The user's selection information is saved on the device and sent to the server.
[0569] server
[0570] The server has the following main functions:
[0571] 1. News data collection:
[0572] The server retrieves the latest news data from various news APIs. This data is aggregated from multiple news sources.
[0573] 2. Filtering and generating a transcript:
[0574] The collected news data is filtered based on the news genre selected by the user, and then a generative AI model is used to automatically generate a reading transcript from the filtered news articles.
[0575] 3. Creating a virtual announcer:
[0576] The virtual announcer reads the script based on the user-customized virtual announcer settings (voice, appearance, etc.), and then generates the final video material for live streaming.
[0577] 4. Live Streaming:
[0578] The generated virtual announcer will broadcast live news videos in real time over the Internet.
[0579] 5. Collecting User Feedback:
[0580] After watching the news, we collect feedback from users to help improve the service. The feedback information is stored on the server and can be used for data analysis.
[0581] Specific examples
[0582] If the user selects sports news
[0583] 1. User:
[0584] Install the news distribution app and create an account. After logging in, select the "Sports" genre on the settings screen and then set your preferred virtual announcer (e.g., a male voice with an active personality).
[0585] 2. Terminal:
[0586] The selection information is stored on the local device and transmitted to the server.
[0587] 3. Server:
[0588] Sports-related news data is collected and filtered. Then, using a generative AI model, a script is automatically generated from the filtered news data. The final video material is generated based on the script and the settings of the virtual announcer selected by the user.
[0589] 4. Terminal:
[0590] The user's device receives the streaming URL from the server and plays the sports news in real time.
[0591] 5. User:
[0592] On the way to work, users can open the app and watch a virtual announcer announce the results of last night's game. After watching, users can provide feedback within the app.
[0593] This system allows for a personalized news experience, increasing viewer engagement.
[0594] The processing flow will be explained below.
[0595] ---
[0596] Step 1:
[0597] Users install the news delivery application on their device and create an account. After creating the account, they log in by entering their email address and password on the login page.
[0598] Step 2:
[0599] The terminal sends the user's login information to the server for authentication. If authentication is successful, the server issues a session ID and returns it to the terminal. The terminal saves the session ID and maintains the logged-in state.
[0600] Step 3:
[0601] Users access the app's settings screen, select their preferred news genre (e.g., sports, technology, politics, etc.), customize their preferred virtual announcer (e.g., gender, tone of voice, appearance, etc.), and click the "Save" button.
[0602] Step 4:
[0603] The terminal transmits the user's selected news genre and virtual announcer setting information to the server, which then stores the received customization information in the user database.
[0604] Step 5:
[0605] The server retrieves the latest news data from multiple news APIs, filters the retrieved news data based on the genre selected by the user, and extracts the corresponding news data.
[0606] Step 6:
[0607] The server uses a generative AI model to automatically generate a speech transcript from the filtered news data, which is then stored in a database.
[0608] Step 7:
[0609] The server retrieves the information of the virtual announcer selected by the user and combines it with the script to generate the final video material. Specifically, it generates the reading voice using a voice synthesizer and creates the character movements using 3D modeling.
[0610] Step 8:
[0611] The server uploads the generated news video of the virtual announcer to the live streaming server and generates a stream URL, which is then sent to the device.
[0612] Step 9:
[0613] The device plays the video in real time based on the stream URL received from the server, and users can open the app at their preferred time to watch customized live news.
[0614] Step 10:
[0615] After watching the news, users can rate the announcer's performance and the content of the news in the feedback form within the app and enter the feedback information on their device, which then sends the feedback information to the server.
[0616] Step 11:
[0617] The server stores user feedback information in a database and uses it for data analysis to help improve generative AI models and services.
[0618] This completes the system's processing flow, allowing users to enjoy a personalized news experience in real time.
[0619] Example 1
[0620] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0621] Conventional news delivery systems make it difficult for viewers to select news that suits their preferences and watch it in a customized format. Furthermore, they lacked a mechanism for effectively collecting viewer feedback and utilizing it to improve services. As a result, they were unable to fully optimize the viewing experience or improve viewer satisfaction.
[0622] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0623] In this invention, the server includes a means for acquiring news data and filtering it based on a genre selected by the user, a means for generating a script from the filtered news data using a generative AI model, and a means for having a virtual announcer read the generated script. This allows users to select news from their favorite genres and watch live news broadcasts by customized virtual announcers. Furthermore, by collecting user feedback and using it to improve the service, viewer satisfaction can be increased.
[0624] 1. "A means for users to select news genres they like" refers to a function that allows users to select news categories that interest them.
[0625] 2. "Means for users to customize the virtual announcer" refers to the functionality that allows users to change settings such as the voice and appearance of the virtual announcer.
[0626] 3. "Means of acquiring news data" refers to the function of collecting the latest news information from external news data providers.
[0627] 4. "Means of filtering based on the genre selected by the user" refers to the function of extracting only information that falls into the category selected by the user from the collected news data.
[0628] 5. "Means for generating a text-to-speech transcript from filtered news data" refers to a function that uses a generative AI model to summarize filtered news information and create text that can be read aloud.
[0629] 6. "Means for having a virtual announcer read the generated script" refers to the function of converting text data into voice using speech synthesis technology and synchronizing it with the actions of the virtual announcer.
[0630] 7. "Means for providing users with generated news videos in a live streaming format" refers to the function of delivering generated audio and video to users in real time via the Internet.
[0631] 8. "Means of collecting user feedback" refers to the function of collecting ratings and comments from viewers and storing them in a database.
[0632] 9. "Means used to improve services" refers to the functions used to analyze collected feedback and improve the quality of news provision services.
[0633] 10. "Generative AI model" refers to an artificial intelligence model that generates text data using natural language processing technology.
[0634] 11. “Prompt” refers to a command or question input to a generative AI model.
[0635] These are the definitions of the important words included in this system.
[0636] The present invention is a system for providing users with a personalized news experience, allowing them to select a news genre and a virtual announcer and watch customized live news. This system has functions that allow users to select news genres and virtual announcers according to their preferences, and functions that automatically generate a text-to-speech script from news data and have the virtual announcer read it aloud. It also has a function to improve the service based on user feedback.
[0637] System configuration
[0638] User terminal
[0639] Users install the news distribution application on their device, create an account, and then log in. Once successfully logged in, users can select news genres and virtual announcers. The user's selection information is saved on the device and sent to the server. This can be done using common mobile devices such as smartphones and tablets.
[0640] server
[0641] The server has the following main functions:
[0642] 1. News data collection:
[0643] The server retrieves the latest news data from various news providers. This data is collected from multiple news sources. Specifically, "News API" and "News Data Provider Services" are used.
[0644] 2. Filtering and generating a transcript:
[0645] The collected news data is filtered based on the news genre selected by the user. Then, a generative AI model is used to automatically generate a speech transcript from the filtered news articles. The generative AI model used here is, for example, GPT-4.
[0646] 3. Creating a virtual announcer:
[0647] The virtual announcer reads the script based on the user's customized settings (voice, appearance, etc.). The final video material for live streaming is then generated. Speech synthesis software (e.g., TTS technology) is used for the speech synthesis, and 3D animation software (e.g., animation generation software) is used to generate the video.
[0648] 4. Live Streaming:
[0649] The generated news video of the virtual announcer is broadcast live. The live broadcast is carried out in real time over the Internet. A "live streaming service," for example, is used as the streaming platform.
[0650] 5. Collecting User Feedback:
[0651] After watching the news, we collect feedback from users to help improve the service. The feedback information is stored on the server and can be used for data analysis.
[0652] Specific examples
[0653] If the user selects sports news
[0654] 1. User:
[0655] Install the news distribution app and create an account. After logging in, select the "Sports" genre on the settings screen and then set your preferred virtual announcer (e.g., a male voice with an active personality).
[0656] 2. Terminal:
[0657] The selection information is stored on the local device and transmitted to the server.
[0658] 3. Server:
[0659] Sports-related news data is collected and filtered. Then, using a generative AI model, a script is automatically generated from the filtered news data. The final video material is generated based on the script and the settings of the virtual announcer selected by the user.
[0660] 4. Terminal:
[0661] The user's device receives the streaming URL from the server and plays the sports news in real time.
[0662] 5. User:
[0663] On the way to work, users can open the app and watch a virtual announcer announce the results of last night's game. After watching, users can provide feedback within the app.
[0664] Prompt Sentence Examples
[0665] "I want to watch sports news. I want the virtual announcer to be a lively male."
[0666] This system will enable viewers to enjoy a news experience that is optimized for each individual, increasing engagement.
[0667] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0668] Step 1:
[0669] A user installs a news delivery app and creates an account
[0670] Input: User's electronic device (smartphone, tablet) and internet connection
[0671] How it works: A user downloads and installs a news delivery app from the app store. They launch the app and create an account by entering their email address and password on the account creation screen that appears when the app is first launched.
[0672] Output: Account information is created and you can log in to the app. After logging in, you will be taken to a screen where you can select the news genre and virtual announcer.
[0673] Step 2:
[0674] The user selects the news genre and virtual announcer.
[0675] Input: User preference information (selected news genre, virtual announcer settings)
[0676] How it works: Within the app, users select "Sports" from a range of news genres, including "Sports," "Politics," and "Entertainment." They then customize the virtual announcer's settings screen, such as selecting a "lively male voice."
[0677] Output: The selection information is temporarily stored in the device's local storage and then sent to the server.
[0678] Step 3:
[0679] The server collects news data
[0680] Input: API requests and responses received by the server from external news providers
[0681] How it works: The server sends a request to an API endpoint such as a "news data provider service" to retrieve the latest news data in JSON format.
[0682] Output: The retrieved news data is stored in a database for further processing.
[0683] Step 4:
[0684] The server filters the news data and generates a reading script.
[0685] Input: News data to be filtered and user genre selection information
[0686] How it works: The server uses a filtering algorithm (e.g., TF-IDF) to extract only news data that falls within the user's selected "sports" genre. It then uses a generative AI model (e.g., GPT-4) to generate a summary of the filtered news data.
[0687] Output: The filtered news data and the generated speech transcripts are stored in a database.
[0688] Step 5:
[0689] The server configures the virtual announcer and generates the news video.
[0690] Input: Reading script and user's virtual announcer setting information
[0691] How it works: The server uses speech synthesis software (e.g., TTS technology) to convert the generated script into audio data, and then uses 3D animation software (e.g., animation generation software) to generate a video that coordinates the voice and movements of the virtual announcer.
[0692] Output: The created news video is stored on the server and ready to be streamed.
[0693] Step 6:
[0694] Server-generated live news video broadcast
[0695] Input: Generated news video data
[0696] How it works: The server uses a live streaming service to create a streaming URL for the generated news video, and then sends this URL to the user's device.
[0697] Output: A response containing a streaming URL is sent to the user's device.
[0698] Step 7:
[0699] User watches news
[0700] Input: Streaming URL
[0701] How it works: The user's device receives the streaming URL and uses the media player component to play the news video in real time.
[0702] Output: User watches live streaming news video.
[0703] Step 8:
[0704] Users provide feedback
[0705] Input: User feedback information (ratings, comments, etc.)
[0706] How it works: After watching a news item, users enter their rating and comments about their viewing experience on the feedback screen displayed within the app. The feedback is then sent from the app to the server.
[0707] Output: The feedback information is sent to the server and stored in a database.
[0708] The above is a detailed description of the specific processing steps of the present system. The present invention is expected to provide users with an individually customized news experience, thereby improving viewer engagement and satisfaction.
[0709] (Application example 1)
[0710] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0711] In today's world, with so much news being distributed, it is difficult for individual users to accurately obtain only the news that interests them. There are also insufficient means to customize the visual and auditory experience of news for each individual user. Furthermore, there is a lack of mechanisms for users to provide feedback on the news they have viewed and use that feedback to improve services. Effective methods to solve these issues are needed.
[0712] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0713] In this invention, the server includes means for acquiring news data and filtering it based on a genre selected by the user, means for generating a script to be read from the filtered news data, and means for having a virtual announcer read the generated script to be read. This allows each user to watch news in the genres they are interested in in real time with a customized virtual announcer. Furthermore, it is possible to stream news to a smartphone application in real time and collect feedback from users to help improve the service.
[0714] "User" refers to an individual who uses this system to receive news distribution services.
[0715] "Genre" refers to the type or category of news, such as a specific field such as sports, economics, or entertainment.
[0716] A "virtual announcer" is a computer-generated character whose role is to read the news aloud and visually.
[0717] "News Data" refers to the collection of information collected from various news sources and refers to the material used for filtering and generating the readings.
[0718] "Filtering" refers to the process of narrowing down news data based on genres selected by the user.
[0719] A "script" is a text automatically generated from filtered news data, and refers to a script that a virtual announcer will read aloud.
[0720] "Live streaming" refers to a format in which news is provided to users in real time, and refers to a method of streaming directly from a server to a user's terminal.
[0721] "Smartphone application" refers to a program that runs on a smartphone and functions as the user interface for this system.
[0722] "Real-time streaming" refers to a technology in which news videos are sent to the user's device as soon as they are generated and played back instantly.
[0723] "Feedback" refers to the opinions and evaluations provided by users regarding the news and services they have viewed, and serves as important data for improving services.
[0724] A "generative AI model" refers to an artificial intelligence algorithm that uses machine learning technology to automatically generate a reading script from news data.
[0725] The present invention is a system that allows users to customize the news genres and virtual announcers that interest them, providing a personalized news experience.
[0726] System configuration
[0727] User terminal
[0728] Users install the corresponding application on their smartphone, create an account, and log in. After logging in, users can configure their news genre and virtual announcer settings. This includes selecting a genre (e.g., sports, economics, entertainment, etc.) and customizing the announcer's appearance and voice. This configuration information is saved on the user's device and sent to the server.
[0729] server
[0730] The server has the following main functions:
[0731] 1. News data collection:
[0732] The server uses news APIs to collect the latest news data from various news sources, examples of which include NewsAPI and Google News API.
[0733] 2. Filtering and generating a transcript:
[0734] The collected news data is filtered based on the genre selected by the user, and a transcription is automatically generated from the filtered news articles using a generative AI model (e.g., GPT-3).
[0735] 3. Creating a virtual announcer:
[0736] A virtual announcer reads the script according to the announcer settings customized by the user, using tools such as Adobe Character Animator and Unity3D. The generated news video is then streamed to the user's device in real time.
[0737] 4. Gathering Feedback:
[0738] After viewing the news, we collect feedback from users and use it to improve our services, thereby continuously improving the quality of news and its delivery methods.
[0739] Process Overview
[0740] 1. User Settings:
[0741] For example, a user selects "Entertainment News" and "An announcer with a female voice and a cheerful personality." The user's selection information is sent from the smartphone terminal to the server.
[0742] 2. News gathering and filtering:
[0743] The server uses NewsAPI to collect the latest entertainment news and filters the collected data.
[0744] 3. Generate a reading transcript:
[0745] A generative AI model (GPT-3) is used to generate a read-aloud transcript from filtered news articles. For example, a sentence such as "A blockbuster film was announced at the film festival yesterday" is generated.
[0746] 4. Creating a Virtual Announcer:
[0747] Using Adobe Character Animator and Unity3D, a user-customized virtual announcer is generated, and a video is created in which the announcer reads the script.
[0748] 5. Real-time Streaming and Viewing:
[0749] The generated news videos are hosted on a server and streamed in real time to smartphones, where users open the app and listen to the news being read by an announcer.
[0750] 6. Collecting Feedback and Improving Our Services:
[0751] After watching, users provide feedback that helps improve the service.
[0752] Examples and prompts
[0753] Examples:
[0754] A user opens the app and selects a sports news item and a calm, male-voiced announcer. The server collects the latest sports news and uses a generative AI model (GPT-3) to create a script, such as "I'll tell you the results of last night's game." Using Unity3D, a virtual announcer reads the script, and the generated video is delivered to a smartphone. The user can watch the news on their way to work and provide feedback.
[0755] Prompt for the generative AI model:
[0756] News Cache: {Sports}
[0757] Announcer Voice: {Male voice, calm}
[0758] The statement read: "Here are the results of last night's match."
[0759] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0760] Step 1:
[0761] A user creates an account and logs into the smartphone application. The user selects a news genre (e.g., sports, economics, entertainment) and virtual announcer settings (e.g., voice type, gender, appearance). The selection information is stored on the user's device and sent to the server.
[0762] Input: User's genre selection information and virtual announcer settings
[0763] Output: Selections sent to the server
[0764] Specific behavior:
[0765] The user opens the app and navigates to the settings screen.
[0766] Use the drop-down menus and sliders to set your preferred news genres and anchor appearance and voice.
[0767] After the setting information is confirmed, pressing the "Save" button saves the selected information in the local storage of the device and simultaneously sends it to the server.
[0768] Step 2:
[0769] The server uses a news API to collect the latest news data for the genre selected by the user. For example, it uses NewsAPI or Google News API.
[0770] Input: News API request
[0771] Output: Collected news data
[0772] Specific behavior:
[0773] The server accesses the news API periodically or upon user request to collect the latest news data for the selected genre.
[0774] The collected news data is stored in a database on the server.
[0775] Step 3:
[0776] The collected news data is filtered based on the genre selected by the user.
[0777] Input: Collected news data, user genre preferences
[0778] Output: Filtered news data
[0779] Specific behavior:
[0780] An algorithm on the server analyzes the collected news data and extracts only articles related to the genre selected by the user.
[0781] The filtered results are stored in a database.
[0782] Step 4:
[0783] A reading script is automatically generated from filtered news data using a generative AI model (e.g., GPT-3).
[0784] Input: Filtered news data
[0785] Output: Speech manuscript
[0786] Specific behavior:
[0787] The server provides a prompt to the generative AI model, which generates a reading script based on the filtered news data.
[0788] For example, the prompt text is:
[0789] News Cache: {Sports}
[0790] Announcer Voice: {Male voice, calm}
[0791] The statement read: "Here are the results of last night's match."
[0792] The generated reading manuscript is stored in a database on the server.
[0793] Step 5:
[0794] Based on the user's customized settings, a video is generated for the virtual announcer to read the script.
[0795] Input: Reading script, user's virtual announcer settings
[0796] Output: News video with a virtual announcer
[0797] Specific behavior:
[0798] Using Adobe Character Animator and Unity3D, a virtual announcer is generated based on the announcer settings selected by the user.
[0799] The script is read aloud by an announcer, and then generated in video format.
[0800] The generated video is stored on the server.
[0801] Step 6:
[0802] The generated news video of the virtual announcer is streamed in real time and provided to users.
[0803] Input: News video with a virtual announcer
[0804] Output: Real-time streaming video played on user devices
[0805] Specific behavior:
[0806] The server hosts the generated news video and generates a streaming URL.
[0807] The user's device accesses this URL and watches the news video in real time.
[0808] Step 7:
[0809] After watching the news, we collect feedback from users to help improve our services.
[0810] Input: User feedback
[0811] Output: Data for service improvement
[0812] Specific behavior:
[0813] After users watch a news video, they can provide their ratings and opinions through an in-app feedback form.
[0814] The server analyzes the collected feedback data and uses it to improve the service.
[0815] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0816] ---
[0817] This invention is a system that achieves even higher levels of personalization by combining a personalized news delivery system with a user-selected news genre and customized virtual announcer with an emotion engine that recognizes the user's emotions. This system recognizes the user's emotions in real time and, based on that information, adjusts the virtual announcer's facial expressions and tone of voice, and recommends related news content.
[0818] System configuration
[0819] User terminal
[0820] Users install the news distribution application on their device, create an account, and then log in. Once successfully logged in, they can select and customize the news genre and virtual announcer on the settings screen. The device is also equipped with a camera and microphone, which are used to recognize the user's emotions in real time.
[0821] server
[0822] The server has the following main functions:
[0823] 1. News data collection and filtering:
[0824] The server retrieves the latest news data from various news APIs and filters it based on the genre selected by the user.
[0825] 2. Generate a reading transcript:
[0826] The server uses a generative AI model to automatically generate a reading transcript from the filtered news articles.
[0827] 3. Creating a virtual announcer:
[0828] The server retrieves the information of the virtual announcer selected by the user and combines it with the script to generate the final video material. Specifically, it generates the reading voice using a voice synthesizer and creates the character movements using 3D modeling.
[0829] 4. Emotion Recognition with Emotion Engine:
[0830] The server receives the user's camera footage and audio data sent from the device and analyzes the user's emotions in real time. The analysis results are classified into emotional states such as smile, surprise, sadness, etc.
[0831] 5. Live streaming and dynamic adjustment:
[0832] The server dynamically adjusts the virtual announcer's facial expressions and tone of voice during live broadcasts based on the analysis results of the emotion engine, and can also recommend relevant news content based on the user's emotional state.
[0833] 6. Collecting User Feedback:
[0834] After watching the news, we collect feedback from users to help improve the service. The feedback information is stored on the server and can be used for data analysis.
[0835] Specific examples
[0836] If the user selects sports news
[0837] 1. User:
[0838] Install the news distribution app, create an account, and log in. On the settings screen, select the "Sports" genre and then set your preferred virtual announcer (e.g., a male voice with an active personality).
[0839] 2. Terminal:
[0840] The selection information is stored on the local device and sent to the server, and facial expression and voice data of the user are collected in real time via the camera and microphone and sent to the server.
[0841] 3. Server:
[0842] Collect and filter sports-related news data. Use a generative AI model to automatically generate a script to read from the news data. Obtain information about the virtual announcer selected by the user and combine it with the script to generate the final video material.
[0843] 4. Emotion Engine:
[0844] At the start, the system analyzes the user's emotions and provides the results to the server, which then adaptively adjusts the virtual announcer's facial expressions and tone of voice based on the analysis results.
[0845] 5. Live Streaming:
[0846] The server uploads the generated video material to the live streaming server and provides a stream URL to the device, which then plays the video in real time.
[0847] 6. Users:
[0848] Users can open the app during their commute and watch a virtual announcer talk about the results of last night's game. While watching, the virtual announcer changes its facial expression and tone of voice depending on the user's emotions. It also provides related news based on the user's emotions.
[0849] 7. Feedback:
[0850] After watching the news, users provide feedback within the app, which is then sent to the server and used to improve the service.
[0851] This system allows users to enjoy a highly personalized, emotionally-driven news experience in real time.
[0852] The processing flow will be explained below.
[0853] ---
[0854] Step 1:
[0855] Users install the news delivery application on their device and create an account. After creating the account, they log in by entering their email address and password on the login page.
[0856] Step 2:
[0857] The terminal sends the user's login information to the server for authentication. If authentication is successful, the server issues a session ID and returns it to the terminal. The terminal saves the session ID and maintains the logged-in state.
[0858] Step 3:
[0859] Users access the app's settings screen, select their preferred news genre (e.g., sports, technology, politics, etc.), customize their preferred virtual announcer (e.g., gender, tone of voice, appearance, etc.), and click the "Save" button.
[0860] Step 4:
[0861] The terminal transmits the user's selected news genre and virtual announcer setting information to the server, which then stores the received customization information in the user database.
[0862] Step 5:
[0863] The server retrieves the latest news data from multiple news APIs, filters the retrieved news data based on the genre selected by the user, and extracts the corresponding news data.
[0864] Step 6:
[0865] The server uses a generative AI model to automatically generate a speech transcript from the filtered news data, which is then stored in a database.
[0866] Step 7:
[0867] The server retrieves the information of the virtual announcer selected by the user and combines it with the script to generate the final video material. Specifically, it generates the reading voice using a voice synthesizer and creates the character movements using 3D modeling.
[0868] Step 8:
[0869] The device uses a camera and microphone to collect the user's facial expressions and voice data in real time and transmits it to a server.
[0870] Step 9:
[0871] The server uses an emotion engine to analyze the user's facial expression and audio data sent from the device and recognize the user's emotional state. The recognition results are classified as emotions such as smile, surprise, sadness, etc.
[0872] Step 10:
[0873] Based on the analysis results of the emotion engine, the server dynamically adjusts the virtual announcer's facial expressions and tone of voice during live broadcasts, enabling a more natural and approachable news delivery that responds to the user's emotions.
[0874] Step 11:
[0875] The server uploads the generated news video of the virtual announcer to the live streaming server and sends the stream URL to the terminal.
[0876] Step 12:
[0877] The device uses the stream URL received from the server to play the video in real time, and the user opens the app to watch the virtual announcer change his or her facial expressions and tone of voice according to the user's emotions.
[0878] Step 13:
[0879] During live streaming, the server dynamically recommends relevant news content based on the user's emotional state, for example, if the user expresses surprise, it will display the latest news related to surprise.
[0880] Step 14:
[0881] After watching the news, the user inputs feedback information using a feedback form, which is then sent to the server by the terminal.
[0882] Step 15:
[0883] The server stores user feedback information in a database and uses it for data analysis to improve services and optimize generative AI models.
[0884] This concludes the specific processing flow of the system, which enables users to enjoy a personalized, emotion-sensitive news experience in real time.
[0885] Example 2
[0886] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0887] Current digital news delivery systems only provide one-way information to users, limiting personalization based on user emotions and preferences. Furthermore, passive news consumption makes it difficult to provide an experience tailored to users' interests and engagement levels. Therefore, there is a need for a news delivery system that can capture users' interest and adjust in real time to reflect their emotions.
[0888] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for selecting information of a genre that the user likes, means for the user to customize a virtual character, means for acquiring information data and filtering it based on the genre selected by the user, means for generating a text-to-speech script from the filtered information data, means for having the virtual character read the generated text-to-speech script, means for recognizing the user's emotions in real time using the camera and microphone of the terminal and transmitting the data, means for analyzing the emotions and dynamically adjusting the virtual character's facial expression and tone of voice based on the analysis results, and means for providing information to the user in a live streaming format. This enables highly personalized news delivery according to the user's emotions.
[0889] "User" refers to a general user who receives information and performs settings and customization.
[0890] A "virtual character" is a digital avatar that provides information to users and has audio and visual representations.
[0891] "Information data" refers to digital information such as text, audio, and video that is acquired based on the genre selected by the user.
[0892] "Filtering" refers to the process of selecting only relevant data from the acquired information data based on user selections.
[0893] A "reading manuscript" refers to text that is analyzed and generated by a generative AI model based on information data and is read aloud by a virtual character.
[0894] A "generative AI model" refers to an artificial intelligence model that generates and analyzes text based on large amounts of data.
[0895] "Emotion recognition" refers to a technology that uses the device's camera and microphone to analyze emotions from the user's facial expressions and voice.
[0896] "Dynamic adjustment" refers to the process of changing a virtual character's facial expression or tone of voice based on the results of a user's emotional analysis in real time.
[0897] "Live streaming" refers to a method of providing generated video material to users in real time.
[0898] "Feedback" refers to opinions and impressions provided by users after viewing.
[0899] "Customization" refers to the process by which a user sets the gender, tone of voice, character style, etc. of a virtual character according to their own preferences.
[0900] This invention is a system that allows users to select information from their favorite genres and delivers personalized information via a customized virtual character. This system achieves even higher levels of personalization by recognizing the user's emotions in real time and dynamically adjusting the virtual character's facial expressions and tone of voice based on that information.
[0901] System configuration
[0902] User
[0903] 1. The user installs a news delivery application on their device (smartphone, tablet, PC, etc.), creates an account, and logs in.
[0904] 2. After successfully logging in, you can select and customize your news genre (e.g., "Sports," "Politics," "Technology," etc.) and virtual character (gender, voice tone, character style, etc.) on the settings screen.
[0905] Terminal
[0906] 1. The device is equipped with a camera and microphone, which are used to recognize the user's emotions (facial expressions and tone of voice) in real time.
[0907] 2. The user's selected settings information (news genre, virtual character customization) is saved in a local database and sent to the server.
[0908] server
[0909] 1. News data collection and filtering: The server retrieves the latest news data from various news APIs (e.g., NewsAPI, Google News API) and filters it based on the genre selected by the user.
[0910] 2. Generate a prompt: The server uses a generative AI model (e.g., OpenAI GPT-3) to automatically generate a prompt from the filtered news articles.
[0911] Sample prompt: "Create a script based on the latest sports news article to be read aloud by a fictional character who is an active and energetic man."
[0912] 3. Virtual character generation: The server acquires the virtual character information (e.g., character model, voice tone) set by the user and combines it with the reading script to generate the final video material. Specifically, the reading voice is generated using a voice synthesizer (e.g., Amazon Polly), and the virtual character's movements are created using 3D modeling software (e.g., Blender).
[0913] 4. Emotion recognition using an emotion engine: The server receives the user's camera footage and audio data sent from the device and analyzes the user's emotions using an emotion recognition engine (e.g., Affectiva SDK). Based on the analysis results (e.g., smile, surprise, sadness), the virtual character's facial expressions and tone of voice are dynamically adjusted.
[0914] 5. Live streaming: The final video material is uploaded to the live streaming server, and a stream URL is generated and sent to the device.
[0915] Specific examples
[0916] 1. User: Install the news distribution app, create an account, and log in. Select the "Sports" genre on the settings screen and set your preferred virtual character (e.g., a male voice with an active personality).
[0917] 2. Device: The device stores the selection information locally and sends it to the server. It also collects the user's facial expressions and voice data in real time via the camera and microphone and sends them to the server.
[0918] 3. Server: Collects and filters sports-related news data. Automatically generates a script to read from the news data using a generative AI model. Information about the virtual character set by the user is acquired, and combined with the script to generate the final video material.
[0919] 4. Emotion Engine: At the start, the system analyzes the user's emotions and provides the results to the server, which then adaptively adjusts the virtual character's facial expressions and tone of voice based on the analysis results.
[0920] 5. Live streaming: The server uploads the generated video material to the live streaming server and provides a stream URL to the device, which then plays the video in real time.
[0921] 6. User: Watch the news and hear a virtual character say, "I'll tell you the results of last night's game." While watching, the virtual character changes its facial expression and tone of voice according to the user's emotions. Related news is also provided.
[0922] 7. Feedback: After watching the news, users can provide feedback within the app, which will be sent to the server and used to improve the service.
[0923] This system allows users to enjoy a highly personalized, emotionally-driven news experience in real time.
[0924] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0925] Step 1: Initial System Setup and Login
[0926] operation:
[0927] Users install the news distribution application on their device, create an account, and then log in.
[0928] input:
[0929] Your account information (email address, password, social media accounts, etc.).
[0930] output:
[0931] Session information indicating that the user logged in successfully.
[0932] Specific behavior:
[0933] A user launches the app and enters their account information on the login screen. The entered information is sent to the server, where authentication is performed. If successful, session information is returned to the device.
[0934] Step 2: Setting the news genre and virtual characters
[0935] operation:
[0936] Users select a news genre on the settings screen and customize their preferred virtual character.
[0937] input:
[0938] Customization information for the news genre and virtual character selected by the user.
[0939] output:
[0940] The setting information is stored on the device and sent to the server.
[0941] Specific behavior:
[0942] Users select a news genre (e.g., sports, politics, technology, etc.) on the app's settings screen and customize the attributes of their virtual character (e.g., gender, tone of voice, appearance, etc.). The selection information is stored in a local database and transmitted to a server over the network.
[0943] Step 3: Collecting and filtering news data
[0944] operation:
[0945] The server retrieves the latest news data from the news API and filters it based on the genre selected by the user.
[0946] input:
[0947] News data obtained from the news API, news genres selected by the user.
[0948] output:
[0949] News data filtered to user-selected genres.
[0950] Specific behavior:
[0951] The server sends a request to a news API (e.g., NewsAPI, Google News API) to retrieve the latest news data, which is then filtered by keywords and categories based on the user's selection to select only relevant news.
[0952] Step 4: Generate a transcript
[0953] operation:
[0954] The server uses a generative AI model to automatically generate a reading script from filtered news data.
[0955] input:
[0956] Filtered news data, generative AI models (e.g., OpenAI GPT-3).
[0957] output:
[0958] The generated reading transcript.
[0959] Specific behavior:
[0960] The server inputs the filtered news data into a generative AI model and generates a reading script using a prompt, such as "Please create a reading script for a lively and energetic male virtual character based on the latest sports news article." The generated script is then stored in a database.
[0961] Step 5: Generate a virtual character
[0962] operation:
[0963] The server combines the information from the virtual character with the reading script to generate the final video material.
[0964] input:
[0965] Virtual character customization information, generated reading script.
[0966] output:
[0967] A video clip of a virtual character reading the news.
[0968] Specific behavior:
[0969] The server retrieves the virtual character's customization information, generates the reading voice using a voice synthesizer (e.g., Amazon Polly), and generates the virtual character's movements and facial expressions using 3D modeling software (e.g., Blender). These are then combined to generate the final video material.
[0970] Step 6: Emotion recognition and dynamic regulation
[0971] operation:
[0972] The device's camera and microphone are used to collect the user's emotions in real time and send the data to a server.
[0973] input:
[0974] User's facial expression data, voice data.
[0975] output:
[0976] User sentiment analysis results.
[0977] Specific behavior:
[0978] The device uses a camera and microphone to collect the user's facial expressions and tone of voice in real time, and this data is sent to a server. An emotion recognition engine (e.g., Affectiva SDK) on the server analyzes this data and classifies emotional states such as smile, sadness, or surprise. Based on the analysis results, the virtual character's facial expressions and tone of voice are dynamically adjusted.
[0979] Step 7: Live Stream
[0980] operation:
[0981] The server uploads the generated video material to a live streaming server, generates a stream URL, and sends it to the terminal.
[0982] input:
[0983] The final video footage produced.
[0984] output:
[0985] The stream URL for live streaming.
[0986] Specific behavior:
[0987] The server uploads the video material to the live streaming server and generates a stream URL. This URL is sent to the device and can be accessed by the user. The device uses this URL to play the video in real time and displays a virtual character reading the news.
[0988] Step 8: Gather feedback
[0989] operation:
[0990] Users provide feedback after watching the news.
[0991] input:
[0992] User feedback information.
[0993] output:
[0994] Feedback information collected.
[0995] Specific behavior:
[0996] Users enter their opinions and thoughts into the feedback form within the app. This feedback information is collected by the device and sent to the server, which stores it in a database and uses it to improve the service in the future.
[0997] (Application example 2)
[0998] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0999] Conventional news delivery systems have limited personalization based on user emotions and provide a uniform news viewing experience, making it difficult to increase user satisfaction. Furthermore, they lack interactivity in news delivery because they are unable to reflect users' real-time reactions.
[1000] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for selecting a news genre that the user likes, means for the user to customize the virtual announcer, means for acquiring news data and filtering it based on the genre selected by the user, means for generating a script to be read from the filtered news data, means for having the virtual announcer read the generated script to be read, means for providing news to the user in a live streaming format, means for analyzing the user's emotions, and means for dynamically adjusting the virtual announcer's facial expressions and tone of voice based on the analyzed emotions. This enables a highly personalized news experience that corresponds to the user's emotions.
[1001] "Means for users to select news genres they like" refers to interfaces and functions that allow users to select news genres that interest them.
[1002] "Means for users to customize their virtual announcers" refers to interfaces and functions that allow users to set the announcer's appearance, voice, personality, etc.
[1003] "Means for acquiring news data and filtering it based on the genre selected by the user" is a function for collecting data from news sources and extracting news according to the genre selected by the user.
[1004] The "means for generating a script to be read aloud from filtered news data" is a function that automatically creates a script to be read aloud by a virtual announcer based on extracted news data.
[1005] The "means for having a virtual announcer read out the generated script" is a function for playing back the generated script in the voice of a virtual announcer.
[1006] The "means for providing news to users in a live distribution format" is a function for distributing news content to users in real time.
[1007] "Means for analyzing user emotions" refers to technology that analyzes a user's facial expressions and tone of voice in real time to recognize their emotional state.
[1008] "Means for dynamically adjusting the facial expressions and tone of voice of a virtual announcer based on analyzed emotions" refers to technology that changes the facial expressions and voice of a virtual announcer in real time based on the results of emotion analysis.
[1009] This invention is a system that achieves even higher levels of personalization by combining a news delivery system with a user-selected news genre and a customized virtual announcer with an emotion engine that recognizes the user's emotions. This system recognizes the user's emotions in real time using a smartphone, and based on that information, adjusts the virtual announcer's facial expressions and tone of voice, and recommends related news content.
[1010] System configuration
[1011] User terminal
[1012] Users install the "Emotion-Responsive News" application on their smartphones, create an account, and then log in. After logging in, they can select and customize the news genre and virtual announcer on the settings screen. The smartphones are also equipped with cameras and microphones, which are used to recognize the user's emotions in real time.
[1013] server
[1014] The server has the following main functions:
[1015] 1. News data collection and filtering
[1016] The server retrieves the latest news data from the news API and filters it based on the user's selected genre.
[1017] 2. Generation of reading script
[1018] The server uses a generative AI model to automatically generate a reading script from filtered news articles.
[1019] 3. Creating a Virtual Announcer
[1020] The server retrieves information about the virtual announcer set by the user, combines it with the script to be read, and generates the final video material using a voice synthesizer and 3D modeling.
[1021] 4. Emotion Recognition by Emotion Engine
[1022] The server receives the user's camera footage and audio data sent from the smartphone and analyzes the user's emotions in real time. The analysis results are classified into emotional states such as smile, surprise, sadness, etc.
[1023] 5. Live streaming and dynamic adjustments
[1024] The server dynamically adjusts the virtual announcer's facial expressions and tone of voice during live broadcasts based on the analysis results of the emotion engine, and can also recommend relevant news content based on the user's emotional state.
[1025] 6. Collecting User Feedback
[1026] After watching the news, feedback is collected from users to help improve the service. The feedback information is stored on the server and used for data analysis.
[1027] Specific examples
[1028] If the user selects sports news
[1029] 1. Users
[1030] Users install the "Emotionally Responsive News" app on their smartphone, create an account, and log in. On the settings screen, they select the "Sports" genre and then set their preferred virtual announcer (e.g., a man with a lively personality).
[1031] 2. Terminal
[1032] The selection information is stored on the local device and sent to the server. In addition, facial expression and voice data of the user are collected in real time via the smartphone's camera and microphone and sent to the server.
[1033] 3. Server
[1034] The server collects sports-related news data and filters it through a news API. It uses a generative AI model to automatically generate a script to read from the news data. It then obtains information about the virtual announcer selected by the user, combines it with the script, and generates the final video material using a voice synthesizer and 3D modeling.
[1035] 4. Emotion Engine
[1036] The server analyzes the video and audio transmitted from the user's smartphone to identify the user's emotional state in real time, and adaptively adjusts the virtual announcer's facial expressions and tone of voice based on the analysis results.
[1037] 5. Live Streaming
[1038] The server uploads the generated video material to a live streaming server and provides a stream URL to the device. The smartphone uses this URL to play the video in real time, and a virtual announcer announces, "I'll tell you the results of last night's game."
[1039] 6. Users
[1040] Users can open the app during their commute and watch a virtual announcer speak. While watching, the virtual announcer changes its facial expression and tone of voice depending on the user's emotions. It also provides relevant news based on the user's emotions.
[1041] Prompt Sentence Examples
[1042] Leveraging a generative AI model, we use prompts like:
[1043] Please create a news script to use in your presentation. It should be based on recent news in the sports genre and should include the following:
[1044] Last night's match results
[1045] Performances of noteworthy players
[1046] Upcoming match schedule
[1047] Please keep it easy to read and in a friendly tone."
[1048] This allows users to enjoy a highly personalized news experience that responds to their emotions in real time.
[1049] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1050] Step 1:
[1051] A user installs the "Emotion Response News" app on their smartphone, creates an account, and logs in. As input, the user provides the necessary registration information, and as output, an account is created. Specifically, information such as a username, password, and email address is entered, and the user's account data is generated based on this.
[1052] Step 2:
[1053] Users select and customize their news genre and virtual announcer on the app's settings screen. The input is the user's genre selection and announcer customization information, and the output is saved to the local device. Specifically, this includes genre selections such as "sports" and "entertainment," as well as the announcer's appearance, voice tone, and personality settings.
[1054] Step 3:
[1055] The terminal sends the setting information of the news genre and virtual announcer to the server. The user's selection information is provided as input, and the server receives this information as output and stores it in a database. Specifically, the setting information is sent to the server as JSON format data.
[1056] Step 4:
[1057] The server retrieves the latest news data from the news API and filters the data based on the genre selected by the user. As input, news data from the news API is provided, and as output, news data narrowed down to a genre is generated. Specifically, the server retrieves news data using an API key and filters the news according to the specified genre.
[1058] Step 5:
[1059] The server uses a generative AI model to automatically generate a reading script from the filtered news articles. The filtered news data is provided as input, and text data for reading is generated as output. Specifically, a prompt sentence is input to the generative AI model (such as OpenAI GPT) and the generated text is obtained.
[1060] Step 6:
[1061] The server combines the generated script with the virtual announcer information set by the user and generates the final video material using a voice synthesizer and 3D modeling. The script and announcer settings are provided as input, and video material is generated as output. Specifically, the voice synthesizer generates audio data and links it to the 3D model to create a video.
[1062] Step 7:
[1063] The device uses the smartphone's camera and microphone to collect the user's facial expressions and voice data in real time and transmits it to a server. The user's real-time video and audio are provided as input, and emotion analysis data is generated as output. Specifically, emotion analysis is performed using the Microsoft Azure Face API and Google Cloud Vision API.
[1064] Step 8:
[1065] The server analyzes the user's emotional data provided in real time and dynamically adjusts the virtual announcer's facial expressions and tone of voice. The emotion analysis results are provided as input, and the announcer's facial expressions and tone of voice are adjusted as output. Specifically, the synthesizer parameters are adjusted based on the emotional data, changing the facial expressions of the 3D model.
[1066] Step 9:
[1067] The server uploads the generated video material to the distribution server in real time and provides a stream URL to the terminal. The completed video material is provided as input, and a stream URL is generated as output. Specifically, the video upload process to the distribution server is executed.
[1068] Step 10:
[1069] The device plays the video in real time based on the stream URL and provides news to the user. The stream URL is provided as input and the video is played as output. Specifically, the video player receives the stream and plays it.
[1070] Step 11:
[1071] After watching the news, the user provides feedback, which the device sends to the server. The user's feedback is provided as input, and the server receives and stores the feedback data as output. Specifically, the user taps the feedback icon on the screen and enters their opinion or impression.
[1072] Step 12:
[1073] The server uses the collected feedback data for data analysis to help improve the service. User feedback data is provided as input, and improvement measures are formulated as output. Specifically, the feedback data is analyzed using a data analysis tool, and areas for improvement are listed.
[1074] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1075] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1076] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[1077] [Third embodiment]
[1078] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[1079] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[1080] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1081] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[1082] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1083] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1084] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1085] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1086] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1087] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1088] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1089] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[1090] ---
[1091] This invention is a system for providing viewers with a personalized news experience, allowing users to select a news genre and a virtual announcer and watch customized live news. This system has functions that allow users to select news genres and virtual announcers according to their preferences, and functions that automatically generate a text-to-speech script from news data and have the virtual announcer read it aloud. It also has a function for improving the service based on user feedback.
[1092] System configuration
[1093] User terminal
[1094] Users install the news distribution application on their device, create an account, and then log in. Once successfully logged in, users can select news genres and virtual announcers. The user's selection information is saved on the device and sent to the server.
[1095] server
[1096] The server has the following main functions:
[1097] 1. News data collection:
[1098] The server retrieves the latest news data from various news APIs. This data is aggregated from multiple news sources.
[1099] 2. Filtering and generating a transcript:
[1100] The collected news data is filtered based on the news genre selected by the user, and then a generative AI model is used to automatically generate a reading transcript from the filtered news articles.
[1101] 3. Creating a virtual announcer:
[1102] The virtual announcer reads the script based on the user-customized virtual announcer settings (voice, appearance, etc.), and then generates the final video material for live streaming.
[1103] 4. Live Streaming:
[1104] The generated virtual announcer will broadcast live news videos in real time over the Internet.
[1105] 5. Collecting User Feedback:
[1106] After watching the news, we collect feedback from users to help improve the service. The feedback information is stored on the server and can be used for data analysis.
[1107] Specific examples
[1108] If the user selects sports news
[1109] 1. User:
[1110] Install the news distribution app and create an account. After logging in, select the "Sports" genre on the settings screen and then set your preferred virtual announcer (e.g., a male voice with an active personality).
[1111] 2. Terminal:
[1112] The selection information is stored on the local device and transmitted to the server.
[1113] 3. Server:
[1114] Sports-related news data is collected and filtered. Then, using a generative AI model, a script is automatically generated from the filtered news data. The final video material is generated based on the script and the settings of the virtual announcer selected by the user.
[1115] 4. Terminal:
[1116] The user's device receives the streaming URL from the server and plays the sports news in real time.
[1117] 5. User:
[1118] On the way to work, users can open the app and watch a virtual announcer announce the results of last night's game. After watching, users can provide feedback within the app.
[1119] This system allows for a personalized news experience, increasing viewer engagement.
[1120] The processing flow will be explained below.
[1121] ---
[1122] Step 1:
[1123] Users install the news delivery application on their device and create an account. After creating the account, they log in by entering their email address and password on the login page.
[1124] Step 2:
[1125] The terminal sends the user's login information to the server for authentication. If authentication is successful, the server issues a session ID and returns it to the terminal. The terminal saves the session ID and maintains the logged-in state.
[1126] Step 3:
[1127] Users access the app's settings screen, select their preferred news genre (e.g., sports, technology, politics, etc.), customize their preferred virtual announcer (e.g., gender, tone of voice, appearance, etc.), and click the "Save" button.
[1128] Step 4:
[1129] The terminal transmits the user's selected news genre and virtual announcer setting information to the server, which then stores the received customization information in the user database.
[1130] Step 5:
[1131] The server retrieves the latest news data from multiple news APIs, filters the retrieved news data based on the genre selected by the user, and extracts the corresponding news data.
[1132] Step 6:
[1133] The server uses a generative AI model to automatically generate a speech transcript from the filtered news data, which is then stored in a database.
[1134] Step 7:
[1135] The server retrieves the information of the virtual announcer selected by the user and combines it with the script to generate the final video material. Specifically, it generates the reading voice using a voice synthesizer and creates the character movements using 3D modeling.
[1136] Step 8:
[1137] The server uploads the generated news video of the virtual announcer to the live streaming server and generates a stream URL, which is then sent to the device.
[1138] Step 9:
[1139] The device plays the video in real time based on the stream URL received from the server, and users can open the app at their preferred time to watch customized live news.
[1140] Step 10:
[1141] After watching the news, users can rate the announcer's performance and the content of the news in the feedback form within the app and enter the feedback information on their device, which then sends the feedback information to the server.
[1142] Step 11:
[1143] The server stores user feedback information in a database and uses it for data analysis to help improve generative AI models and services.
[1144] This completes the system's processing flow, allowing users to enjoy a personalized news experience in real time.
[1145] Example 1
[1146] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1147] Conventional news delivery systems make it difficult for viewers to select news that suits their preferences and watch it in a customized format. Furthermore, they lacked a mechanism for effectively collecting viewer feedback and utilizing it to improve services. As a result, they were unable to fully optimize the viewing experience or improve viewer satisfaction.
[1148] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1149] In this invention, the server includes a means for acquiring news data and filtering it based on a genre selected by the user, a means for generating a script from the filtered news data using a generative AI model, and a means for having a virtual announcer read the generated script. This allows users to select news from their favorite genres and watch live news broadcasts by customized virtual announcers. Furthermore, by collecting user feedback and using it to improve the service, viewer satisfaction can be increased.
[1150] 1. "A means for users to select news genres they like" refers to a function that allows users to select news categories that interest them.
[1151] 2. "Means for users to customize the virtual announcer" refers to the functionality that allows users to change settings such as the voice and appearance of the virtual announcer.
[1152] 3. "Means of acquiring news data" refers to the function of collecting the latest news information from external news data providers.
[1153] 4. "Means of filtering based on the genre selected by the user" refers to the function of extracting only information that falls into the category selected by the user from the collected news data.
[1154] 5. "Means for generating a text-to-speech transcript from filtered news data" refers to a function that uses a generative AI model to summarize filtered news information and create text that can be read aloud.
[1155] 6. "Means for having a virtual announcer read the generated script" refers to the function of converting text data into voice using speech synthesis technology and synchronizing it with the actions of the virtual announcer.
[1156] 7. "Means for providing users with generated news videos in a live streaming format" refers to the function of delivering generated audio and video to users in real time via the Internet.
[1157] 8. "Means of collecting user feedback" refers to the function of collecting ratings and comments from viewers and storing them in a database.
[1158] 9. "Means used to improve services" refers to the functions used to analyze collected feedback and improve the quality of news provision services.
[1159] 10. "Generative AI model" refers to an artificial intelligence model that generates text data using natural language processing technology.
[1160] 11. “Prompt” refers to a command or question input to a generative AI model.
[1161] These are the definitions of the important words included in this system.
[1162] The present invention is a system for providing users with a personalized news experience, allowing them to select a news genre and a virtual announcer and watch customized live news. This system has functions that allow users to select news genres and virtual announcers according to their preferences, and functions that automatically generate a text-to-speech script from news data and have the virtual announcer read it aloud. It also has a function to improve the service based on user feedback.
[1163] System configuration
[1164] User terminal
[1165] Users install the news distribution application on their device, create an account, and then log in. Once successfully logged in, users can select news genres and virtual announcers. The user's selection information is saved on the device and sent to the server. This can be done using common mobile devices such as smartphones and tablets.
[1166] server
[1167] The server has the following main functions:
[1168] 1. News data collection:
[1169] The server retrieves the latest news data from various news providers. This data is collected from multiple news sources. Specifically, "News API" and "News Data Provider Services" are used.
[1170] 2. Filtering and generating a transcript:
[1171] The collected news data is filtered based on the news genre selected by the user. Then, a generative AI model is used to automatically generate a speech transcript from the filtered news articles. The generative AI model used here is, for example, GPT-4.
[1172] 3. Creating a virtual announcer:
[1173] The virtual announcer reads the script based on the user's customized settings (voice, appearance, etc.). The final video material for live streaming is then generated. Speech synthesis software (e.g., TTS technology) is used for the speech synthesis, and 3D animation software (e.g., animation generation software) is used to generate the video.
[1174] 4. Live Streaming:
[1175] The generated news video of the virtual announcer is broadcast live. The live broadcast is carried out in real time over the Internet. A "live streaming service," for example, is used as the streaming platform.
[1176] 5. Collecting User Feedback:
[1177] After watching the news, we collect feedback from users to help improve the service. The feedback information is stored on the server and can be used for data analysis.
[1178] Specific examples
[1179] If the user selects sports news
[1180] 1. User:
[1181] Install the news distribution app and create an account. After logging in, select the "Sports" genre on the settings screen and then set your preferred virtual announcer (e.g., a male voice with an active personality).
[1182] 2. Terminal:
[1183] The selection information is stored on the local device and transmitted to the server.
[1184] 3. Server:
[1185] Sports-related news data is collected and filtered. Then, using a generative AI model, a script is automatically generated from the filtered news data. The final video material is generated based on the script and the settings of the virtual announcer selected by the user.
[1186] 4. Terminal:
[1187] The user's device receives the streaming URL from the server and plays the sports news in real time.
[1188] 5. User:
[1189] On the way to work, users can open the app and watch a virtual announcer announce the results of last night's game. After watching, users can provide feedback within the app.
[1190] Prompt Sentence Examples
[1191] "I want to watch sports news. I want the virtual announcer to be a lively male."
[1192] This system will enable viewers to enjoy a news experience that is optimized for each individual, increasing engagement.
[1193] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1194] Step 1:
[1195] A user installs a news delivery app and creates an account
[1196] Input: User's electronic device (smartphone, tablet) and internet connection
[1197] How it works: A user downloads and installs a news delivery app from the app store. They launch the app and create an account by entering their email address and password on the account creation screen that appears when the app is first launched.
[1198] Output: Account information is created and you can log in to the app. After logging in, you will be taken to a screen where you can select the news genre and virtual announcer.
[1199] Step 2:
[1200] The user selects the news genre and virtual announcer.
[1201] Input: User preference information (selected news genre, virtual announcer settings)
[1202] How it works: Within the app, users select "Sports" from a range of news genres, including "Sports," "Politics," and "Entertainment." They then customize the virtual announcer's settings screen, such as selecting a "lively male voice."
[1203] Output: The selection information is temporarily stored in the device's local storage and then sent to the server.
[1204] Step 3:
[1205] The server collects news data
[1206] Input: API requests and responses received by the server from external news providers
[1207] How it works: The server sends a request to an API endpoint such as a "news data provider service" to retrieve the latest news data in JSON format.
[1208] Output: The retrieved news data is stored in a database for further processing.
[1209] Step 4:
[1210] The server filters the news data and generates a reading script.
[1211] Input: News data to be filtered and user genre selection information
[1212] How it works: The server uses a filtering algorithm (e.g., TF-IDF) to extract only news data that falls within the user's selected "sports" genre. It then uses a generative AI model (e.g., GPT-4) to generate a summary of the filtered news data.
[1213] Output: The filtered news data and the generated speech transcripts are stored in a database.
[1214] Step 5:
[1215] The server configures the virtual announcer and generates the news video.
[1216] Input: Reading script and user's virtual announcer setting information
[1217] How it works: The server uses speech synthesis software (e.g., TTS technology) to convert the generated script into audio data, and then uses 3D animation software (e.g., animation generation software) to generate a video that coordinates the voice and movements of the virtual announcer.
[1218] Output: The created news video is stored on the server and ready to be streamed.
[1219] Step 6:
[1220] Server-generated live news video broadcast
[1221] Input: Generated news video data
[1222] How it works: The server uses a live streaming service to create a streaming URL for the generated news video, and then sends this URL to the user's device.
[1223] Output: A response containing a streaming URL is sent to the user's device.
[1224] Step 7:
[1225] User watches news
[1226] Input: Streaming URL
[1227] How it works: The user's device receives the streaming URL and uses the media player component to play the news video in real time.
[1228] Output: User watches live streaming news video.
[1229] Step 8:
[1230] Users provide feedback
[1231] Input: User feedback information (ratings, comments, etc.)
[1232] How it works: After watching a news item, users enter their rating and comments about their viewing experience on the feedback screen displayed within the app. The feedback is then sent from the app to the server.
[1233] Output: The feedback information is sent to the server and stored in a database.
[1234] The above is a detailed description of the specific processing steps of the present system. The present invention is expected to provide users with an individually customized news experience, thereby improving viewer engagement and satisfaction.
[1235] (Application example 1)
[1236] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1237] In today's world, with so much news being distributed, it is difficult for individual users to accurately obtain only the news that interests them. There are also insufficient means to customize the visual and auditory experience of news for each individual user. Furthermore, there is a lack of mechanisms for users to provide feedback on the news they have viewed and use that feedback to improve services. Effective methods to solve these issues are needed.
[1238] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1239] In this invention, the server includes means for acquiring news data and filtering it based on a genre selected by the user, means for generating a script to be read from the filtered news data, and means for having a virtual announcer read the generated script to be read. This allows each user to watch news in the genres they are interested in in real time with a customized virtual announcer. Furthermore, it is possible to stream news to a smartphone application in real time and collect feedback from users to help improve the service.
[1240] "User" refers to an individual who uses this system to receive news distribution services.
[1241] "Genre" refers to the type or category of news, such as a specific field such as sports, economics, or entertainment.
[1242] A "virtual announcer" is a computer-generated character whose role is to read the news aloud and visually.
[1243] "News Data" refers to the collection of information collected from various news sources and refers to the material used for filtering and generating the readings.
[1244] "Filtering" refers to the process of narrowing down news data based on genres selected by the user.
[1245] A "script" is a text automatically generated from filtered news data, and refers to a script that a virtual announcer will read aloud.
[1246] "Live streaming" refers to a format in which news is provided to users in real time, and refers to a method of streaming directly from a server to a user's terminal.
[1247] "Smartphone application" refers to a program that runs on a smartphone and functions as the user interface for this system.
[1248] "Real-time streaming" refers to a technology in which news videos are sent to the user's device as soon as they are generated and played back instantly.
[1249] "Feedback" refers to the opinions and evaluations provided by users regarding the news and services they have viewed, and serves as important data for improving services.
[1250] A "generative AI model" refers to an artificial intelligence algorithm that uses machine learning technology to automatically generate a reading script from news data.
[1251] The present invention is a system that allows users to customize the news genres and virtual announcers that interest them, providing a personalized news experience.
[1252] System configuration
[1253] User terminal
[1254] Users install the corresponding application on their smartphone, create an account, and log in. After logging in, users can configure their news genre and virtual announcer settings. This includes selecting a genre (e.g., sports, economics, entertainment, etc.) and customizing the announcer's appearance and voice. This configuration information is saved on the user's device and sent to the server.
[1255] server
[1256] The server has the following main functions:
[1257] 1. News data collection:
[1258] The server uses news APIs to collect the latest news data from various news sources, examples of which include NewsAPI and Google News API.
[1259] 2. Filtering and generating a transcript:
[1260] The collected news data is filtered based on the genre selected by the user, and a transcription is automatically generated from the filtered news articles using a generative AI model (e.g., GPT-3).
[1261] 3. Creating a virtual announcer:
[1262] A virtual announcer reads the script according to the announcer settings customized by the user, using tools such as Adobe Character Animator and Unity3D. The generated news video is then streamed to the user's device in real time.
[1263] 4. Gathering Feedback:
[1264] After viewing the news, we collect feedback from users and use it to improve our services, thereby continuously improving the quality of news and its delivery methods.
[1265] Process Overview
[1266] 1. User Settings:
[1267] For example, a user selects "Entertainment News" and "An announcer with a female voice and a cheerful personality." The user's selection information is sent from the smartphone terminal to the server.
[1268] 2. News gathering and filtering:
[1269] The server uses NewsAPI to collect the latest entertainment news and filters the collected data.
[1270] 3. Generate a reading transcript:
[1271] A generative AI model (GPT-3) is used to generate a read-aloud transcript from filtered news articles. For example, a sentence such as "A blockbuster film was announced at the film festival yesterday" is generated.
[1272] 4. Creating a Virtual Announcer:
[1273] Using Adobe Character Animator and Unity3D, a user-customized virtual announcer is generated, and a video is created in which the announcer reads the script.
[1274] 5. Real-time Streaming and Viewing:
[1275] The generated news videos are hosted on a server and streamed in real time to smartphones, where users open the app and listen to the news being read by an announcer.
[1276] 6. Collecting Feedback and Improving Our Services:
[1277] After watching, users provide feedback that helps improve the service.
[1278] Examples and prompts
[1279] Examples:
[1280] A user opens the app and selects a sports news item and a calm, male-voiced announcer. The server collects the latest sports news and uses a generative AI model (GPT-3) to create a script, such as "I'll tell you the results of last night's game." Using Unity3D, a virtual announcer reads the script, and the generated video is delivered to a smartphone. The user can watch the news on their way to work and provide feedback.
[1281] Prompt for the generative AI model:
[1282] News Cache: {Sports}
[1283] Announcer Voice: {Male voice, calm}
[1284] The statement read: "Here are the results of last night's match."
[1285] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1286] Step 1:
[1287] A user creates an account and logs into the smartphone application. The user selects a news genre (e.g., sports, economics, entertainment) and virtual announcer settings (e.g., voice type, gender, appearance). The selection information is stored on the user's device and sent to the server.
[1288] Input: User's genre selection information and virtual announcer settings
[1289] Output: Selections sent to the server
[1290] Specific behavior:
[1291] The user opens the app and navigates to the settings screen.
[1292] Use the drop-down menus and sliders to set your preferred news genres and anchor appearance and voice.
[1293] After the setting information is confirmed, pressing the "Save" button saves the selected information in the local storage of the device and simultaneously sends it to the server.
[1294] Step 2:
[1295] The server uses a news API to collect the latest news data for the genre selected by the user. For example, it uses NewsAPI or Google News API.
[1296] Input: News API request
[1297] Output: Collected news data
[1298] Specific behavior:
[1299] The server accesses the news API periodically or upon user request to collect the latest news data for the selected genre.
[1300] The collected news data is stored in a database on the server.
[1301] Step 3:
[1302] The collected news data is filtered based on the genre selected by the user.
[1303] Input: Collected news data, user genre preferences
[1304] Output: Filtered news data
[1305] Specific behavior:
[1306] An algorithm on the server analyzes the collected news data and extracts only articles related to the genre selected by the user.
[1307] The filtered results are stored in a database.
[1308] Step 4:
[1309] A reading script is automatically generated from filtered news data using a generative AI model (e.g., GPT-3).
[1310] Input: Filtered news data
[1311] Output: Speech manuscript
[1312] Specific behavior:
[1313] The server provides a prompt to the generative AI model, which generates a reading script based on the filtered news data.
[1314] For example, the prompt text is:
[1315] News Cache: {Sports}
[1316] Announcer Voice: {Male voice, calm}
[1317] The statement read: "Here are the results of last night's match."
[1318] The generated reading manuscript is stored in a database on the server.
[1319] Step 5:
[1320] Based on the user's customized settings, a video is generated for the virtual announcer to read the script.
[1321] Input: Reading script, user's virtual announcer settings
[1322] Output: News video with a virtual announcer
[1323] Specific behavior:
[1324] Using Adobe Character Animator and Unity3D, a virtual announcer is generated based on the announcer settings selected by the user.
[1325] The script is read aloud by an announcer, and then generated in video format.
[1326] The generated video is stored on the server.
[1327] Step 6:
[1328] The generated news video of the virtual announcer is streamed in real time and provided to users.
[1329] Input: News video with a virtual announcer
[1330] Output: Real-time streaming video played on user devices
[1331] Specific behavior:
[1332] The server hosts the generated news video and generates a streaming URL.
[1333] The user's device accesses this URL and watches the news video in real time.
[1334] Step 7:
[1335] After watching the news, we collect feedback from users to help improve our services.
[1336] Input: User feedback
[1337] Output: Data for service improvement
[1338] Specific behavior:
[1339] After users watch a news video, they can provide their ratings and opinions through an in-app feedback form.
[1340] The server analyzes the collected feedback data and uses it to improve the service.
[1341] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1342] ---
[1343] This invention is a system that achieves even higher levels of personalization by combining a personalized news delivery system with a user-selected news genre and customized virtual announcer with an emotion engine that recognizes the user's emotions. This system recognizes the user's emotions in real time and, based on that information, adjusts the virtual announcer's facial expressions and tone of voice, and recommends related news content.
[1344] System configuration
[1345] User terminal
[1346] Users install the news distribution application on their device, create an account, and then log in. Once successfully logged in, they can select and customize the news genre and virtual announcer on the settings screen. The device is also equipped with a camera and microphone, which are used to recognize the user's emotions in real time.
[1347] server
[1348] The server has the following main functions:
[1349] 1. News data collection and filtering:
[1350] The server retrieves the latest news data from various news APIs and filters it based on the genre selected by the user.
[1351] 2. Generate a reading transcript:
[1352] The server uses a generative AI model to automatically generate a reading transcript from the filtered news articles.
[1353] 3. Creating a virtual announcer:
[1354] The server retrieves the information of the virtual announcer selected by the user and combines it with the script to generate the final video material. Specifically, it generates the reading voice using a voice synthesizer and creates the character movements using 3D modeling.
[1355] 4. Emotion Recognition with Emotion Engine:
[1356] The server receives the user's camera footage and audio data sent from the device and analyzes the user's emotions in real time. The analysis results are classified into emotional states such as smile, surprise, sadness, etc.
[1357] 5. Live streaming and dynamic adjustment:
[1358] The server dynamically adjusts the virtual announcer's facial expressions and tone of voice during live broadcasts based on the analysis results of the emotion engine, and can also recommend relevant news content based on the user's emotional state.
[1359] 6. Collecting User Feedback:
[1360] After watching the news, we collect feedback from users to help improve the service. The feedback information is stored on the server and can be used for data analysis.
[1361] Specific examples
[1362] If the user selects sports news
[1363] 1. User:
[1364] Install the news distribution app, create an account, and log in. On the settings screen, select the "Sports" genre and then set your preferred virtual announcer (e.g., a male voice with an active personality).
[1365] 2. Terminal:
[1366] The selection information is stored on the local device and sent to the server, and facial expression and voice data of the user are collected in real time via the camera and microphone and sent to the server.
[1367] 3. Server:
[1368] Collect and filter sports-related news data. Use a generative AI model to automatically generate a script to read from the news data. Obtain information about the virtual announcer selected by the user and combine it with the script to generate the final video material.
[1369] 4. Emotion Engine:
[1370] At the start, the system analyzes the user's emotions and provides the results to the server, which then adaptively adjusts the virtual announcer's facial expressions and tone of voice based on the analysis results.
[1371] 5. Live Streaming:
[1372] The server uploads the generated video material to the live streaming server and provides a stream URL to the device, which then plays the video in real time.
[1373] 6. Users:
[1374] Users can open the app during their commute and watch a virtual announcer talk about the results of last night's game. While watching, the virtual announcer changes its facial expression and tone of voice depending on the user's emotions. It also provides related news based on the user's emotions.
[1375] 7. Feedback:
[1376] After watching the news, users provide feedback within the app, which is then sent to the server and used to improve the service.
[1377] This system allows users to enjoy a highly personalized, emotionally-driven news experience in real time.
[1378] The processing flow will be explained below.
[1379] ---
[1380] Step 1:
[1381] Users install the news delivery application on their device and create an account. After creating the account, they log in by entering their email address and password on the login page.
[1382] Step 2:
[1383] The terminal sends the user's login information to the server for authentication. If authentication is successful, the server issues a session ID and returns it to the terminal. The terminal saves the session ID and maintains the logged-in state.
[1384] Step 3:
[1385] Users access the app's settings screen, select their preferred news genre (e.g., sports, technology, politics, etc.), customize their preferred virtual announcer (e.g., gender, tone of voice, appearance, etc.), and click the "Save" button.
[1386] Step 4:
[1387] The terminal transmits the user's selected news genre and virtual announcer setting information to the server, which then stores the received customization information in the user database.
[1388] Step 5:
[1389] The server retrieves the latest news data from multiple news APIs, filters the retrieved news data based on the genre selected by the user, and extracts the corresponding news data.
[1390] Step 6:
[1391] The server uses a generative AI model to automatically generate a speech transcript from the filtered news data, which is then stored in a database.
[1392] Step 7:
[1393] The server retrieves the information of the virtual announcer selected by the user and combines it with the script to generate the final video material. Specifically, it generates the reading voice using a voice synthesizer and creates the character movements using 3D modeling.
[1394] Step 8:
[1395] The device uses a camera and microphone to collect the user's facial expressions and voice data in real time and transmits it to a server.
[1396] Step 9:
[1397] The server uses an emotion engine to analyze the user's facial expression and audio data sent from the device and recognize the user's emotional state. The recognition results are classified as emotions such as smile, surprise, sadness, etc.
[1398] Step 10:
[1399] Based on the analysis results of the emotion engine, the server dynamically adjusts the virtual announcer's facial expressions and tone of voice during live broadcasts, enabling a more natural and approachable news delivery that responds to the user's emotions.
[1400] Step 11:
[1401] The server uploads the generated news video of the virtual announcer to the live streaming server and sends the stream URL to the terminal.
[1402] Step 12:
[1403] The device uses the stream URL received from the server to play the video in real time, and the user opens the app to watch the virtual announcer change his or her facial expressions and tone of voice according to the user's emotions.
[1404] Step 13:
[1405] During live streaming, the server dynamically recommends relevant news content based on the user's emotional state, for example, if the user expresses surprise, it will display the latest news related to surprise.
[1406] Step 14:
[1407] After watching the news, the user inputs feedback information using a feedback form, which is then sent to the server by the terminal.
[1408] Step 15:
[1409] The server stores user feedback information in a database and uses it for data analysis to improve services and optimize generative AI models.
[1410] This concludes the specific processing flow of the system, which enables users to enjoy a personalized, emotion-sensitive news experience in real time.
[1411] Example 2
[1412] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1413] Current digital news delivery systems only provide one-way information to users, limiting personalization based on user emotions and preferences. Furthermore, passive news consumption makes it difficult to provide an experience tailored to users' interests and engagement levels. Therefore, there is a need for a news delivery system that can capture users' interest and adjust in real time to reflect their emotions.
[1414] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for selecting information of a genre that the user likes, means for the user to customize a virtual character, means for acquiring information data and filtering it based on the genre selected by the user, means for generating a text-to-speech script from the filtered information data, means for having the virtual character read the generated text-to-speech script, means for recognizing the user's emotions in real time using the camera and microphone of the terminal and transmitting the data, means for analyzing the emotions and dynamically adjusting the virtual character's facial expression and tone of voice based on the analysis results, and means for providing information to the user in a live streaming format. This enables highly personalized news delivery according to the user's emotions.
[1415] "User" refers to a general user who receives information and performs settings and customization.
[1416] A "virtual character" is a digital avatar that provides information to users and has audio and visual representations.
[1417] "Information data" refers to digital information such as text, audio, and video that is acquired based on the genre selected by the user.
[1418] "Filtering" refers to the process of selecting only relevant data from the acquired information data based on user selections.
[1419] A "reading manuscript" refers to text that is analyzed and generated by a generative AI model based on information data and is read aloud by a virtual character.
[1420] A "generative AI model" refers to an artificial intelligence model that generates and analyzes text based on large amounts of data.
[1421] "Emotion recognition" refers to a technology that uses the device's camera and microphone to analyze emotions from the user's facial expressions and voice.
[1422] "Dynamic adjustment" refers to the process of changing a virtual character's facial expression or tone of voice based on the results of a user's emotional analysis in real time.
[1423] "Live streaming" refers to a method of providing generated video material to users in real time.
[1424] "Feedback" refers to opinions and impressions provided by users after viewing.
[1425] "Customization" refers to the process by which a user sets the gender, tone of voice, character style, etc. of a virtual character according to their own preferences.
[1426] This invention is a system that allows users to select information from their favorite genres and delivers personalized information via a customized virtual character. This system achieves even higher levels of personalization by recognizing the user's emotions in real time and dynamically adjusting the virtual character's facial expressions and tone of voice based on that information.
[1427] System configuration
[1428] User
[1429] 1. The user installs a news delivery application on their device (smartphone, tablet, PC, etc.), creates an account, and logs in.
[1430] 2. After successfully logging in, you can select and customize your news genre (e.g., "Sports," "Politics," "Technology," etc.) and virtual character (gender, voice tone, character style, etc.) on the settings screen.
[1431] Terminal
[1432] 1. The device is equipped with a camera and microphone, which are used to recognize the user's emotions (facial expressions and tone of voice) in real time.
[1433] 2. The user's selected settings information (news genre, virtual character customization) is saved in a local database and sent to the server.
[1434] server
[1435] 1. News data collection and filtering: The server retrieves the latest news data from various news APIs (e.g., NewsAPI, Google News API) and filters it based on the genre selected by the user.
[1436] 2. Generate a prompt: The server uses a generative AI model (e.g., OpenAI GPT-3) to automatically generate a prompt from the filtered news articles.
[1437] Sample prompt: "Create a script based on the latest sports news article to be read aloud by a fictional character who is an active and energetic man."
[1438] 3. Virtual character generation: The server acquires the virtual character information (e.g., character model, voice tone) set by the user and combines it with the reading script to generate the final video material. Specifically, the reading voice is generated using a voice synthesizer (e.g., Amazon Polly), and the virtual character's movements are created using 3D modeling software (e.g., Blender).
[1439] 4. Emotion recognition using an emotion engine: The server receives the user's camera footage and audio data sent from the device and analyzes the user's emotions using an emotion recognition engine (e.g., Affectiva SDK). Based on the analysis results (e.g., smile, surprise, sadness), the virtual character's facial expressions and tone of voice are dynamically adjusted.
[1440] 5. Live streaming: The final video material is uploaded to the live streaming server, and a stream URL is generated and sent to the device.
[1441] Specific examples
[1442] 1. User: Install the news distribution app, create an account, and log in. Select the "Sports" genre on the settings screen and set your preferred virtual character (e.g., a male voice with an active personality).
[1443] 2. Device: The device stores the selection information locally and sends it to the server. It also collects the user's facial expressions and voice data in real time via the camera and microphone and sends them to the server.
[1444] 3. Server: Collects and filters sports-related news data. Automatically generates a script to read from the news data using a generative AI model. Information about the virtual character set by the user is acquired, and combined with the script to generate the final video material.
[1445] 4. Emotion Engine: At the start, the system analyzes the user's emotions and provides the results to the server, which then adaptively adjusts the virtual character's facial expressions and tone of voice based on the analysis results.
[1446] 5. Live streaming: The server uploads the generated video material to the live streaming server and provides a stream URL to the device, which then plays the video in real time.
[1447] 6. User: Watch the news and hear a virtual character say, "I'll tell you the results of last night's game." While watching, the virtual character changes its facial expression and tone of voice according to the user's emotions. Related news is also provided.
[1448] 7. Feedback: After watching the news, users can provide feedback within the app, which will be sent to the server and used to improve the service.
[1449] This system allows users to enjoy a highly personalized, emotionally-driven news experience in real time.
[1450] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1451] Step 1: Initial System Setup and Login
[1452] operation:
[1453] Users install the news distribution application on their device, create an account, and then log in.
[1454] input:
[1455] Your account information (email address, password, social media accounts, etc.).
[1456] output:
[1457] Session information indicating that the user logged in successfully.
[1458] Specific behavior:
[1459] A user launches the app and enters their account information on the login screen. The entered information is sent to the server, where authentication is performed. If successful, session information is returned to the device.
[1460] Step 2: Setting the news genre and virtual characters
[1461] operation:
[1462] Users select a news genre on the settings screen and customize their preferred virtual character.
[1463] input:
[1464] Customization information for the news genre and virtual character selected by the user.
[1465] output:
[1466] The setting information is stored on the device and sent to the server.
[1467] Specific behavior:
[1468] Users select a news genre (e.g., sports, politics, technology, etc.) on the app's settings screen and customize the attributes of their virtual character (e.g., gender, tone of voice, appearance, etc.). The selection information is stored in a local database and transmitted to a server over the network.
[1469] Step 3: Collecting and filtering news data
[1470] operation:
[1471] The server retrieves the latest news data from the news API and filters it based on the genre selected by the user.
[1472] input:
[1473] News data obtained from the news API, news genres selected by the user.
[1474] output:
[1475] News data filtered to user-selected genres.
[1476] Specific behavior:
[1477] The server sends a request to a news API (e.g., NewsAPI, Google News API) to retrieve the latest news data, which is then filtered by keywords and categories based on the user's selection to select only relevant news.
[1478] Step 4: Generate a transcript
[1479] operation:
[1480] The server uses a generative AI model to automatically generate a reading script from filtered news data.
[1481] input:
[1482] Filtered news data, generative AI models (e.g., OpenAI GPT-3).
[1483] output:
[1484] The generated reading transcript.
[1485] Specific behavior:
[1486] The server inputs the filtered news data into a generative AI model and generates a reading script using a prompt, such as "Please create a reading script for a lively and energetic male virtual character based on the latest sports news article." The generated script is then stored in a database.
[1487] Step 5: Generate a virtual character
[1488] operation:
[1489] The server combines the information from the virtual character with the reading script to generate the final video material.
[1490] input:
[1491] Virtual character customization information, generated reading script.
[1492] output:
[1493] A video clip of a virtual character reading the news.
[1494] Specific behavior:
[1495] The server retrieves the virtual character's customization information, generates the reading voice using a voice synthesizer (e.g., Amazon Polly), and generates the virtual character's movements and facial expressions using 3D modeling software (e.g., Blender). These are then combined to generate the final video material.
[1496] Step 6: Emotion recognition and dynamic regulation
[1497] operation:
[1498] The device's camera and microphone are used to collect the user's emotions in real time and send the data to a server.
[1499] input:
[1500] User's facial expression data, voice data.
[1501] output:
[1502] User sentiment analysis results.
[1503] Specific behavior:
[1504] The device uses a camera and microphone to collect the user's facial expressions and tone of voice in real time, and this data is sent to a server. An emotion recognition engine (e.g., Affectiva SDK) on the server analyzes this data and classifies emotional states such as smile, sadness, or surprise. Based on the analysis results, the virtual character's facial expressions and tone of voice are dynamically adjusted.
[1505] Step 7: Live Stream
[1506] operation:
[1507] The server uploads the generated video material to a live streaming server, generates a stream URL, and sends it to the terminal.
[1508] input:
[1509] The final video footage produced.
[1510] output:
[1511] The stream URL for live streaming.
[1512] Specific behavior:
[1513] The server uploads the video material to the live streaming server and generates a stream URL. This URL is sent to the device and can be accessed by the user. The device uses this URL to play the video in real time and displays a virtual character reading the news.
[1514] Step 8: Gather feedback
[1515] operation:
[1516] Users provide feedback after watching the news.
[1517] input:
[1518] User feedback information.
[1519] output:
[1520] Feedback information collected.
[1521] Specific behavior:
[1522] Users enter their opinions and thoughts into the feedback form within the app. This feedback information is collected by the device and sent to the server, which stores it in a database and uses it to improve the service in the future.
[1523] (Application example 2)
[1524] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1525] Conventional news delivery systems have limited personalization based on user emotions and provide a uniform news viewing experience, making it difficult to increase user satisfaction. Furthermore, they lack interactivity in news delivery because they are unable to reflect users' real-time reactions.
[1526] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for selecting a news genre that the user likes, means for the user to customize the virtual announcer, means for acquiring news data and filtering it based on the genre selected by the user, means for generating a script to be read from the filtered news data, means for having the virtual announcer read the generated script to be read, means for providing news to the user in a live streaming format, means for analyzing the user's emotions, and means for dynamically adjusting the virtual announcer's facial expressions and tone of voice based on the analyzed emotions. This enables a highly personalized news experience that corresponds to the user's emotions.
[1527] "Means for users to select news genres they like" refers to interfaces and functions that allow users to select news genres that interest them.
[1528] "Means for users to customize their virtual announcers" refers to interfaces and functions that allow users to set the announcer's appearance, voice, personality, etc.
[1529] "Means for acquiring news data and filtering it based on the genre selected by the user" is a function for collecting data from news sources and extracting news according to the genre selected by the user.
[1530] The "means for generating a script to be read aloud from filtered news data" is a function that automatically creates a script to be read aloud by a virtual announcer based on extracted news data.
[1531] The "means for having a virtual announcer read out the generated script" is a function for playing back the generated script in the voice of a virtual announcer.
[1532] The "means for providing news to users in a live distribution format" is a function for distributing news content to users in real time.
[1533] "Means for analyzing user emotions" refers to technology that analyzes a user's facial expressions and tone of voice in real time to recognize their emotional state.
[1534] "Means for dynamically adjusting the facial expressions and tone of voice of a virtual announcer based on analyzed emotions" refers to technology that changes the facial expressions and voice of a virtual announcer in real time based on the results of emotion analysis.
[1535] This invention is a system that achieves even higher levels of personalization by combining a news delivery system with a user-selected news genre and a customized virtual announcer with an emotion engine that recognizes the user's emotions. This system recognizes the user's emotions in real time using a smartphone, and based on that information, adjusts the virtual announcer's facial expressions and tone of voice, and recommends related news content.
[1536] System configuration
[1537] User terminal
[1538] Users install the "Emotion-Responsive News" application on their smartphones, create an account, and then log in. After logging in, they can select and customize the news genre and virtual announcer on the settings screen. The smartphones are also equipped with cameras and microphones, which are used to recognize the user's emotions in real time.
[1539] server
[1540] The server has the following main functions:
[1541] 1. News data collection and filtering
[1542] The server retrieves the latest news data from the news API and filters it based on the user's selected genre.
[1543] 2. Generation of reading script
[1544] The server uses a generative AI model to automatically generate a reading script from filtered news articles.
[1545] 3. Creating a Virtual Announcer
[1546] The server retrieves information about the virtual announcer set by the user, combines it with the script to be read, and generates the final video material using a voice synthesizer and 3D modeling.
[1547] 4. Emotion Recognition by Emotion Engine
[1548] The server receives the user's camera footage and audio data sent from the smartphone and analyzes the user's emotions in real time. The analysis results are classified into emotional states such as smile, surprise, sadness, etc.
[1549] 5. Live streaming and dynamic adjustments
[1550] The server dynamically adjusts the virtual announcer's facial expressions and tone of voice during live broadcasts based on the analysis results of the emotion engine, and can also recommend relevant news content based on the user's emotional state.
[1551] 6. Collecting User Feedback
[1552] After watching the news, feedback is collected from users to help improve the service. The feedback information is stored on the server and used for data analysis.
[1553] Specific examples
[1554] If the user selects sports news
[1555] 1. Users
[1556] Users install the "Emotionally Responsive News" app on their smartphone, create an account, and log in. On the settings screen, they select the "Sports" genre and then set their preferred virtual announcer (e.g., a man with a lively personality).
[1557] 2. Terminal
[1558] The selection information is stored on the local device and sent to the server. In addition, facial expression and voice data of the user are collected in real time via the smartphone's camera and microphone and sent to the server.
[1559] 3. Server
[1560] The server collects sports-related news data and filters it through a news API. It uses a generative AI model to automatically generate a script to read from the news data. It then obtains information about the virtual announcer selected by the user, combines it with the script, and generates the final video material using a voice synthesizer and 3D modeling.
[1561] 4. Emotion Engine
[1562] The server analyzes the video and audio transmitted from the user's smartphone to identify the user's emotional state in real time, and adaptively adjusts the virtual announcer's facial expressions and tone of voice based on the analysis results.
[1563] 5. Live Streaming
[1564] The server uploads the generated video material to a live streaming server and provides a stream URL to the device. The smartphone uses this URL to play the video in real time, and a virtual announcer announces, "I'll tell you the results of last night's game."
[1565] 6. Users
[1566] Users can open the app during their commute and watch a virtual announcer speak. While watching, the virtual announcer changes its facial expression and tone of voice depending on the user's emotions. It also provides relevant news based on the user's emotions.
[1567] Prompt Sentence Examples
[1568] Leveraging a generative AI model, we use prompts like:
[1569] Please create a news script to use in your presentation. It should be based on recent news in the sports genre and should include the following:
[1570] Last night's match results
[1571] Performances of noteworthy players
[1572] Upcoming match schedule
[1573] Please keep it easy to read and in a friendly tone."
[1574] This allows users to enjoy a highly personalized news experience that responds to their emotions in real time.
[1575] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1576] Step 1:
[1577] A user installs the "Emotion Response News" app on their smartphone, creates an account, and logs in. As input, the user provides the necessary registration information, and as output, an account is created. Specifically, information such as a username, password, and email address is entered, and the user's account data is generated based on this.
[1578] Step 2:
[1579] Users select and customize their news genre and virtual announcer on the app's settings screen. The input is the user's genre selection and announcer customization information, and the output is saved to the local device. Specifically, this includes genre selections such as "sports" and "entertainment," as well as the announcer's appearance, voice tone, and personality settings.
[1580] Step 3:
[1581] The terminal sends the setting information of the news genre and virtual announcer to the server. The user's selection information is provided as input, and the server receives this information as output and stores it in a database. Specifically, the setting information is sent to the server as JSON format data.
[1582] Step 4:
[1583] The server retrieves the latest news data from the news API and filters the data based on the genre selected by the user. As input, news data from the news API is provided, and as output, news data narrowed down to a genre is generated. Specifically, the server retrieves news data using an API key and filters the news according to the specified genre.
[1584] Step 5:
[1585] The server uses a generative AI model to automatically generate a reading script from the filtered news articles. The filtered news data is provided as input, and text data for reading is generated as output. Specifically, a prompt sentence is input to the generative AI model (such as OpenAI GPT) and the generated text is obtained.
[1586] Step 6:
[1587] The server combines the generated script with the virtual announcer information set by the user and generates the final video material using a voice synthesizer and 3D modeling. The script and announcer settings are provided as input, and video material is generated as output. Specifically, the voice synthesizer generates audio data and links it to the 3D model to create a video.
[1588] Step 7:
[1589] The device uses the smartphone's camera and microphone to collect the user's facial expressions and voice data in real time and transmits it to a server. The user's real-time video and audio are provided as input, and emotion analysis data is generated as output. Specifically, emotion analysis is performed using the Microsoft Azure Face API and Google Cloud Vision API.
[1590] Step 8:
[1591] The server analyzes the user's emotional data provided in real time and dynamically adjusts the virtual announcer's facial expressions and tone of voice. The emotion analysis results are provided as input, and the announcer's facial expressions and tone of voice are adjusted as output. Specifically, the synthesizer parameters are adjusted based on the emotional data, changing the facial expressions of the 3D model.
[1592] Step 9:
[1593] The server uploads the generated video material to the distribution server in real time and provides a stream URL to the terminal. The completed video material is provided as input, and a stream URL is generated as output. Specifically, the video upload process to the distribution server is executed.
[1594] Step 10:
[1595] The device plays the video in real time based on the stream URL and provides news to the user. The stream URL is provided as input and the video is played as output. Specifically, the video player receives the stream and plays it.
[1596] Step 11:
[1597] After watching the news, the user provides feedback, which the device sends to the server. The user's feedback is provided as input, and the server receives and stores the feedback data as output. Specifically, the user taps the feedback icon on the screen and enters their opinion or impression.
[1598] Step 12:
[1599] The server uses the collected feedback data for data analysis to help improve the service. User feedback data is provided as input, and improvement measures are formulated as output. Specifically, the feedback data is analyzed using a data analysis tool, and areas for improvement are listed.
[1600] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1601] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1602] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1603] [Fourth embodiment]
[1604] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1605] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1606] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1607] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1608] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1609] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1610] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1611] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1612] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1613] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1614] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1615] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1616] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1617] ---
[1618] This invention is a system for providing viewers with a personalized news experience, allowing users to select a news genre and a virtual announcer and watch customized live news. This system has functions that allow users to select news genres and virtual announcers according to their preferences, and functions that automatically generate a text-to-speech script from news data and have the virtual announcer read it aloud. It also has a function for improving the service based on user feedback.
[1619] System configuration
[1620] User terminal
[1621] Users install the news distribution application on their device, create an account, and then log in. Once successfully logged in, users can select news genres and virtual announcers. The user's selection information is saved on the device and sent to the server.
[1622] server
[1623] The server has the following main functions:
[1624] 1. News data collection:
[1625] The server retrieves the latest news data from various news APIs. This data is aggregated from multiple news sources.
[1626] 2. Filtering and generating a transcript:
[1627] The collected news data is filtered based on the news genre selected by the user, and then a generative AI model is used to automatically generate a reading transcript from the filtered news articles.
[1628] 3. Creating a virtual announcer:
[1629] The virtual announcer reads the script based on the user-customized virtual announcer settings (voice, appearance, etc.), and then generates the final video material for live streaming.
[1630] 4. Live Streaming:
[1631] The generated virtual announcer will broadcast live news videos in real time over the Internet.
[1632] 5. Collecting User Feedback:
[1633] After watching the news, we collect feedback from users to help improve the service. The feedback information is stored on the server and can be used for data analysis.
[1634] Specific examples
[1635] If the user selects sports news
[1636] 1. User:
[1637] Install the news distribution app and create an account. After logging in, select the "Sports" genre on the settings screen and then set your preferred virtual announcer (e.g., a male voice with an active personality).
[1638] 2. Terminal:
[1639] The selection information is stored on the local device and transmitted to the server.
[1640] 3. Server:
[1641] Sports-related news data is collected and filtered. Then, using a generative AI model, a script is automatically generated from the filtered news data. The final video material is generated based on the script and the settings of the virtual announcer selected by the user.
[1642] 4. Terminal:
[1643] The user's device receives the streaming URL from the server and plays the sports news in real time.
[1644] 5. User:
[1645] On the way to work, users can open the app and watch a virtual announcer announce the results of last night's game. After watching, users can provide feedback within the app.
[1646] This system allows for a personalized news experience, increasing viewer engagement.
[1647] The processing flow will be explained below.
[1648] ---
[1649] Step 1:
[1650] Users install the news delivery application on their device and create an account. After creating the account, they log in by entering their email address and password on the login page.
[1651] Step 2:
[1652] The terminal sends the user's login information to the server for authentication. If authentication is successful, the server issues a session ID and returns it to the terminal. The terminal saves the session ID and maintains the logged-in state.
[1653] Step 3:
[1654] Users access the app's settings screen, select their preferred news genre (e.g., sports, technology, politics, etc.), customize their preferred virtual announcer (e.g., gender, tone of voice, appearance, etc.), and click the "Save" button.
[1655] Step 4:
[1656] The terminal transmits the user's selected news genre and virtual announcer setting information to the server, which then stores the received customization information in the user database.
[1657] Step 5:
[1658] The server retrieves the latest news data from multiple news APIs, filters the retrieved news data based on the genre selected by the user, and extracts the corresponding news data.
[1659] Step 6:
[1660] The server uses a generative AI model to automatically generate a speech transcript from the filtered news data, which is then stored in a database.
[1661] Step 7:
[1662] The server retrieves the information of the virtual announcer selected by the user and combines it with the script to generate the final video material. Specifically, it generates the reading voice using a voice synthesizer and creates the character movements using 3D modeling.
[1663] Step 8:
[1664] The server uploads the generated news video of the virtual announcer to the live streaming server and generates a stream URL, which is then sent to the device.
[1665] Step 9:
[1666] The device plays the video in real time based on the stream URL received from the server, and users can open the app at their preferred time to watch customized live news.
[1667] Step 10:
[1668] After watching the news, users can rate the announcer's performance and the content of the news in the feedback form within the app and enter the feedback information on their device, which then sends the feedback information to the server.
[1669] Step 11:
[1670] The server stores user feedback information in a database and uses it for data analysis to help improve generative AI models and services.
[1671] This completes the system's processing flow, allowing users to enjoy a personalized news experience in real time.
[1672] Example 1
[1673] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1674] Conventional news delivery systems make it difficult for viewers to select news that suits their preferences and watch it in a customized format. Furthermore, they lacked a mechanism for effectively collecting viewer feedback and utilizing it to improve services. As a result, they were unable to fully optimize the viewing experience or improve viewer satisfaction.
[1675] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1676] In this invention, the server includes a means for acquiring news data and filtering it based on a genre selected by the user, a means for generating a script from the filtered news data using a generative AI model, and a means for having a virtual announcer read the generated script. This allows users to select news from their favorite genres and watch live news broadcasts by customized virtual announcers. Furthermore, by collecting user feedback and using it to improve the service, viewer satisfaction can be increased.
[1677] 1. "A means for users to select news genres they like" refers to a function that allows users to select news categories that interest them.
[1678] 2. "Means for users to customize the virtual announcer" refers to the functionality that allows users to change settings such as the voice and appearance of the virtual announcer.
[1679] 3. "Means of acquiring news data" refers to the function of collecting the latest news information from external news data providers.
[1680] 4. "Means of filtering based on the genre selected by the user" refers to the function of extracting only information that falls into the category selected by the user from the collected news data.
[1681] 5. "Means for generating a text-to-speech transcript from filtered news data" refers to a function that uses a generative AI model to summarize filtered news information and create text that can be read aloud.
[1682] 6. "Means for having a virtual announcer read the generated script" refers to the function of converting text data into voice using speech synthesis technology and synchronizing it with the actions of the virtual announcer.
[1683] 7. "Means for providing users with generated news videos in a live streaming format" refers to the function of delivering generated audio and video to users in real time via the Internet.
[1684] 8. "Means of collecting user feedback" refers to the function of collecting ratings and comments from viewers and storing them in a database.
[1685] 9. "Means used to improve services" refers to the functions used to analyze collected feedback and improve the quality of news provision services.
[1686] 10. "Generative AI model" refers to an artificial intelligence model that generates text data using natural language processing technology.
[1687] 11. “Prompt” refers to a command or question input to a generative AI model.
[1688] These are the definitions of the important words included in this system.
[1689] The present invention is a system for providing users with a personalized news experience, allowing them to select a news genre and a virtual announcer and watch customized live news. This system has functions that allow users to select news genres and virtual announcers according to their preferences, and functions that automatically generate a text-to-speech script from news data and have the virtual announcer read it aloud. It also has a function to improve the service based on user feedback.
[1690] System configuration
[1691] User terminal
[1692] Users install the news distribution application on their device, create an account, and then log in. Once successfully logged in, users can select news genres and virtual announcers. The user's selection information is saved on the device and sent to the server. This can be done using common mobile devices such as smartphones and tablets.
[1693] server
[1694] The server has the following main functions:
[1695] 1. News data collection:
[1696] The server retrieves the latest news data from various news providers. This data is collected from multiple news sources. Specifically, "News API" and "News Data Provider Services" are used.
[1697] 2. Filtering and generating a transcript:
[1698] The collected news data is filtered based on the news genre selected by the user. Then, a generative AI model is used to automatically generate a speech transcript from the filtered news articles. The generative AI model used here is, for example, GPT-4.
[1699] 3. Creating a virtual announcer:
[1700] The virtual announcer reads the script based on the user's customized settings (voice, appearance, etc.). The final video material for live streaming is then generated. Speech synthesis software (e.g., TTS technology) is used for the speech synthesis, and 3D animation software (e.g., animation generation software) is used to generate the video.
[1701] 4. Live Streaming:
[1702] The generated news video of the virtual announcer is broadcast live. The live broadcast is carried out in real time over the Internet. A "live streaming service," for example, is used as the streaming platform.
[1703] 5. Collecting User Feedback:
[1704] After watching the news, we collect feedback from users to help improve the service. The feedback information is stored on the server and can be used for data analysis.
[1705] Specific examples
[1706] If the user selects sports news
[1707] 1. User:
[1708] Install the news distribution app and create an account. After logging in, select the "Sports" genre on the settings screen and then set your preferred virtual announcer (e.g., a male voice with an active personality).
[1709] 2. Terminal:
[1710] The selection information is stored on the local device and transmitted to the server.
[1711] 3. Server:
[1712] Sports-related news data is collected and filtered. Then, using a generative AI model, a script is automatically generated from the filtered news data. The final video material is generated based on the script and the settings of the virtual announcer selected by the user.
[1713] 4. Terminal:
[1714] The user's device receives the streaming URL from the server and plays the sports news in real time.
[1715] 5. User:
[1716] On the way to work, users can open the app and watch a virtual announcer announce the results of last night's game. After watching, users can provide feedback within the app.
[1717] Prompt Sentence Examples
[1718] "I want to watch sports news. I want the virtual announcer to be a lively male."
[1719] This system will enable viewers to enjoy a news experience that is optimized for each individual, increasing engagement.
[1720] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1721] Step 1:
[1722] A user installs a news delivery app and creates an account
[1723] Input: User's electronic device (smartphone, tablet) and internet connection
[1724] How it works: A user downloads and installs a news delivery app from the app store. They launch the app and create an account by entering their email address and password on the account creation screen that appears when the app is first launched.
[1725] Output: Account information is created and you can log in to the app. After logging in, you will be taken to a screen where you can select the news genre and virtual announcer.
[1726] Step 2:
[1727] The user selects the news genre and virtual announcer.
[1728] Input: User preference information (selected news genre, virtual announcer settings)
[1729] How it works: Within the app, users select "Sports" from a range of news genres, including "Sports," "Politics," and "Entertainment." They then customize the virtual announcer's settings screen, such as selecting a "lively male voice."
[1730] Output: The selection information is temporarily stored in the device's local storage and then sent to the server.
[1731] Step 3:
[1732] The server collects news data
[1733] Input: API requests and responses received by the server from external news providers
[1734] How it works: The server sends a request to an API endpoint such as a "news data provider service" to retrieve the latest news data in JSON format.
[1735] Output: The retrieved news data is stored in a database for further processing.
[1736] Step 4:
[1737] The server filters the news data and generates a reading script.
[1738] Input: News data to be filtered and user genre selection information
[1739] How it works: The server uses a filtering algorithm (e.g., TF-IDF) to extract only news data that falls within the user's selected "sports" genre. It then uses a generative AI model (e.g., GPT-4) to generate a summary of the filtered news data.
[1740] Output: The filtered news data and the generated speech transcripts are stored in a database.
[1741] Step 5:
[1742] The server configures the virtual announcer and generates the news video.
[1743] Input: Reading script and user's virtual announcer setting information
[1744] How it works: The server uses speech synthesis software (e.g., TTS technology) to convert the generated script into audio data, and then uses 3D animation software (e.g., animation generation software) to generate a video that coordinates the voice and movements of the virtual announcer.
[1745] Output: The created news video is stored on the server and ready to be streamed.
[1746] Step 6:
[1747] Server-generated live news video broadcast
[1748] Input: Generated news video data
[1749] How it works: The server uses a live streaming service to create a streaming URL for the generated news video, and then sends this URL to the user's device.
[1750] Output: A response containing a streaming URL is sent to the user's device.
[1751] Step 7:
[1752] User watches news
[1753] Input: Streaming URL
[1754] How it works: The user's device receives the streaming URL and uses the media player component to play the news video in real time.
[1755] Output: User watches live streaming news video.
[1756] Step 8:
[1757] Users provide feedback
[1758] Input: User feedback information (ratings, comments, etc.)
[1759] How it works: After watching a news item, users enter their rating and comments about their viewing experience on the feedback screen displayed within the app. The feedback is then sent from the app to the server.
[1760] Output: The feedback information is sent to the server and stored in a database.
[1761] The above is a detailed description of the specific processing steps of the present system. The present invention is expected to provide users with an individually customized news experience, thereby improving viewer engagement and satisfaction.
[1762] (Application example 1)
[1763] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1764] In today's world, with so much news being distributed, it is difficult for individual users to accurately obtain only the news that interests them. There are also insufficient means to customize the visual and auditory experience of news for each individual user. Furthermore, there is a lack of mechanisms for users to provide feedback on the news they have viewed and use that feedback to improve services. Effective methods to solve these issues are needed.
[1765] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1766] In this invention, the server includes means for acquiring news data and filtering it based on a genre selected by the user, means for generating a script to be read from the filtered news data, and means for having a virtual announcer read the generated script to be read. This allows each user to watch news in the genres they are interested in in real time with a customized virtual announcer. Furthermore, it is possible to stream news to a smartphone application in real time and collect feedback from users to help improve the service.
[1767] "User" refers to an individual who uses this system to receive news distribution services.
[1768] "Genre" refers to the type or category of news, such as a specific field such as sports, economics, or entertainment.
[1769] A "virtual announcer" is a computer-generated character whose role is to read the news aloud and visually.
[1770] "News Data" refers to the collection of information collected from various news sources and refers to the material used for filtering and generating the readings.
[1771] "Filtering" refers to the process of narrowing down news data based on genres selected by the user.
[1772] A "script" is a text automatically generated from filtered news data, and refers to a script that a virtual announcer will read aloud.
[1773] "Live streaming" refers to a format in which news is provided to users in real time, and refers to a method of streaming directly from a server to a user's terminal.
[1774] "Smartphone application" refers to a program that runs on a smartphone and functions as the user interface for this system.
[1775] "Real-time streaming" refers to a technology in which news videos are sent to the user's device as soon as they are generated and played back instantly.
[1776] "Feedback" refers to the opinions and evaluations provided by users regarding the news and services they have viewed, and serves as important data for improving services.
[1777] A "generative AI model" refers to an artificial intelligence algorithm that uses machine learning technology to automatically generate a reading script from news data.
[1778] The present invention is a system that allows users to customize the news genres and virtual announcers that interest them, providing a personalized news experience.
[1779] System configuration
[1780] User terminal
[1781] Users install the corresponding application on their smartphone, create an account, and log in. After logging in, users can configure their news genre and virtual announcer settings. This includes selecting a genre (e.g., sports, economics, entertainment, etc.) and customizing the announcer's appearance and voice. This configuration information is saved on the user's device and sent to the server.
[1782] server
[1783] The server has the following main functions:
[1784] 1. News data collection:
[1785] The server uses news APIs to collect the latest news data from various news sources, examples of which include NewsAPI and Google News API.
[1786] 2. Filtering and generating a transcript:
[1787] The collected news data is filtered based on the genre selected by the user, and a transcription is automatically generated from the filtered news articles using a generative AI model (e.g., GPT-3).
[1788] 3. Creating a virtual announcer:
[1789] A virtual announcer reads the script according to the announcer settings customized by the user, using tools such as Adobe Character Animator and Unity3D. The generated news video is then streamed to the user's device in real time.
[1790] 4. Gathering Feedback:
[1791] After viewing the news, we collect feedback from users and use it to improve our services, thereby continuously improving the quality of news and its delivery methods.
[1792] Process Overview
[1793] 1. User Settings:
[1794] For example, a user selects "Entertainment News" and "An announcer with a female voice and a cheerful personality." The user's selection information is sent from the smartphone terminal to the server.
[1795] 2. News gathering and filtering:
[1796] The server uses NewsAPI to collect the latest entertainment news and filters the collected data.
[1797] 3. Generate a reading transcript:
[1798] A generative AI model (GPT-3) is used to generate a read-aloud transcript from filtered news articles. For example, a sentence such as "A blockbuster film was announced at the film festival yesterday" is generated.
[1799] 4. Creating a Virtual Announcer:
[1800] Using Adobe Character Animator and Unity3D, a user-customized virtual announcer is generated, and a video is created in which the announcer reads the script.
[1801] 5. Real-time Streaming and Viewing:
[1802] The generated news videos are hosted on a server and streamed in real time to smartphones, where users open the app and listen to the news being read by an announcer.
[1803] 6. Collecting Feedback and Improving Our Services:
[1804] After watching, users provide feedback that helps improve the service.
[1805] Examples and prompts
[1806] Examples:
[1807] A user opens the app and selects a sports news item and a calm, male-voiced announcer. The server collects the latest sports news and uses a generative AI model (GPT-3) to create a script, such as "I'll tell you the results of last night's game." Using Unity3D, a virtual announcer reads the script, and the generated video is delivered to a smartphone. The user can watch the news on their way to work and provide feedback.
[1808] Prompt for the generative AI model:
[1809] News Cache: {Sports}
[1810] Announcer Voice: {Male voice, calm}
[1811] The statement read: "Here are the results of last night's match."
[1812] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1813] Step 1:
[1814] A user creates an account and logs into the smartphone application. The user selects a news genre (e.g., sports, economics, entertainment) and virtual announcer settings (e.g., voice type, gender, appearance). The selection information is stored on the user's device and sent to the server.
[1815] Input: User's genre selection information and virtual announcer settings
[1816] Output: Selections sent to the server
[1817] Specific behavior:
[1818] The user opens the app and navigates to the settings screen.
[1819] Use the drop-down menus and sliders to set your preferred news genres and anchor appearance and voice.
[1820] After the setting information is confirmed, pressing the "Save" button saves the selected information in the local storage of the device and simultaneously sends it to the server.
[1821] Step 2:
[1822] The server uses a news API to collect the latest news data for the genre selected by the user. For example, it uses NewsAPI or Google News API.
[1823] Input: News API request
[1824] Output: Collected news data
[1825] Specific behavior:
[1826] The server accesses the news API periodically or upon user request to collect the latest news data for the selected genre.
[1827] The collected news data is stored in a database on the server.
[1828] Step 3:
[1829] The collected news data is filtered based on the genre selected by the user.
[1830] Input: Collected news data, user genre preferences
[1831] Output: Filtered news data
[1832] Specific behavior:
[1833] An algorithm on the server analyzes the collected news data and extracts only articles related to the genre selected by the user.
[1834] The filtered results are stored in a database.
[1835] Step 4:
[1836] A reading script is automatically generated from filtered news data using a generative AI model (e.g., GPT-3).
[1837] Input: Filtered news data
[1838] Output: Speech manuscript
[1839] Specific behavior:
[1840] The server provides a prompt to the generative AI model, which generates a reading script based on the filtered news data.
[1841] For example, the prompt text is:
[1842] News Cache: {Sports}
[1843] Announcer Voice: {Male voice, calm}
[1844] The statement read: "Here are the results of last night's match."
[1845] The generated reading manuscript is stored in a database on the server.
[1846] Step 5:
[1847] Based on the user's customized settings, a video is generated for the virtual announcer to read the script.
[1848] Input: Reading script, user's virtual announcer settings
[1849] Output: News video with a virtual announcer
[1850] Specific behavior:
[1851] Using Adobe Character Animator and Unity3D, a virtual announcer is generated based on the announcer settings selected by the user.
[1852] The script is read aloud by an announcer, and then generated in video format.
[1853] The generated video is stored on the server.
[1854] Step 6:
[1855] The generated news video of the virtual announcer is streamed in real time and provided to users.
[1856] Input: News video with a virtual announcer
[1857] Output: Real-time streaming video played on user devices
[1858] Specific behavior:
[1859] The server hosts the generated news video and generates a streaming URL.
[1860] The user's device accesses this URL and watches the news video in real time.
[1861] Step 7:
[1862] After watching the news, we collect feedback from users to help improve our services.
[1863] Input: User feedback
[1864] Output: Data for service improvement
[1865] Specific behavior:
[1866] After users watch a news video, they can provide their ratings and opinions through an in-app feedback form.
[1867] The server analyzes the collected feedback data and uses it to improve the service.
[1868] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1869] ---
[1870] This invention is a system that achieves even higher levels of personalization by combining a personalized news delivery system with a user-selected news genre and customized virtual announcer with an emotion engine that recognizes the user's emotions. This system recognizes the user's emotions in real time and, based on that information, adjusts the virtual announcer's facial expressions and tone of voice, and recommends related news content.
[1871] System configuration
[1872] User terminal
[1873] Users install the news distribution application on their device, create an account, and then log in. Once successfully logged in, they can select and customize the news genre and virtual announcer on the settings screen. The device is also equipped with a camera and microphone, which are used to recognize the user's emotions in real time.
[1874] server
[1875] The server has the following main functions:
[1876] 1. News data collection and filtering:
[1877] The server retrieves the latest news data from various news APIs and filters it based on the genre selected by the user.
[1878] 2. Generate a reading transcript:
[1879] The server uses a generative AI model to automatically generate a reading transcript from the filtered news articles.
[1880] 3. Creating a virtual announcer:
[1881] The server retrieves the information of the virtual announcer selected by the user and combines it with the script to generate the final video material. Specifically, it generates the reading voice using a voice synthesizer and creates the character movements using 3D modeling.
[1882] 4. Emotion Recognition with Emotion Engine:
[1883] The server receives the user's camera footage and audio data sent from the device and analyzes the user's emotions in real time. The analysis results are classified into emotional states such as smile, surprise, sadness, etc.
[1884] 5. Live streaming and dynamic adjustment:
[1885] The server dynamically adjusts the virtual announcer's facial expressions and tone of voice during live broadcasts based on the analysis results of the emotion engine, and can also recommend relevant news content based on the user's emotional state.
[1886] 6. Collecting User Feedback:
[1887] After watching the news, we collect feedback from users to help improve the service. The feedback information is stored on the server and can be used for data analysis.
[1888] Specific examples
[1889] If the user selects sports news
[1890] 1. User:
[1891] Install the news distribution app, create an account, and log in. On the settings screen, select the "Sports" genre and then set your preferred virtual announcer (e.g., a male voice with an active personality).
[1892] 2. Terminal:
[1893] The selection information is stored on the local device and sent to the server, and facial expression and voice data of the user are collected in real time via the camera and microphone and sent to the server.
[1894] 3. Server:
[1895] Collect and filter sports-related news data. Use a generative AI model to automatically generate a script to read from the news data. Obtain information about the virtual announcer selected by the user and combine it with the script to generate the final video material.
[1896] 4. Emotion Engine:
[1897] At the start, the system analyzes the user's emotions and provides the results to the server, which then adaptively adjusts the virtual announcer's facial expressions and tone of voice based on the analysis results.
[1898] 5. Live Streaming:
[1899] The server uploads the generated video material to the live streaming server and provides a stream URL to the device, which then plays the video in real time.
[1900] 6. Users:
[1901] Users can open the app during their commute and watch a virtual announcer talk about the results of last night's game. While watching, the virtual announcer changes its facial expression and tone of voice depending on the user's emotions. It also provides related news based on the user's emotions.
[1902] 7. Feedback:
[1903] After watching the news, users provide feedback within the app, which is then sent to the server and used to improve the service.
[1904] This system allows users to enjoy a highly personalized, emotionally-driven news experience in real time.
[1905] The processing flow will be explained below.
[1906] ---
[1907] Step 1:
[1908] Users install the news delivery application on their device and create an account. After creating the account, they log in by entering their email address and password on the login page.
[1909] Step 2:
[1910] The terminal sends the user's login information to the server for authentication. If authentication is successful, the server issues a session ID and returns it to the terminal. The terminal saves the session ID and maintains the logged-in state.
[1911] Step 3:
[1912] Users access the app's settings screen, select their preferred news genre (e.g., sports, technology, politics, etc.), customize their preferred virtual announcer (e.g., gender, tone of voice, appearance, etc.), and click the "Save" button.
[1913] Step 4:
[1914] The terminal transmits the user's selected news genre and virtual announcer setting information to the server, which then stores the received customization information in the user database.
[1915] Step 5:
[1916] The server retrieves the latest news data from multiple news APIs, filters the retrieved news data based on the genre selected by the user, and extracts the corresponding news data.
[1917] Step 6:
[1918] The server uses a generative AI model to automatically generate a speech transcript from the filtered news data, which is then stored in a database.
[1919] Step 7:
[1920] The server retrieves the information of the virtual announcer selected by the user and combines it with the script to generate the final video material. Specifically, it generates the reading voice using a voice synthesizer and creates the character movements using 3D modeling.
[1921] Step 8:
[1922] The device uses a camera and microphone to collect the user's facial expressions and voice data in real time and transmits it to a server.
[1923] Step 9:
[1924] The server uses an emotion engine to analyze the user's facial expression and audio data sent from the device and recognize the user's emotional state. The recognition results are classified as emotions such as smile, surprise, sadness, etc.
[1925] Step 10:
[1926] Based on the analysis results of the emotion engine, the server dynamically adjusts the virtual announcer's facial expressions and tone of voice during live broadcasts, enabling a more natural and approachable news delivery that responds to the user's emotions.
[1927] Step 11:
[1928] The server uploads the generated news video of the virtual announcer to the live streaming server and sends the stream URL to the terminal.
[1929] Step 12:
[1930] The device uses the stream URL received from the server to play the video in real time, and the user opens the app to watch the virtual announcer change his or her facial expressions and tone of voice according to the user's emotions.
[1931] Step 13:
[1932] During live streaming, the server dynamically recommends relevant news content based on the user's emotional state, for example, if the user expresses surprise, it will display the latest news related to surprise.
[1933] Step 14:
[1934] After watching the news, the user inputs feedback information using a feedback form, which is then sent to the server by the terminal.
[1935] Step 15:
[1936] The server stores user feedback information in a database and uses it for data analysis to improve services and optimize generative AI models.
[1937] This concludes the specific processing flow of the system, which enables users to enjoy a personalized, emotion-sensitive news experience in real time.
[1938] Example 2
[1939] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1940] Current digital news delivery systems only provide one-way information to users, limiting personalization based on user emotions and preferences. Furthermore, passive news consumption makes it difficult to provide an experience tailored to users' interests and engagement levels. Therefore, there is a need for a news delivery system that can capture users' interest and adjust in real time to reflect their emotions.
[1941] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for selecting information of a genre that the user likes, means for the user to customize a virtual character, means for acquiring information data and filtering it based on the genre selected by the user, means for generating a text-to-speech script from the filtered information data, means for having the virtual character read the generated text-to-speech script, means for recognizing the user's emotions in real time using the camera and microphone of the terminal and transmitting the data, means for analyzing the emotions and dynamically adjusting the virtual character's facial expression and tone of voice based on the analysis results, and means for providing information to the user in a live streaming format. This enables highly personalized news delivery according to the user's emotions.
[1942] "User" refers to a general user who receives information and performs settings and customization.
[1943] A "virtual character" is a digital avatar that provides information to users and has audio and visual representations.
[1944] "Information data" refers to digital information such as text, audio, and video that is acquired based on the genre selected by the user.
[1945] "Filtering" refers to the process of selecting only relevant data from the acquired information data based on user selections.
[1946] A "reading manuscript" refers to text that is analyzed and generated by a generative AI model based on information data and is read aloud by a virtual character.
[1947] A "generative AI model" refers to an artificial intelligence model that generates and analyzes text based on large amounts of data.
[1948] "Emotion recognition" refers to a technology that uses the device's camera and microphone to analyze emotions from the user's facial expressions and voice.
[1949] "Dynamic adjustment" refers to the process of changing a virtual character's facial expression or tone of voice based on the results of a user's emotional analysis in real time.
[1950] "Live streaming" refers to a method of providing generated video material to users in real time.
[1951] "Feedback" refers to opinions and impressions provided by users after viewing.
[1952] "Customization" refers to the process by which a user sets the gender, tone of voice, character style, etc. of a virtual character according to their own preferences.
[1953] This invention is a system that allows users to select information from their favorite genres and delivers personalized information via a customized virtual character. This system achieves even higher levels of personalization by recognizing the user's emotions in real time and dynamically adjusting the virtual character's facial expressions and tone of voice based on that information.
[1954] System configuration
[1955] User
[1956] 1. The user installs a news delivery application on their device (smartphone, tablet, PC, etc.), creates an account, and logs in.
[1957] 2. After successfully logging in, you can select and customize your news genre (e.g., "Sports," "Politics," "Technology," etc.) and virtual character (gender, voice tone, character style, etc.) on the settings screen.
[1958] Terminal
[1959] 1. The device is equipped with a camera and microphone, which are used to recognize the user's emotions (facial expressions and tone of voice) in real time.
[1960] 2. The user's selected settings information (news genre, virtual character customization) is saved in a local database and sent to the server.
[1961] server
[1962] 1. News data collection and filtering: The server retrieves the latest news data from various news APIs (e.g., NewsAPI, Google News API) and filters it based on the genre selected by the user.
[1963] 2. Generate a prompt: The server uses a generative AI model (e.g., OpenAI GPT-3) to automatically generate a prompt from the filtered news articles.
[1964] Sample prompt: "Create a script based on the latest sports news article to be read aloud by a fictional character who is an active and energetic man."
[1965] 3. Virtual character generation: The server acquires the virtual character information (e.g., character model, voice tone) set by the user and combines it with the reading script to generate the final video material. Specifically, the reading voice is generated using a voice synthesizer (e.g., Amazon Polly), and the virtual character's movements are created using 3D modeling software (e.g., Blender).
[1966] 4. Emotion recognition using an emotion engine: The server receives the user's camera footage and audio data sent from the device and analyzes the user's emotions using an emotion recognition engine (e.g., Affectiva SDK). Based on the analysis results (e.g., smile, surprise, sadness), the virtual character's facial expressions and tone of voice are dynamically adjusted.
[1967] 5. Live streaming: The final video material is uploaded to the live streaming server, and a stream URL is generated and sent to the device.
[1968] Specific examples
[1969] 1. User: Install the news distribution app, create an account, and log in. Select the "Sports" genre on the settings screen and set your preferred virtual character (e.g., a male voice with an active personality).
[1970] 2. Device: The device stores the selection information locally and sends it to the server. It also collects the user's facial expressions and voice data in real time via the camera and microphone and sends them to the server.
[1971] 3. Server: Collects and filters sports-related news data. Automatically generates a script to read from the news data using a generative AI model. Information about the virtual character set by the user is acquired, and combined with the script to generate the final video material.
[1972] 4. Emotion Engine: At the start, the system analyzes the user's emotions and provides the results to the server, which then adaptively adjusts the virtual character's facial expressions and tone of voice based on the analysis results.
[1973] 5. Live streaming: The server uploads the generated video material to the live streaming server and provides a stream URL to the device, which then plays the video in real time.
[1974] 6. User: Watch the news and hear a virtual character say, "I'll tell you the results of last night's game." While watching, the virtual character changes its facial expression and tone of voice according to the user's emotions. Related news is also provided.
[1975] 7. Feedback: After watching the news, users can provide feedback within the app, which will be sent to the server and used to improve the service.
[1976] This system allows users to enjoy a highly personalized, emotionally-driven news experience in real time.
[1977] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1978] Step 1: Initial System Setup and Login
[1979] operation:
[1980] Users install the news distribution application on their device, create an account, and then log in.
[1981] input:
[1982] Your account information (email address, password, social media accounts, etc.).
[1983] output:
[1984] Session information indicating that the user logged in successfully.
[1985] Specific behavior:
[1986] A user launches the app and enters their account information on the login screen. The entered information is sent to the server, where authentication is performed. If successful, session information is returned to the device.
[1987] Step 2: Setting the news genre and virtual characters
[1988] operation:
[1989] Users select a news genre on the settings screen and customize their preferred virtual character.
[1990] input:
[1991] Customization information for the news genre and virtual character selected by the user.
[1992] output:
[1993] The setting information is stored on the device and sent to the server.
[1994] Specific behavior:
[1995] Users select a news genre (e.g., sports, politics, technology, etc.) on the app's settings screen and customize the attributes of their virtual character (e.g., gender, tone of voice, appearance, etc.). The selection information is stored in a local database and transmitted to a server over the network.
[1996] Step 3: Collecting and filtering news data
[1997] operation:
[1998] The server retrieves the latest news data from the news API and filters it based on the genre selected by the user.
[1999] input:
[2000] News data obtained from the news API, news genres selected by the user.
[2001] output:
[2002] News data filtered to user-selected genres.
[2003] Specific behavior:
[2004] The server sends a request to a news API (e.g., NewsAPI, Google News API) to retrieve the latest news data, which is then filtered by keywords and categories based on the user's selection to select only relevant news.
[2005] Step 4: Generate a transcript
[2006] operation:
[2007] The server uses a generative AI model to automatically generate a reading script from filtered news data.
[2008] input:
[2009] Filtered news data, generative AI models (e.g., OpenAI GPT-3).
[2010] output:
[2011] The generated reading transcript.
[2012] Specific behavior:
[2013] The server inputs the filtered news data into a generative AI model and generates a reading script using a prompt, such as "Please create a reading script for a lively and energetic male virtual character based on the latest sports news article." The generated script is then stored in a database.
[2014] Step 5: Generate a virtual character
[2015] operation:
[2016] The server combines the information from the virtual character with the reading script to generate the final video material.
[2017] input:
[2018] Virtual character customization information, generated reading script.
[2019] output:
[2020] A video clip of a virtual character reading the news.
[2021] Specific behavior:
[2022] The server retrieves the virtual character's customization information, generates the reading voice using a voice synthesizer (e.g., Amazon Polly), and generates the virtual character's movements and facial expressions using 3D modeling software (e.g., Blender). These are then combined to generate the final video material.
[2023] Step 6: Emotion recognition and dynamic regulation
[2024] operation:
[2025] The device's camera and microphone are used to collect the user's emotions in real time and send the data to a server.
[2026] input:
[2027] User's facial expression data, voice data.
[2028] output:
[2029] User sentiment analysis results.
[2030] Specific behavior:
[2031] The device uses a camera and microphone to collect the user's facial expressions and tone of voice in real time, and this data is sent to a server. An emotion recognition engine (e.g., Affectiva SDK) on the server analyzes this data and classifies emotional states such as smile, sadness, or surprise. Based on the analysis results, the virtual character's facial expressions and tone of voice are dynamically adjusted.
[2032] Step 7: Live Stream
[2033] operation:
[2034] The server uploads the generated video material to a live streaming server, generates a stream URL, and sends it to the terminal.
[2035] input:
[2036] The final video footage produced.
[2037] output:
[2038] The stream URL for live streaming.
[2039] Specific behavior:
[2040] The server uploads the video material to the live streaming server and generates a stream URL. This URL is sent to the device and can be accessed by the user. The device uses this URL to play the video in real time and displays a virtual character reading the news.
[2041] Step 8: Gather feedback
[2042] operation:
[2043] Users provide feedback after watching the news.
[2044] input:
[2045] User feedback information.
[2046] output:
[2047] Feedback information collected.
[2048] Specific behavior:
[2049] Users enter their opinions and thoughts into the feedback form within the app. This feedback information is collected by the device and sent to the server, which stores it in a database and uses it to improve the service in the future.
[2050] (Application example 2)
[2051] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[2052] Conventional news delivery systems have limited personalization based on user emotions and provide a uniform news viewing experience, making it difficult to increase user satisfaction. Furthermore, they lack interactivity in news delivery because they are unable to reflect users' real-time reactions.
[2053] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for selecting a news genre that the user likes, means for the user to customize the virtual announcer, means for acquiring news data and filtering it based on the genre selected by the user, means for generating a script to be read from the filtered news data, means for having the virtual announcer read the generated script to be read, means for providing news to the user in a live streaming format, means for analyzing the user's emotions, and means for dynamically adjusting the virtual announcer's facial expressions and tone of voice based on the analyzed emotions. This enables a highly personalized news experience that corresponds to the user's emotions.
[2054] "Means for users to select news genres they like" refers to interfaces and functions that allow users to select news genres that interest them.
[2055] "Means for users to customize their virtual announcers" refers to interfaces and functions that allow users to set the announcer's appearance, voice, personality, etc.
[2056] "Means for acquiring news data and filtering it based on the genre selected by the user" is a function for collecting data from news sources and extracting news according to the genre selected by the user.
[2057] The "means for generating a script to be read aloud from filtered news data" is a function that automatically creates a script to be read aloud by a virtual announcer based on extracted news data.
[2058] The "means for having a virtual announcer read out the generated script" is a function for playing back the generated script in the voice of a virtual announcer.
[2059] The "means for providing news to users in a live distribution format" is a function for distributing news content to users in real time.
[2060] "Means for analyzing user emotions" refers to technology that analyzes a user's facial expressions and tone of voice in real time to recognize their emotional state.
[2061] "Means for dynamically adjusting the facial expressions and tone of voice of a virtual announcer based on analyzed emotions" refers to technology that changes the facial expressions and voice of a virtual announcer in real time based on the results of emotion analysis.
[2062] This invention is a system that achieves even higher levels of personalization by combining a news delivery system with a user-selected news genre and a customized virtual announcer with an emotion engine that recognizes the user's emotions. This system recognizes the user's emotions in real time using a smartphone, and based on that information, adjusts the virtual announcer's facial expressions and tone of voice, and recommends related news content.
[2063] System configuration
[2064] User terminal
[2065] Users install the "Emotion-Responsive News" application on their smartphones, create an account, and then log in. After logging in, they can select and customize the news genre and virtual announcer on the settings screen. The smartphones are also equipped with cameras and microphones, which are used to recognize the user's emotions in real time.
[2066] server
[2067] The server has the following main functions:
[2068] 1. News data collection and filtering
[2069] The server retrieves the latest news data from the news API and filters it based on the user's selected genre.
[2070] 2. Generation of reading script
[2071] The server uses a generative AI model to automatically generate a reading script from filtered news articles.
[2072] 3. Creating a Virtual Announcer
[2073] The server retrieves information about the virtual announcer set by the user, combines it with the script to be read, and generates the final video material using a voice synthesizer and 3D modeling.
[2074] 4. Emotion Recognition by Emotion Engine
[2075] The server receives the user's camera footage and audio data sent from the smartphone and analyzes the user's emotions in real time. The analysis results are classified into emotional states such as smile, surprise, sadness, etc.
[2076] 5. Live streaming and dynamic adjustments
[2077] The server dynamically adjusts the virtual announcer's facial expressions and tone of voice during live broadcasts based on the analysis results of the emotion engine, and can also recommend relevant news content based on the user's emotional state.
[2078] 6. Collecting User Feedback
[2079] After watching the news, feedback is collected from users to help improve the service. The feedback information is stored on the server and used for data analysis.
[2080] Specific examples
[2081] If the user selects sports news
[2082] 1. Users
[2083] Users install the "Emotionally Responsive News" app on their smartphone, create an account, and log in. On the settings screen, they select the "Sports" genre and then set their preferred virtual announcer (e.g., a man with a lively personality).
[2084] 2. Terminal
[2085] The selection information is stored on the local device and sent to the server. In addition, facial expression and voice data of the user are collected in real time via the smartphone's camera and microphone and sent to the server.
[2086] 3. Server
[2087] The server collects sports-related news data and filters it through a news API. It uses a generative AI model to automatically generate a script to read from the news data. It then obtains information about the virtual announcer selected by the user, combines it with the script, and generates the final video material using a voice synthesizer and 3D modeling.
[2088] 4. Emotion Engine
[2089] The server analyzes the video and audio transmitted from the user's smartphone to identify the user's emotional state in real time, and adaptively adjusts the virtual announcer's facial expressions and tone of voice based on the analysis results.
[2090] 5. Live Streaming
[2091] The server uploads the generated video material to a live streaming server and provides a stream URL to the device. The smartphone uses this URL to play the video in real time, and a virtual announcer announces, "I'll tell you the results of last night's game."
[2092] 6. Users
[2093] Users can open the app during their commute and watch a virtual announcer speak. While watching, the virtual announcer changes its facial expression and tone of voice depending on the user's emotions. It also provides relevant news based on the user's emotions.
[2094] Prompt Sentence Examples
[2095] Leveraging a generative AI model, we use prompts like:
[2096] Please create a news script to use in your presentation. It should be based on recent news in the sports genre and should include the following:
[2097] Last night's match results
[2098] Performances of noteworthy players
[2099] Upcoming match schedule
[2100] Please keep it easy to read and in a friendly tone."
[2101] This allows users to enjoy a highly personalized news experience that responds to their emotions in real time.
[2102] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[2103] Step 1:
[2104] A user installs the "Emotion Response News" app on their smartphone, creates an account, and logs in. As input, the user provides the necessary registration information, and as output, an account is created. Specifically, information such as a username, password, and email address is entered, and the user's account data is generated based on this.
[2105] Step 2:
[2106] Users select and customize their news genre and virtual announcer on the app's settings screen. The input is the user's genre selection and announcer customization information, and the output is saved to the local device. Specifically, this includes genre selections such as "sports" and "entertainment," as well as the announcer's appearance, voice tone, and personality settings.
[2107] Step 3:
[2108] The terminal sends the setting information of the news genre and virtual announcer to the server. The user's selection information is provided as input, and the server receives this information as output and stores it in a database. Specifically, the setting information is sent to the server as JSON format data.
[2109] Step 4:
[2110] The server retrieves the latest news data from the news API and filters the data based on the genre selected by the user. As input, news data from the news API is provided, and as output, news data narrowed down to a genre is generated. Specifically, the server retrieves news data using an API key and filters the news according to the specified genre.
[2111] Step 5:
[2112] The server uses a generative AI model to automatically generate a reading script from the filtered news articles. The filtered news data is provided as input, and text data for reading is generated as output. Specifically, a prompt sentence is input to the generative AI model (such as OpenAI GPT) and the generated text is obtained.
[2113] Step 6:
[2114] The server combines the generated script with the virtual announcer information set by the user and generates the final video material using a voice synthesizer and 3D modeling. The script and announcer settings are provided as input, and video material is generated as output. Specifically, the voice synthesizer generates audio data and links it to the 3D model to create a video.
[2115] Step 7:
[2116] The device uses the smartphone's camera and microphone to collect the user's facial expressions and voice data in real time and transmits it to a server. The user's real-time video and audio are provided as input, and emotion analysis data is generated as output. Specifically, emotion analysis is performed using the Microsoft Azure Face API and Google Cloud Vision API.
[2117] Step 8:
[2118] The server analyzes the user's emotional data provided in real time and dynamically adjusts the virtual announcer's facial expressions and tone of voice. The emotion analysis results are provided as input, and the announcer's facial expressions and tone of voice are adjusted as output. Specifically, the synthesizer parameters are adjusted based on the emotional data, changing the facial expressions of the 3D model.
[2119] Step 9:
[2120] The server uploads the generated video material to the distribution server in real time and provides a stream URL to the terminal. The completed video material is provided as input, and a stream URL is generated as output. Specifically, the video upload process to the distribution server is executed.
[2121] Step 10:
[2122] The device plays the video in real time based on the stream URL and provides news to the user. The stream URL is provided as input and the video is played as output. Specifically, the video player receives the stream and plays it.
[2123] Step 11:
[2124] After watching the news, the user provides feedback, which the device sends to the server. The user's feedback is provided as input, and the server receives and stores the feedback data as output. Specifically, the user taps the feedback icon on the screen and enters their opinion or impression.
[2125] Step 12:
[2126] The server uses the collected feedback data for data analysis to help improve the service. User feedback data is provided as input, and improvement measures are formulated as output. Specifically, the feedback data is analyzed using a data analysis tool, and areas for improvement are listed.
[2127] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[2128] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[2129] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[2130] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[2131] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[2132] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[2133] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[2134] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[2135] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[2136] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[2137] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[2138] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[2139] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[2140] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[2141] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[2142] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[2143] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[2144] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[2145] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[2146] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[2147] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[2148] The following is further disclosed regarding the above embodiment.
[2149] (Claim 1)
[2150] A means for the user to select news genres that they like;
[2151] a means for a user to customize the virtual announcer;
[2152] means for retrieving news data and filtering it based on a user-selected genre;
[2153] a means for generating a reading script from the filtered news data;
[2154] A means for having a virtual announcer read the generated reading script;
[2155] A means for providing news to users in a live streaming format;
[2156] A system including:
[2157] (Claim 2)
[2158] 10. The system of claim 1, further comprising means for collecting and utilizing user feedback to improve the service.
[2159] (Claim 3)
[2160] 10. The system of claim 1, further comprising means for automatically generating a reading script from news data using a generative AI model.
[2161] "Example 1"
[2162] (Claim 1)
[2163] A means for the user to select news genres that they like;
[2164] a means for a user to customize the virtual announcer;
[2165] means for retrieving news data and filtering it based on a user-selected genre;
[2166] a means for generating a reading script from the filtered news data;
[2167] A means for having a virtual announcer read the generated reading script;
[2168] a means for providing the generated news video to users in a live streaming format;
[2169] A system including:
[2170] (Claim 2)
[2171] 10. The system of claim 1, further comprising means for collecting and utilizing user feedback to improve the service.
[2172] (Claim 3)
[2173] 10. The system of claim 1, further comprising means for automatically generating a reading script from news data using a generative AI model.
[2174] "Application Example 1"
[2175] (Claim 1)
[2176] A means for the user to select news genres that they like;
[2177] a means for a user to customize the virtual announcer;
[2178] means for retrieving news data and filtering it based on a user-selected genre;
[2179] a means for generating a reading script from the filtered news data;
[2180] A means for having a virtual announcer read the generated reading script;
[2181] A means for providing news to users in a live streaming format;
[2182] a means of streaming news in real time to smartphone applications;
[2183] A system including:
[2184] (Claim 2)
[2185] 10. The system of claim 1, further comprising means for collecting and utilizing user feedback to improve the service.
[2186] (Claim 3)
[2187] 10. The system of claim 1, further comprising means for automatically generating a reading script from news data using a generative AI model.
[2188] "Example 2: Combining Emotion Engines"
[2189] (Claim 1)
[2190] A means for a user to select information of a genre that the user likes;
[2191] a means for a user to customize a virtual character;
[2192] means for acquiring informat...
Claims
1. A means for the user to select news genres that they like; a means for a user to customize the virtual announcer; means for retrieving news data and filtering it based on a user-selected genre; a means for generating a reading script from the filtered news data; A means for having a virtual announcer read the generated reading script; A means for providing news to users in a live streaming format; A system including:
2. The system of claim 1 further comprising means for collecting user feedback and using it to improve the service.
3. The system of claim 1 , further comprising means for automatically generating a reading script from news data using a generative AI model.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A