system

The system addresses the challenge of personalized content and advertising by analyzing user interests and emotions, converting content into audio/video formats, and optimizing delivery, resulting in enhanced user experiences and advertising effectiveness.

JP2026070939APending Publication Date: 2026-04-28SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-16
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Existing information delivery systems struggle to efficiently provide personalized digital content and targeted advertisements based on user interests and emotional states, leading to suboptimal user experiences and advertising effectiveness.

Method used

A system that utilizes a server to collect and analyze user interests and emotional data, converting selected digital content into audio and video formats using speech and video generation technologies, and delivering it to user terminals while optimizing future content and advertising strategies based on viewing data.

Benefits of technology

Enables personalized content delivery and effective advertising by tailoring information to users' interests and emotional states, enhancing user experience and advertising efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026070939000001_ABST
    Figure 2026070939000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] A means of receiving category information based on user interests and using that information to select digital content, A means for converting selected digital content into audio and video formats using speech synthesis technology and video generation technology, A means for delivering converted audio and video content to the user's terminal, A means of collecting user content viewing and advertising viewing data and optimizing future content delivery based on this data, An information distribution system that includes this.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of the chatbot's character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In recent years, with the development of information technology, a large amount of digital content is generated every day, but means for efficiently delivering this to consumers are required. Also, it is important to provide personalized content according to the interests and preferences of individual users in order to increase consumer interest. Furthermore, for companies, it is required to increase profits by realizing effective advertisement distribution, but there is a problem that it is difficult to provide appropriate advertisements according to user interests.

Means for Solving the Problems

[0005] This invention provides a means for receiving category information based on user interests and using it to select digital content. The selected digital content is converted into audio and video formats using speech synthesis and video generation technologies and delivered to the user's terminal. Furthermore, by collecting data on the user's content viewing and advertising viewing and optimizing the next content delivery based on this data, the invention realizes a system that enables the provision of useful information to users and effective advertising delivery for companies.

[0006] A "user" is a person who uses an information distribution system to receive content and advertisements based on their personal interests and preferences.

[0007] "Category information" refers to information that indicates the classification of digital content related to the user's interests and concerns.

[0008] "Digital content" refers to information created or distributed electronically, and includes formats such as text, audio, and video.

[0009] "Selection" is the process of classifying digital content based on established criteria and selecting those that are most relevant to the user.

[0010] "Speech synthesis technology" is a technology that uses computers to convert text data into sounds that sound like human speech.

[0011] "Video generation technology" is a technology that automatically creates visual video content using digital data.

[0012] "Distribution" refers to the process of delivering specific digital content to a user's device via a communication network.

[0013] A "terminal" is an electronic device used by users to receive and view content from an information distribution system.

[0014] "Viewing data" is a record of which content and advertisements a user has viewed through an information distribution system.

[0015] "Optimization" is the process of analyzing the collected data to more efficiently and effectively adjust the distribution of content and advertisements in subsequent times.

Brief Explanation of Drawings

[0016] [Figure 1] It is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] It is a conceptual diagram showing an example of the main functions of a data processing device and a smart device according to the first embodiment. [Figure 3] It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which a plurality of emotions are mapped. [Figure 10] It shows an emotion map to which a plurality of emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13]It is a sequence diagram showing the processing flow of the data processing system in Embodiment 2 when the emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when the emotion engine is combined.

Mode for Carrying Out the Invention

[0017] Hereinafter, an example of an embodiment of the system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0018] First, the terms used in the following description will be explained.

[0019] In the following embodiments, the numbered processor (hereinafter simply referred to as "processor") may be one arithmetic unit or a combination of a plurality of arithmetic units. Also, the processor may be one type of arithmetic unit or a combination of a plurality of types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0020] In the following embodiments, the numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0021] In the following embodiments, the numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, etc.

[0022] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0023] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0024] [First Embodiment]

[0025] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0026] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0027] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0028] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0029] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0030] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0031] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0032] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0033] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0034] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0035] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0036] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0037] This invention is a system that delivers digital content in audio and video formats based on user interests. This system mainly consists of a server, terminals, and users.

[0038] The server uses APIs to collect information daily from contracted media platforms to obtain the latest digital content. The server analyzes this collected digital content using AI to determine the topic and importance of each article. In doing so, it refers to the user's profile information, i.e., category information indicating their interests, to select the content most relevant to the user.

[0039] A terminal is a device owned by the user and used to view content. Users can set categories of interest through their terminal. This setting allows the server to filter content to suit individual users.

[0040] The server can convert the selected content into speech that sounds like a human voice using speech synthesis technology, and then further convert it into video format using video generation technology. This audio and video content is sent to the user's device when it is deemed appropriate for distribution.

[0041] For example, for users interested in sports, the latest match results and related news are provided as audio, and match highlight videos are generated. Users can then view these on their devices while commuting.

[0042] Furthermore, after delivery, the server collects user viewing data and uses it to further personalize content. This data includes playback time, completion rate, and ad response, and contributes to optimizing future content delivery and advertising.

[0043] In this way, the present invention is designed to enable users to efficiently consume meaningful content tailored to their daily interests through the system. For advertisers, it also enables targeted advertising, leading to improved profitability.

[0044] The following describes the processing flow.

[0045] Step 1:

[0046] As part of the initial setup, users use the device and set their categories of interest. This information is stored on the server as a user profile.

[0047] Step 2:

[0048] The server accesses the APIs of contracted media outlets at specified times to collect the latest news articles and digital content. This ensures that the information is always up-to-date.

[0049] Step 3:

[0050] The server uses AI to analyze the collected content. The AI ​​understands the content of the articles and extracts each topic and related keywords.

[0051] Step 4:

[0052] The server filters relevant content based on the user's profile. The filtered content is selected to best match the user's interests.

[0053] Step 5:

[0054] The server uses speech synthesis technology to convert filtered articles into speech. At this stage, it is adjusted to produce natural-sounding speech.

[0055] Step 6:

[0056] The server utilizes video generation technology as needed to create video that corresponds to the audio content. This enables the delivery of visual content.

[0057] Step 7:

[0058] The server delivers the generated audio and video content to the user's device. The delivery timing is adjusted according to the user's settings.

[0059] Step 8:

[0060] Users view content delivered to their devices. Data such as their actions during viewing and their reactions to advertisements are recorded.

[0061] Step 9:

[0062] The server collects viewing data and analyzes it for future content delivery. This allows for a more personalized content experience.

[0063] (Example 1)

[0064] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0065] In modern society, users are surrounded by a vast amount of digital information, making it difficult to efficiently acquire information that matches their interests. Furthermore, traditional information delivery systems have not fully utilized individual users' interests and behavioral history, making personalized information delivery difficult. Additionally, insufficient advertising targeting has posed a challenge in improving advertising effectiveness.

[0066] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0067] In this invention, the server includes means for receiving information based on the user's interests, analyzing the digital information using a generative AI model, and selecting highly relevant information; means for converting the selected digital information into audio and video formats using speech synthesis technology and video generation technology; and means for transmitting the converted audio and video information to the user's terminal. This enables the provision of personalized information based on the user's interests and improves the effectiveness of targeted advertising.

[0068] "Information based on user interests" refers to information that indicates a user's interests and concerns, obtained through their profile information and behavioral history.

[0069] A "generative AI model" is an algorithmic model that uses artificial intelligence technology to analyze data and make predictions or generate data.

[0070] "Digital information" refers to information that exists in electronic form and includes a variety of media such as text, images, audio, and video.

[0071] "Speech synthesis technology" is a technology that converts text data into speech data and generates speech that closely resembles a human voice.

[0072] "Video generation technology" refers to technology used to convert content such as text and still images into video format.

[0073] A "user terminal" is an electronic device owned by a user and used for transmitting or receiving information.

[0074] "Notification" refers to a means of sending messages to inform users of information or events.

[0075] "Interest-based personalization" refers to identifying and optimizing information and services according to the individual interests and preferences of users.

[0076] "Targeted advertising effectiveness" refers to the efficiency of advertising activities that target specific consumers or market segments.

[0077] This invention is an information distribution system that efficiently delivers digital information in audio and video formats based on user interests. This system primarily consists of a server, terminals, and users.

[0078] The server periodically acquires the latest digital information via the internet, based on contracts with multiple online information sources and providers. This includes database access and automated information retrieval using APIs. The acquired information is analyzed using a generative AI model. This model is trained on a large dataset and is used, for example, to understand the meaning of text, summarize it, and classify it as a topic. Specifically, the server uses prompts such as "Output a summary and topic of this article" to provide the generative AI model.

[0079] The server uses the output of this generating AI model to combine it with user profile information, i.e., data indicating their interests and preferences, and selects highly relevant information. This results in information optimized for each individual user.

[0080] The user's device is a smartphone, tablet, or other similar electronic device that functions as a receiver for information transmitted from the server. The device also provides an interface for the user to set their interest categories. This setting allows the server to select information with greater precision.

[0081] The selected information is converted into audio data on the server using speech synthesis technology. Furthermore, video data is generated by adding the relevant video material and visual representations using video generation technology. Specifically, general text-to-speech software is used for speech synthesis, and simple video editing software is used for video generation.

[0082] The audio and video information generated in this way is transmitted to the user's device. For example, a user can receive a summary of the latest news and related highlight videos during their commute. Furthermore, this system collects the user's viewing data after delivery and continues to use it for future deliveries and advertisements, thereby continuously improving its personalization capabilities.

[0083] By using such an information distribution system, users can efficiently consume important information based on their daily interests, and advertisers can develop effective advertising strategies with clearly defined targets.

[0084] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0085] Step 1:

[0086] The server retrieves the latest digital information from the internet via the API of its contracted information source. It sends HTTP requests using the API endpoint URL and authentication information as input. As output, it receives JSON data of articles and news, which it stores in its internal storage.

[0087] Step 2:

[0088] The server converts the acquired digital information into a format necessary for inputting it into the generating AI model. Specifically, it parses JSON data and extracts the article text and metadata. The stored JSON data is used as input, and prompts for the generating AI model are generated as output.

[0089] Step 3:

[0090] The server sends prompts to the generating AI model to analyze the digital information. The input is a prompt (e.g., "Please output the summary and topic of this article"). The output is a summary and topic classification results returned by the AI ​​model.

[0091] Step 4:

[0092] The server matches the output of the AI ​​model with the user's profile information to select information of interest. The inputs used are the analysis results of the AI ​​model and the user's interest data. The output is a list of information highly relevant to the user.

[0093] Step 5:

[0094] The server processes the selected information for visualization and audio production using speech synthesis and video generation technologies. The selected information is used as input, and audio and video data are created based on its title and content. The output is data converted into auditory and visual media formats.

[0095] Step 6:

[0096] The server transmits the generated audio and video data to the user's device. Visual and auditory media data are used as input, and these are delivered to the user's device as output. Delivery is performed according to the delivery schedule set by the user.

[0097] Step 7:

[0098] The device presents the received content to the user and records viewing data. Received audio and video data are used as input, and user data such as viewing start time, completion rate, and ad response are returned to the server as output.

[0099] Step 8:

[0100] The server uses the collected viewing data to refine its next delivery strategy. The collected viewing data is used as input, and the output is a new content recommendation list based on user preferences and behavior, optimizing future delivery and advertising strategies.

[0101] (Application Example 1)

[0102] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0103] In today's world, where digital content is becoming increasingly abundant and diverse, it is becoming difficult for users to efficiently acquire information that matches their interests. Furthermore, there is a growing need for easy ways to enjoy content of interest while on the go or amidst busy daily life. In addition, optimizing content delivery based on viewing data, including ad viewing, is becoming increasingly important.

[0104] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0105] In this invention, the server includes means for receiving category information based on the user's interests and using that information to select digital content; means for converting the selected digital content into audio and video formats using speech synthesis technology and video generation technology; and means for transmitting the generated content to a portable display device and providing audio and video. As a result, users can always efficiently receive content based on their interests and easily acquire information while on the go or in their busy daily lives.

[0106] "Category information based on user interests" refers to information that indicates the themes and genres that users are interested in, and is data used for content selection.

[0107] "Digital content" refers to information resources such as audio, video, and text that are stored and used electronically.

[0108] "Speech synthesis technology" is a technology that converts text data into speech data, and is used to generate speech that sounds like a human voice.

[0109] "Video generation technology" is a technology that expresses still images and text as videos, and is a technology for creating visually appealing videos.

[0110] A "visual display device" is a device used to visually display digital content and to provide images and videos to users.

[0111] A "portable display device" is a device for displaying content that can be carried and used by the user.

[0112] "Content viewing data" refers to statistical data about the content that users have played, including viewing time and viewing completion rate.

[0113] The system for implementing this invention mainly consists of a server, a terminal, and a user. The server is equipped with means for selecting digital content based on category information derived from the user's interests. This information is collected by referring to the user's set interest data, and the latest digital content is collected from contracted information sources via API, and the content is analyzed using a generative AI model.

[0114] Furthermore, the server uses speech synthesis technology to convert the selected content into speech that closely resembles a human voice, and then processes it into video format using video generation technology. This process utilizes technologies such as "Google® Text-to-Speech" for speech synthesis and "OpenCV" for video generation. The generated media content is then transmitted to a portable display device (e.g., smart glasses), which is a visual display device, and delivered to the user so that it can be viewed in real time.

[0115] In this system, the device functions as a platform for users to enjoy highly personalized services while maintaining a portable form. Users can access the latest information and content of interest presented in audio and video formats in a way that is easy to use while on the go. Their viewing and advertising response data is then sent back to the server to help optimize future content and advertising delivery.

[0116] As an example, consider a scenario where a user selects sports news. The server filters the latest match results and player information, providing them with audio along with match highlight videos. A specific prompt might be: "Summarize the latest soccer match and generate audio commentary along with key match highlight videos. The audio should be natural and easy to listen to, and include match results and information on notable players." In this way, users can efficiently view content tailored to their interests and maximize their viewing time.

[0117] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0118] Step 1:

[0119] The server collects the latest digital content through the APIs of contracted information sources. It uses category information and access information for the information sources as input. Based on this, it executes API requests and receives the latest digital content as output.

[0120] Step 2:

[0121] The server analyzes the collected digital content using a generative AI model. The input consists of content data obtained in step 1 and category information based on user interests. The AI ​​model classifies the content by topic and evaluates its importance. The output generates a list of content highly relevant to the user.

[0122] Step 3:

[0123] The server converts selected content into audio data using speech synthesis technology. Text-based content data is used as input. The speech synthesis engine converts the text into natural-sounding speech data and outputs an audio file.

[0124] Step 4:

[0125] The server generates video content using video generation technology along with audio data. It uses audio files and associated images or video footage as input. Video editing software is used to synchronize the audio and video, and the completed video file is output.

[0126] Step 5:

[0127] The server transmits the generated audio and video content to the user's portable display device. Using the user's terminal information and video files as input, the server transmits data to the terminal via the network. Upon receiving the content on the terminal, the user can immediately view it.

[0128] Step 6:

[0129] The terminal provides the received content to the user visually and aurally. It uses audio and video data received from the server as input. Playback is performed through the terminal's built-in display and speaker, completing the presentation on the visual display device. The user then utilizes the content in this step.

[0130] Step 7:

[0131] The server collects user content viewing data. It uses data such as viewing completion rates and ad responses as input. This data is then analyzed to generate output data for optimizing future content delivery and advertising.

[0132] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0133] This invention incorporates an emotion engine into an information distribution system that provides digital content tailored to user interests, thereby realizing a personalized content experience that responds to the user's emotional state. This system mainly consists of a server, terminals, and users, and operates with integrated emotion recognition technology.

[0134] The server retrieves the latest news and articles from contracted media outlets and analyzes them using generative AI. This allows it to select highly relevant digital content, taking into account the user's profile information and emotional state.

[0135] A terminal is a device that allows users to receive information and view content. By incorporating an emotion engine that detects the user's facial expressions and voice into this terminal, the user's emotional state can be recognized in real time.

[0136] User emotional data is collected through facial expression and voice analysis. Based on this data, the server utilizes speech synthesis and video generation technologies to adjust the tone and style of the content. For example, if the user is relaxed, content with a calm tone will be provided; if they are excited, an energetic tone will be selected.

[0137] As a concrete example, suppose a sports-loving user receives the latest match results. If the emotion engine detects that the user is disappointed because their favorite team lost, the server can supplement the information with positive news or uplifting content to capture their interest.

[0138] Furthermore, the device sends this emotional data to the server, which analyzes it along with viewing data to optimize future content delivery. This enriches the experience for each individual user and contributes to the appropriate delivery of advertisements.

[0139] This invention enables the provision of information that resonates with users' emotions, resulting in more effective content experiences and advertising services.

[0140] The following describes the processing flow.

[0141] Step 1:

[0142] Users use their devices to set category information based on their interests and preferences. This sends the user's profile to the server.

[0143] Step 2:

[0144] The server collects the latest digital content using the APIs of contracted media platforms. This information is then subjected to topic analysis and keyword extraction by generative AI.

[0145] Step 3:

[0146] An emotion engine built into the device analyzes the user's facial expressions and voice to recognize the user's emotional state in real time.

[0147] Step 4:

[0148] The server receives user emotion data sent from the emotion engine and selects content according to the user's current emotional state. This includes selecting calming content when the user is relaxed and energetic content when the user is excited.

[0149] Step 5:

[0150] The server generates narration using speech synthesis technology for the selected content and creates video content using video generation technology as needed.

[0151] Step 6:

[0152] The server delivers the generated audio and video content to the user's device. It can also send push notifications at the optimal time, taking into account changes in the user's emotional state.

[0153] Step 7:

[0154] Users view the delivered content using their devices. The emotion engine continues to recognize changes in the user's emotions while they are viewing the content, and this data is sent to the server.

[0155] Step 8:

[0156] The server collects viewing and sentiment data, which is then used to improve future content delivery. This process makes the user experience more personalized.

[0157] (Example 2)

[0158] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0159] Conventional information distribution systems are insufficient in providing personalized content based on users' interests and emotional states, limiting their ability to improve user experience and advertising effectiveness. Furthermore, they are unable to properly analyze emotional states and reflect them in real time, creating a need for optimal information delivery tailored to each user's needs.

[0160] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0161] This invention includes a server that receives information based on the user's interests and emotional state and uses that information to select content; a server that converts the selected content into audio and video formats using speech synthesis and video creation technologies; and a server that analyzes the user's facial expressions and voice data to recognize their emotional state. This enables the provision of personalized content that is tailored to each user's emotions, resulting in a richer experience and more effective advertising delivery.

[0162] "User interests" refer to the content and topics that users are interested in, and serve as the criteria for selecting the information to provide based on these interests.

[0163] "Emotional state" refers to data that indicates the user's emotional condition, such as joy, sadness, excitement, or relaxation.

[0164] "Content selection" refers to the process of choosing appropriate digital information based on the user's interests and emotional state.

[0165] "Speech synthesis technology" refers to the technology that synthesizes human voices using analog or digital data and outputs them in speech format.

[0166] "Video creation technology" refers to the technology used to process visual digital data and create content in video format.

[0167] "Facial and voice data" refers to information that captures the user's facial movements and voice tone, and is used to analyze their emotional state.

[0168] A "contracted source" refers to a source of data from which a server retrieves information for content delivery, and is a medium or platform that is officially granted access rights.

[0169] A "generative AI model" is a type of artificial intelligence that refers to an algorithm capable of generating new data and content based on large amounts of existing data.

[0170] This information distribution system consists of servers, terminals, and users, and is designed to provide personalized content based on users' interests and emotions. The specific roles of each component are shown below.

[0171] The server collects the latest news and articles from contracted information sources using APIs and web scraping techniques. The collected information is analyzed using generative AI models. Specifically, the server uses this information to classify topics and perform sentiment analysis, selecting the most relevant content based on the user's profile information and emotional state. Python libraries and machine learning algorithms are used in this selection process.

[0172] The device is where users view content and is equipped with an emotion engine. This engine has the ability to detect the user's facial expressions and voice in real time and recognize their emotional state. For example, it uses a camera and microphone to capture facial expressions and voice tone, and then analyzes this data with a machine learning model.

[0173] If a user prefers sports news, the server selects the latest match results based on that interest, and the generated content is delivered to the device. If the emotion engine determines that the user is disappointed, the server improves the quality of the user experience by adding positive information to the delivery.

[0174] As a concrete example, consider the following prompt statements: "If the user is relaxed while reading this news article, create a summary in a calm tone," or "Based on the user's emotional response, suggest ways to optimize the selection of content to deliver next." This allows for the delivery of information that matches the user's emotions and interests, resulting in a better content experience.

[0175] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0176] Step 1:

[0177] The server retrieves the latest news and articles from contracted information sources using APIs and web scraping techniques. Input requires access information such as the URL of the information source and an API key. The retrieved information is stored in a database and used in the next processing step. This step specifically involves running a script that automatically retrieves information every hour.

[0178] Step 2:

[0179] The server analyzes the acquired articles and news using a generative AI model. The input is the text data of the articles acquired in step 1. The generative AI model classifies topics, performs sentiment analysis, and labels the content of the articles. The program outputs these results and selects content that matches the user's interests. Specifically, this involves data analysis using scripts that leverage natural language processing techniques.

[0180] Step 3:

[0181] The device uses an emotion engine to collect the user's facial expressions and voice in real time. Inputs include image and audio data acquired from the device's built-in camera and microphone. This data is analyzed using machine learning algorithms to recognize the user's emotional state and output an evaluation result. Specific operations include the execution of facial recognition software and voice analysis software.

[0182] Step 4:

[0183] The server creates personalized content based on the user's emotional state and interests. Its inputs include content information generated in step 2 and the emotional evaluation results obtained in step 3. Using these, it employs speech synthesis and video creation technologies to output digital content in audio and video formats preferred by the user. Specific operations include the process of adding selected music and visual effects.

[0184] Step 5:

[0185] The user views the provided personalized content on their device. The input is the audio and video content delivered from the server in step 4. The output is the user's viewing experience, and their emotional reactions based on that experience. Specific actions include using a video player and providing interactive feedback.

[0186] Step 6:

[0187] The device sends user viewing data and reactions to the server. Input includes log data of the content the user viewed and real-time sentiment data. This data is aggregated and analyzed on the server to optimize the content delivered next time. Specific actions in this step include secure data transfer and statistical analysis within the server.

[0188] (Application Example 2)

[0189] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0190] In digital content distribution, it is essential to consider not only users' interests but also their emotional states in real time to provide appropriate content and advertising experiences. However, conventional systems do not adequately consider users' emotional states, making it difficult to provide personalized experiences based on those emotions. To solve this problem, a content distribution method that is attentive to users' emotions is necessary.

[0191] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0192] This invention includes a server that receives category information based on the user's interests and emotional state, and uses that information to select digital content; a server that converts the selected digital content into audio and video formats using speech synthesis and video generation technologies, and adjusts the tone and style according to the user's emotional state; and a server that delivers the converted audio and video content to the user's terminal and receives feedback based on the emotional state. This makes it possible to provide personalized content and advertising experiences based on the user's emotions.

[0193] "User interests" refer to information that indicates the degree of interest a user has in a particular field or topic, based on their past behavioral history and preferences.

[0194] "Emotional state" refers to the psychological or physiological state that a user is currently experiencing, and is usually determined by analyzing facial expressions and voice.

[0195] "Category information" refers to information used to classify digital content according to a specific theme or topic.

[0196] "Digital content" is a general term for media that can be distributed in digital format, such as audio, video, and text.

[0197] "Speech synthesis technology" is a technology that converts text into speech data and is widely used in assistive technologies for the visually impaired and in artificial intelligence-based dialogue systems.

[0198] "Video generation technology" is a technology that converts abstract data and information into a visual video format, and is used when creating animations and computer graphics (CG).

[0199] "Adjusting tone and style" refers to the process of changing the atmosphere and presentation style according to the content, and is particularly important for providing a personalized experience that responds to the user's emotions.

[0200] "User terminal" refers to equipment or devices used by users to receive or view information, and includes smartphones, tablets, and other similar devices.

[0201] "Feedback" refers to the reactions and opinions received from users, and is information used to improve and optimize the system.

[0202] The system based on this invention consists of a terminal equipped with an emotion engine and an associated server. In order to implement the invention, it is necessary to build a system that can analyze the user's emotional state in real time. The emotion engine built into the terminal analyzes facial expressions and voice and generates the user's emotional data. In this process, emotion recognition technologies such as OpenCV and MediaPipe are utilized to determine emotional states such as relaxation and stress from the facial expression data and voice data.

[0203] The server uses a generative AI model to select digital content, taking into account emotional data and user interests. This generative AI has the ability to retrieve and analyze the latest news and articles from contracted content providers. The analyzed content is then adjusted to match the user's emotional state using speech synthesis and video generation technologies. For example, if the user is feeling tired, relaxing videos with soothing background music will be provided.

[0204] The server then delivers the adjusted content to the device and collects user feedback. This feedback data is used to optimize future content delivery. By monitoring user reactions in real time and having the generative AI model learn from this information, the quality of the content delivered improves over time.

[0205] For example, if a user opens a fitness app on a holiday and facial recognition determines that the user is in need of energy, the server can select and deliver an energetic training video to boost their spirits.

[0206] An example of a prompt message would be, "Please provide fitness content that matches the perceived emotion. If you are relaxed, please select yoga; if you are stressed, please select full-body stretches." This system operates on user devices such as smartphones, enabling flexible enhancement of the user's content experience.

[0207] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0208] Step 1:

[0209] The device collects the user's facial expressions and voice in real time using an emotion engine. The input is the user's video and audio data. This data is processed into emotional state data using emotion recognition libraries such as OpenCV and MediaPipe. The output is the user's emotional state (e.g., relaxed, stressed).

[0210] Step 2:

[0211] The server receives emotional states transmitted from the terminal and pre-stored user interest data as input. This input is fed to a generative AI model, which selects relevant digital content based on conditions indicated by prompts. News and articles obtained in real-time from contracted media outlets are used for content selection. The output is digital content that matches the user's emotional state.

[0212] Step 3:

[0213] The server converts selected digital content into audio and video formats using speech synthesis and video generation technologies. During this process, it adjusts the tone and style according to the emotional state. The input is the selected digital content. The output is emotion-sensitive audio and video data for delivery to the user.

[0214] Step 4:

[0215] The server delivers the converted audio and video content to the user's device. The device receives this data as input and plays it back as content for the user. The output obtained from this playback is feedback data such as the user's satisfaction level and emotional changes.

[0216] Step 5:

[0217] The device collects data about the user's content viewing and sends it to the server. This data includes playback time, skipped sections, and the user's emotional state. The input is the user's viewing behavior and emotional data, and the output is analytical data sent to the server.

[0218] Step 6:

[0219] The server uses feedback and collected data to optimize future content delivery. This process involves a generative AI model learning to provide even more personalized content. The input is user feedback data, and the output is the improved content delivery strategy.

[0220] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0221] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0222] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0223] [Second Embodiment]

[0224] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0225] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0226] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0227] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0228] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0229] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0230] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0231] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0232] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0233] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0234] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0235] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0236] This invention is a system that delivers digital content in audio and video formats based on user interests. This system mainly consists of a server, terminals, and users.

[0237] The server uses APIs to collect information daily from contracted media platforms to obtain the latest digital content. The server analyzes this collected digital content using AI to determine the topic and importance of each article. In doing so, it refers to the user's profile information, i.e., category information indicating their interests, to select the content most relevant to the user.

[0238] A terminal is a device owned by the user and used to view content. Users can set categories of interest through their terminal. This setting allows the server to filter content to suit individual users.

[0239] The server can convert the selected content into speech that sounds like a human voice using speech synthesis technology, and then further convert it into video format using video generation technology. This audio and video content is sent to the user's device when it is deemed appropriate for distribution.

[0240] For example, for users interested in sports, the latest match results and related news are provided as audio, and match highlight videos are generated. Users can then view these on their devices while commuting.

[0241] Furthermore, after delivery, the server collects user viewing data and uses it to further personalize content. This data includes playback time, completion rate, and ad response, and contributes to optimizing future content delivery and advertising.

[0242] In this way, the present invention is designed to enable users to efficiently consume meaningful content tailored to their daily interests through the system. For advertisers, it also enables targeted advertising, leading to improved profitability.

[0243] The following describes the processing flow.

[0244] Step 1:

[0245] As part of the initial setup, users use the device and set their categories of interest. This information is stored on the server as a user profile.

[0246] Step 2:

[0247] The server accesses the APIs of contracted media outlets at specified times to collect the latest news articles and digital content. This ensures that the information is always up-to-date.

[0248] Step 3:

[0249] The server uses AI to analyze the collected content. The AI ​​understands the content of the articles and extracts each topic and related keywords.

[0250] Step 4:

[0251] The server filters relevant content based on the user's profile. The filtered content is selected to best match the user's interests.

[0252] Step 5:

[0253] The server uses speech synthesis technology to convert filtered articles into speech. At this stage, it is adjusted to produce natural-sounding speech.

[0254] Step 6:

[0255] The server utilizes video generation technology as needed to create video that corresponds to the audio content. This enables the delivery of visual content.

[0256] Step 7:

[0257] The server delivers the generated audio and video content to the user's device. The delivery timing is adjusted according to the user's settings.

[0258] Step 8:

[0259] Users view content delivered to their devices. Data such as their actions during viewing and their reactions to advertisements are recorded.

[0260] Step 9:

[0261] The server collects viewing data and analyzes it for future content delivery. This allows for a more personalized content experience.

[0262] (Example 1)

[0263] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0264] In modern society, users are surrounded by a vast amount of digital information, making it difficult to efficiently acquire information that matches their interests. Furthermore, traditional information delivery systems have not fully utilized individual users' interests and behavioral history, making personalized information delivery difficult. Additionally, insufficient advertising targeting has posed a challenge in improving advertising effectiveness.

[0265] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0266] In this invention, the server includes means for receiving information based on the user's interests, analyzing the digital information using a generative AI model, and selecting highly relevant information; means for converting the selected digital information into audio and video formats using speech synthesis technology and video generation technology; and means for transmitting the converted audio and video information to the user's terminal. This enables the provision of personalized information based on the user's interests and improves the effectiveness of targeted advertising.

[0267] "Information based on user interests" refers to information that indicates a user's interests and concerns, obtained through their profile information and behavioral history.

[0268] A "generative AI model" is an algorithmic model that uses artificial intelligence technology to analyze data and make predictions or generate data.

[0269] "Digital information" refers to information that exists in electronic form and includes a variety of media such as text, images, audio, and video.

[0270] "Speech synthesis technology" is a technology that converts text data into speech data and generates speech that closely resembles a human voice.

[0271] "Video generation technology" refers to technology used to convert content such as text and still images into video format.

[0272] A "user terminal" is an electronic device owned by a user and used for transmitting or receiving information.

[0273] "Notification" refers to a means of sending messages to inform users of information or events.

[0274] "Interest-based personalization" refers to identifying and optimizing information and services according to the individual interests and preferences of users.

[0275] "Targeted advertising effectiveness" refers to the efficiency of advertising activities that target specific consumers or market segments.

[0276] This invention is an information distribution system that efficiently delivers digital information in audio and video formats based on user interests. This system primarily consists of a server, terminals, and users.

[0277] The server periodically acquires the latest digital information via the internet, based on contracts with multiple online information sources and providers. This includes database access and automated information retrieval using APIs. The acquired information is analyzed using a generative AI model. This model is trained on a large dataset and is used, for example, to understand the meaning of text, summarize it, and classify it as a topic. Specifically, the server uses prompts such as "Output a summary and topic of this article" to provide the generative AI model.

[0278] The server uses the output of this generating AI model to combine it with user profile information, i.e., data indicating their interests and preferences, and selects highly relevant information. This results in information optimized for each individual user.

[0279] The user's device is a smartphone, tablet, or other similar electronic device that functions as a receiver for information transmitted from the server. The device also provides an interface for the user to set their interest categories. This setting allows the server to select information with greater precision.

[0280] The selected information is converted into audio data on the server using speech synthesis technology. Furthermore, video data is generated by adding the relevant video material and visual representations using video generation technology. Specifically, general text-to-speech software is used for speech synthesis, and simple video editing software is used for video generation.

[0281] The audio and video information generated in this way is transmitted to the user's device. For example, a user can receive a summary of the latest news and related highlight videos during their commute. Furthermore, this system collects the user's viewing data after delivery and continues to use it for future deliveries and advertisements, thereby continuously improving its personalization capabilities.

[0282] By using such an information distribution system, users can efficiently consume important information based on their daily interests, and for advertisers, it also enables an effective advertising strategy with a clear target.

[0283] The flow of the specific process in Example 1 will be described using FIG. 11.

[0284] Step 1:

[0285] The server obtains the latest digital information from the Internet via the API of the contract information source. Using the API endpoint URL and authentication information as input, it sends an HTTP request. As output, it receives JSON data of articles and news and saves this in the server's internal storage.

[0286] Step 2:

[0287] The server converts the acquired digital information into the format necessary for input to the generation AI model. Specifically, it analyzes the JSON data and extracts the article text and metadata. Using the saved JSON data as input, it generates a prompt sentence for the generation AI model as output.

[0288] Step 3:

[0289] The server sends a prompt sentence to the generation AI model and analyzes the digital information. Using the prompt sentence (e.g., "Please output the summary and topics of this article") as input, it sends it to the AI model. As output, a summary sentence and topic classification result are returned from the AI model.

[0290] Step 4:

[0291] The server collates the output of the AI model with the user's profile information and selects the information of interest. Using the analysis result of the AI model and the user's interest data as input, it creates a list of information highly relevant to the user as output.

[0292] Step 5:

[0293] The server processes the selected information for visualization and audio production using speech synthesis and video generation technologies. The selected information is used as input, and audio and video data are created based on its title and content. The output is data converted into auditory and visual media formats.

[0294] Step 6:

[0295] The server transmits the generated audio and video data to the user's device. Visual and auditory media data are used as input, and these are delivered to the user's device as output. Delivery is performed according to the delivery schedule set by the user.

[0296] Step 7:

[0297] The device presents the received content to the user and records viewing data. Received audio and video data are used as input, and user data such as viewing start time, completion rate, and ad response are returned to the server as output.

[0298] Step 8:

[0299] The server uses the collected viewing data to refine its next delivery strategy. The collected viewing data is used as input, and the output is a new content recommendation list based on user preferences and behavior, optimizing future delivery and advertising strategies.

[0300] (Application Example 1)

[0301] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0302] In modern times when the quantity and variety of digital content are increasing, it has become difficult for users to efficiently obtain information that matches their interests. Also, in the midst of a busy daily life or while on the move, there is a need for means to easily enjoy content of interest. Furthermore, the optimization of content distribution based on viewing data including advertisement viewing has also increased in importance.

[0303] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0304] In this invention, the server includes means for receiving category information based on the interests of the user and using that information to select digital content, means for converting the selected digital content into audio and video formats using voice synthesis technology and video generation technology, and means for transmitting the generated content to a portable display device to provide audio and video. As a result, the user can always efficiently receive content based on their own interests and can easily obtain information even while on the move or in a busy daily life.

[0305] The "category information based on the interests of the user" is information indicating themes or genres in which the user is interested, and is data used for content selection.

[0306] "Digital content" is an information resource such as audio, video, and text that is electronically stored and used.

[0307] "Voice synthesis technology" is a technology for converting text data into voice data and is a technology for generating voice that sounds like a human voice.

[0308] "Video generation technology" is a technology for expressing still images and text as video and is a technology for creating visually appealing videos.

[0309] A "visual display device" is a device used to visually display digital content and to provide images and videos to users.

[0310] A "portable display device" is a device for displaying content that can be carried and used by the user.

[0311] "Content viewing data" refers to statistical data about the content that users have played, including viewing time and viewing completion rate.

[0312] The system for implementing this invention mainly consists of a server, a terminal, and a user. The server is equipped with means for selecting digital content based on category information derived from the user's interests. This information is collected by referring to the user's set interest data, and the latest digital content is collected from contracted information sources via API, and the content is analyzed using a generative AI model.

[0313] Furthermore, the server uses speech synthesis technology to convert the selected content into speech that closely resembles a human voice, and then processes it into video format using video generation technology. This process utilizes technologies such as "Google Text-to-Speech" for speech synthesis and "OpenCV" for video generation. The generated media content is then transmitted to a portable display device (e.g., smart glasses), which is a visual display device, and delivered to the user for real-time viewing.

[0314] In this system, the device functions as a platform for users to enjoy highly personalized services while maintaining a portable form. Users can access the latest information and content of interest presented in audio and video formats in a way that is easy to use while on the go. Their viewing and advertising response data is then sent back to the server to help optimize future content and advertising delivery.

[0315] As an example, consider a scenario where a user selects sports news. The server filters the latest match results and player information, providing them with audio along with match highlight videos. A specific prompt might be: "Summarize the latest soccer match and generate audio commentary along with key match highlight videos. The audio should be natural and easy to listen to, and include match results and information on notable players." In this way, users can efficiently view content tailored to their interests and maximize their viewing time.

[0316] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0317] Step 1:

[0318] The server collects the latest digital content through the APIs of contracted information sources. It uses category information and access information for the information sources as input. Based on this, it executes API requests and receives the latest digital content as output.

[0319] Step 2:

[0320] The server analyzes the collected digital content using a generative AI model. The input consists of content data obtained in step 1 and category information based on user interests. The AI ​​model classifies the content by topic and evaluates its importance. The output generates a list of content highly relevant to the user.

[0321] Step 3:

[0322] The server converts selected content into audio data using speech synthesis technology. Text-based content data is used as input. The speech synthesis engine converts the text into natural-sounding speech data and outputs an audio file.

[0323] Step 4:

[0324] The server generates video content using video generation technology along with audio data. It uses audio files and associated images or video footage as input. Video editing software is used to synchronize the audio and video, and the completed video file is output.

[0325] Step 5:

[0326] The server transmits the generated audio and video content to the user's portable display device. Using the user's terminal information and video files as input, the server transmits data to the terminal via the network. Upon receiving the content on the terminal, the user can immediately view it.

[0327] Step 6:

[0328] The terminal provides the received content to the user visually and aurally. It uses audio and video data received from the server as input. Playback is performed through the terminal's built-in display and speaker, completing the presentation on the visual display device. The user then utilizes the content in this step.

[0329] Step 7:

[0330] The server collects user content viewing data. It uses data such as viewing completion rates and ad responses as input. This data is then analyzed to generate output data for optimizing future content delivery and advertising.

[0331] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0332] This invention incorporates an emotion engine into an information distribution system that provides digital content tailored to user interests, thereby realizing a personalized content experience that responds to the user's emotional state. This system mainly consists of a server, terminals, and users, and operates with integrated emotion recognition technology.

[0333] The server retrieves the latest news and articles from contracted media outlets and analyzes them using generative AI. This allows it to select highly relevant digital content, taking into account the user's profile information and emotional state.

[0334] A terminal is a device that allows users to receive information and view content. By incorporating an emotion engine that detects the user's facial expressions and voice into this terminal, the user's emotional state can be recognized in real time.

[0335] User emotional data is collected through facial expression and voice analysis. Based on this data, the server utilizes speech synthesis and video generation technologies to adjust the tone and style of the content. For example, if the user is relaxed, content with a calm tone will be provided; if they are excited, an energetic tone will be selected.

[0336] As a concrete example, suppose a sports-loving user receives the latest match results. If the emotion engine detects that the user is disappointed because their favorite team lost, the server can supplement the information with positive news or uplifting content to capture their interest.

[0337] Furthermore, the device sends this emotional data to the server, which analyzes it along with viewing data to optimize future content delivery. This enriches the experience for each individual user and contributes to the appropriate delivery of advertisements.

[0338] This invention enables the provision of information that resonates with users' emotions, resulting in more effective content experiences and advertising services.

[0339] The following describes the processing flow.

[0340] Step 1:

[0341] Users use their devices to set category information based on their interests and preferences. This sends the user's profile to the server.

[0342] Step 2:

[0343] The server collects the latest digital content using the APIs of contracted media platforms. This information is then subjected to topic analysis and keyword extraction by generative AI.

[0344] Step 3:

[0345] An emotion engine built into the device analyzes the user's facial expressions and voice to recognize the user's emotional state in real time.

[0346] Step 4:

[0347] The server receives user emotion data sent from the emotion engine and selects content according to the user's current emotional state. This includes selecting calming content when the user is relaxed and energetic content when the user is excited.

[0348] Step 5:

[0349] The server generates narration using speech synthesis technology for the selected content and creates video content using video generation technology as needed.

[0350] Step 6:

[0351] The server delivers the generated audio and video content to the user's device. It can also send push notifications at the optimal time, taking into account changes in the user's emotional state.

[0352] Step 7:

[0353] Users view the delivered content using their devices. The emotion engine continues to recognize changes in the user's emotions while they are viewing the content, and this data is sent to the server.

[0354] Step 8:

[0355] The server collects viewing and sentiment data, which is then used to improve future content delivery. This process makes the user experience more personalized.

[0356] (Example 2)

[0357] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0358] Conventional information distribution systems are insufficient in providing personalized content based on users' interests and emotional states, limiting their ability to improve user experience and advertising effectiveness. Furthermore, they are unable to properly analyze emotional states and reflect them in real time, creating a need for optimal information delivery tailored to each user's needs.

[0359] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0360] This invention includes a server that receives information based on the user's interests and emotional state and uses that information to select content; a server that converts the selected content into audio and video formats using speech synthesis and video creation technologies; and a server that analyzes the user's facial expressions and voice data to recognize their emotional state. This enables the provision of personalized content that is tailored to each user's emotions, resulting in a richer experience and more effective advertising delivery.

[0361] "User interests" refer to the content and topics that users are interested in, and serve as the criteria for selecting the information to provide based on these interests.

[0362] "Emotional state" refers to data that indicates the user's emotional condition, such as joy, sadness, excitement, or relaxation.

[0363] "Content selection" refers to the process of choosing appropriate digital information based on the user's interests and emotional state.

[0364] "Speech synthesis technology" refers to the technology that synthesizes human voices using analog or digital data and outputs them in speech format.

[0365] "Video creation technology" refers to the technology used to process visual digital data and create content in video format.

[0366] "Facial and voice data" refers to information that captures the user's facial movements and voice tone, and is used to analyze their emotional state.

[0367] A "contracted source" refers to a source of data from which a server retrieves information for content delivery, and is a medium or platform that is officially granted access rights.

[0368] A "generative AI model" is a type of artificial intelligence that refers to an algorithm capable of generating new data and content based on large amounts of existing data.

[0369] This information distribution system consists of servers, terminals, and users, and is designed to provide personalized content based on users' interests and emotions. The specific roles of each component are shown below.

[0370] The server collects the latest news and articles from contracted information sources using APIs and web scraping techniques. The collected information is analyzed using generative AI models. Specifically, the server uses this information to classify topics and perform sentiment analysis, selecting the most relevant content based on the user's profile information and emotional state. Python libraries and machine learning algorithms are used in this selection process.

[0371] The device is where users view content and is equipped with an emotion engine. This engine has the ability to detect the user's facial expressions and voice in real time and recognize their emotional state. For example, it uses a camera and microphone to capture facial expressions and voice tone, and then analyzes this data with a machine learning model.

[0372] If a user prefers sports news, the server selects the latest match results based on that interest, and the generated content is delivered to the device. If the emotion engine determines that the user is disappointed, the server improves the quality of the user experience by adding positive information to the delivery.

[0373] As a concrete example, consider the following prompt statements: "If the user is relaxed while reading this news article, create a summary in a calm tone," or "Based on the user's emotional response, suggest ways to optimize the selection of content to deliver next." This allows for the delivery of information that matches the user's emotions and interests, resulting in a better content experience.

[0374] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0375] Step 1:

[0376] The server retrieves the latest news and articles from contracted information sources using APIs and web scraping techniques. Input requires access information such as the URL of the information source and an API key. The retrieved information is stored in a database and used in the next processing step. This step specifically involves running a script that automatically retrieves information every hour.

[0377] Step 2:

[0378] The server analyzes the acquired articles and news using a generative AI model. The input is the text data of the articles acquired in step 1. The generative AI model classifies topics, performs sentiment analysis, and labels the content of the articles. The program outputs these results and selects content that matches the user's interests. Specifically, this involves data analysis using scripts that leverage natural language processing techniques.

[0379] Step 3:

[0380] The device uses an emotion engine to collect the user's facial expressions and voice in real time. Inputs include image and audio data acquired from the device's built-in camera and microphone. This data is analyzed using machine learning algorithms to recognize the user's emotional state and output an evaluation result. Specific operations include the execution of facial recognition software and voice analysis software.

[0381] Step 4:

[0382] The server creates personalized content based on the user's emotional state and interests. Its inputs include content information generated in step 2 and the emotional evaluation results obtained in step 3. Using these, it employs speech synthesis and video creation technologies to output digital content in audio and video formats preferred by the user. Specific operations include the process of adding selected music and visual effects.

[0383] Step 5:

[0384] The user views the provided personalized content on their device. The input is the audio and video content delivered from the server in step 4. The output is the user's viewing experience, and their emotional reactions based on that experience. Specific actions include using a video player and providing interactive feedback.

[0385] Step 6:

[0386] The device sends user viewing data and reactions to the server. Input includes log data of the content the user viewed and real-time sentiment data. This data is aggregated and analyzed on the server to optimize the content delivered next time. Specific actions in this step include secure data transfer and statistical analysis within the server.

[0387] (Application Example 2)

[0388] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0389] In digital content distribution, it is essential to consider not only users' interests but also their emotional states in real time to provide appropriate content and advertising experiences. However, conventional systems do not adequately consider users' emotional states, making it difficult to provide personalized experiences based on those emotions. To solve this problem, a content distribution method that is attentive to users' emotions is necessary.

[0390] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0391] This invention includes a server that receives category information based on the user's interests and emotional state, and uses that information to select digital content; a server that converts the selected digital content into audio and video formats using speech synthesis and video generation technologies, and adjusts the tone and style according to the user's emotional state; and a server that delivers the converted audio and video content to the user's terminal and receives feedback based on the emotional state. This makes it possible to provide personalized content and advertising experiences based on the user's emotions.

[0392] "User interests" refer to information that indicates the degree of interest a user has in a particular field or topic, based on their past behavioral history and preferences.

[0393] "Emotional state" refers to the psychological or physiological state that a user is currently experiencing, and is usually determined by analyzing facial expressions and voice.

[0394] "Category information" refers to information used to classify digital content according to a specific theme or topic.

[0395] "Digital content" is a general term for media that can be distributed in digital format, such as audio, video, and text.

[0396] "Speech synthesis technology" is a technology that converts text into speech data and is widely used in assistive technologies for the visually impaired and in artificial intelligence-based dialogue systems.

[0397] "Video generation technology" is a technology that converts abstract data and information into a visual video format, and is used when creating animations and computer graphics (CG).

[0398] "Adjusting tone and style" refers to the process of changing the atmosphere and presentation style according to the content, and is particularly important for providing a personalized experience that responds to the user's emotions.

[0399] "User terminal" refers to equipment or devices used by users to receive or view information, and includes smartphones, tablets, and other similar devices.

[0400] "Feedback" refers to the reactions and opinions received from users, and is information used to improve and optimize the system.

[0401] The system based on this invention consists of a terminal equipped with an emotion engine and an associated server. In order to implement the invention, it is necessary to build a system that can analyze the user's emotional state in real time. The emotion engine built into the terminal analyzes facial expressions and voice and generates the user's emotional data. In this process, emotion recognition technologies such as OpenCV and MediaPipe are utilized to determine emotional states such as relaxation and stress from the facial expression data and voice data.

[0402] The server uses a generative AI model to select digital content, taking into account emotional data and user interests. This generative AI has the ability to retrieve and analyze the latest news and articles from contracted content providers. The analyzed content is then adjusted to match the user's emotional state using speech synthesis and video generation technologies. For example, if the user is feeling tired, relaxing videos with soothing background music will be provided.

[0403] The server then delivers the adjusted content to the device and collects user feedback. This feedback data is used to optimize future content delivery. By monitoring user reactions in real time and having the generative AI model learn from this information, the quality of the content delivered improves over time.

[0404] For example, if a user opens a fitness app on a holiday and facial recognition determines that the user is in need of energy, the server can select and deliver an energetic training video to boost their spirits.

[0405] An example of a prompt message would be, "Please provide fitness content that matches the perceived emotion. If you are relaxed, please select yoga; if you are stressed, please select full-body stretches." This system operates on user devices such as smartphones, enabling flexible enhancement of the user's content experience.

[0406] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0407] Step 1:

[0408] The device collects the user's facial expressions and voice in real time using an emotion engine. The input is the user's video and audio data. This data is processed into emotional state data using emotion recognition libraries such as OpenCV and MediaPipe. The output is the user's emotional state (e.g., relaxed, stressed).

[0409] Step 2:

[0410] The server receives emotional states transmitted from the terminal and pre-stored user interest data as input. This input is fed to a generative AI model, which selects relevant digital content based on conditions indicated by prompts. News and articles obtained in real-time from contracted media outlets are used for content selection. The output is digital content that matches the user's emotional state.

[0411] Step 3:

[0412] The server converts selected digital content into audio and video formats using speech synthesis and video generation technologies. During this process, it adjusts the tone and style according to the emotional state. The input is the selected digital content. The output is emotion-sensitive audio and video data for delivery to the user.

[0413] Step 4:

[0414] The server delivers the converted audio and video content to the user's device. The device receives this data as input and plays it back as content for the user. The output obtained from this playback is feedback data such as the user's satisfaction level and emotional changes.

[0415] Step 5:

[0416] The device collects data about the user's content viewing and sends it to the server. This data includes playback time, skipped sections, and the user's emotional state. The input is the user's viewing behavior and emotional data, and the output is analytical data sent to the server.

[0417] Step 6:

[0418] The server uses feedback and collected data to optimize future content delivery. This process involves a generative AI model learning to provide even more personalized content. The input is user feedback data, and the output is the improved content delivery strategy.

[0419] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0420] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0421] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0422] [Third Embodiment]

[0423] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0424] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0425] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0426] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0427] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0428] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0429] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0430] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0431] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0432] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0433] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0434] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0435] This invention is a system that delivers digital content in audio and video formats based on user interests. This system mainly consists of a server, terminals, and users.

[0436] The server uses APIs to collect information daily from contracted media platforms to obtain the latest digital content. The server analyzes this collected digital content using AI to determine the topic and importance of each article. In doing so, it refers to the user's profile information, i.e., category information indicating their interests, to select the content most relevant to the user.

[0437] A terminal is a device owned by the user and used to view content. Users can set categories of interest through their terminal. This setting allows the server to filter content to suit individual users.

[0438] The server can convert the selected content into speech that sounds like a human voice using speech synthesis technology, and then further convert it into video format using video generation technology. This audio and video content is sent to the user's device when it is deemed appropriate for distribution.

[0439] For example, for users interested in sports, the latest match results and related news are provided as audio, and match highlight videos are generated. Users can then view these on their devices while commuting.

[0440] Furthermore, after delivery, the server collects user viewing data and uses it to further personalize content. This data includes playback time, completion rate, and ad response, and contributes to optimizing future content delivery and advertising.

[0441] In this way, the present invention is designed to enable users to efficiently consume meaningful content tailored to their daily interests through the system. For advertisers, it also enables targeted advertising, leading to improved profitability.

[0442] The following describes the processing flow.

[0443] Step 1:

[0444] As part of the initial setup, users use the device and set their categories of interest. This information is stored on the server as a user profile.

[0445] Step 2:

[0446] The server accesses the APIs of contracted media outlets at specified times to collect the latest news articles and digital content. This ensures that the information is always up-to-date.

[0447] Step 3:

[0448] The server uses AI to analyze the collected content. The AI ​​understands the content of the articles and extracts each topic and related keywords.

[0449] Step 4:

[0450] The server filters relevant content based on the user's profile. The filtered content is selected to best match the user's interests.

[0451] Step 5:

[0452] The server uses speech synthesis technology to convert filtered articles into speech. At this stage, it is adjusted to produce natural-sounding speech.

[0453] Step 6:

[0454] The server utilizes video generation technology as needed to create video that corresponds to the audio content. This enables the delivery of visual content.

[0455] Step 7:

[0456] The server delivers the generated audio and video content to the user's device. The delivery timing is adjusted according to the user's settings.

[0457] Step 8:

[0458] Users view content delivered to their devices. Data such as their actions during viewing and their reactions to advertisements are recorded.

[0459] Step 9:

[0460] The server collects viewing data and analyzes it for future content delivery. This allows for a more personalized content experience.

[0461] (Example 1)

[0462] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0463] In modern society, users are surrounded by a vast amount of digital information, making it difficult to efficiently acquire information that matches their interests. Furthermore, traditional information delivery systems have not fully utilized individual users' interests and behavioral history, making personalized information delivery difficult. Additionally, insufficient advertising targeting has posed a challenge in improving advertising effectiveness.

[0464] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0465] In this invention, the server includes means for receiving information based on the user's interests, analyzing the digital information using a generative AI model, and selecting highly relevant information; means for converting the selected digital information into audio and video formats using speech synthesis technology and video generation technology; and means for transmitting the converted audio and video information to the user's terminal. This enables the provision of personalized information based on the user's interests and improves the effectiveness of targeted advertising.

[0466] "Information based on user interests" refers to information that indicates a user's interests and concerns, obtained through their profile information and behavioral history.

[0467] A "generative AI model" is an algorithmic model that uses artificial intelligence technology to analyze data and make predictions or generate data.

[0468] "Digital information" refers to information that exists in electronic form and includes a variety of media such as text, images, audio, and video.

[0469] "Speech synthesis technology" is a technology that converts text data into speech data and generates speech that closely resembles a human voice.

[0470] "Video generation technology" refers to technology used to convert content such as text and still images into video format.

[0471] A "user terminal" is an electronic device owned by a user and used for transmitting or receiving information.

[0472] "Notification" refers to a means of sending messages to inform users of information or events.

[0473] "Interest-based personalization" refers to identifying and optimizing information and services according to the individual interests and preferences of users.

[0474] "Targeted advertising effectiveness" refers to the efficiency of advertising activities that target specific consumers or market segments.

[0475] This invention is an information distribution system that efficiently delivers digital information in audio and video formats based on user interests. This system primarily consists of a server, terminals, and users.

[0476] The server periodically acquires the latest digital information via the internet, based on contracts with multiple online information sources and providers. This includes database access and automated information retrieval using APIs. The acquired information is analyzed using a generative AI model. This model is trained on a large dataset and is used, for example, to understand the meaning of text, summarize it, and classify it as a topic. Specifically, the server uses prompts such as "Output a summary and topic of this article" to provide the generative AI model.

[0477] The server uses the output of this generating AI model to combine it with user profile information, i.e., data indicating their interests and preferences, and selects highly relevant information. This results in information optimized for each individual user.

[0478] The user's device is a smartphone, tablet, or other similar electronic device that functions as a receiver for information transmitted from the server. The device also provides an interface for the user to set their interest categories. This setting allows the server to select information with greater precision.

[0479] The selected information is converted into audio data on the server using speech synthesis technology. Furthermore, video data is generated by adding the relevant video material and visual representations using video generation technology. Specifically, general text-to-speech software is used for speech synthesis, and simple video editing software is used for video generation.

[0480] The audio and video information generated in this way is transmitted to the user's device. For example, a user can receive a summary of the latest news and related highlight videos during their commute. Furthermore, this system collects the user's viewing data after delivery and continues to use it for future deliveries and advertisements, thereby continuously improving its personalization capabilities.

[0481] By using such an information distribution system, users can efficiently consume important information based on their daily interests, and advertisers can develop effective advertising strategies with clearly defined targets.

[0482] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0483] Step 1:

[0484] The server retrieves the latest digital information from the internet via the API of its contracted information source. It sends HTTP requests using the API endpoint URL and authentication information as input. As output, it receives JSON data of articles and news, which it stores in its internal storage.

[0485] Step 2:

[0486] The server converts the acquired digital information into a format necessary for inputting it into the generating AI model. Specifically, it parses JSON data and extracts the article text and metadata. The stored JSON data is used as input, and prompts for the generating AI model are generated as output.

[0487] Step 3:

[0488] The server sends prompts to the generating AI model to analyze the digital information. The input is a prompt (e.g., "Please output the summary and topic of this article"). The output is a summary and topic classification results returned by the AI ​​model.

[0489] Step 4:

[0490] The server matches the output of the AI ​​model with the user's profile information to select information of interest. The inputs used are the analysis results of the AI ​​model and the user's interest data. The output is a list of information highly relevant to the user.

[0491] Step 5:

[0492] The server processes the selected information for visualization and audio production using speech synthesis and video generation technologies. The selected information is used as input, and audio and video data are created based on its title and content. The output is data converted into auditory and visual media formats.

[0493] Step 6:

[0494] The server transmits the generated audio and video data to the user's device. Visual and auditory media data are used as input, and these are delivered to the user's device as output. Delivery is performed according to the delivery schedule set by the user.

[0495] Step 7:

[0496] The device presents the received content to the user and records viewing data. Received audio and video data are used as input, and user data such as viewing start time, completion rate, and ad response are returned to the server as output.

[0497] Step 8:

[0498] The server uses the collected viewing data to refine its next delivery strategy. The collected viewing data is used as input, and the output is a new content recommendation list based on user preferences and behavior, optimizing future delivery and advertising strategies.

[0499] (Application Example 1)

[0500] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0501] In today's world, where digital content is becoming increasingly abundant and diverse, it is becoming difficult for users to efficiently acquire information that matches their interests. Furthermore, there is a growing need for easy ways to enjoy content of interest while on the go or amidst busy daily life. In addition, optimizing content delivery based on viewing data, including ad viewing, is becoming increasingly important.

[0502] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0503] In this invention, the server includes means for receiving category information based on the user's interests and using that information to select digital content; means for converting the selected digital content into audio and video formats using speech synthesis technology and video generation technology; and means for transmitting the generated content to a portable display device and providing audio and video. As a result, users can always efficiently receive content based on their interests and easily acquire information while on the go or in their busy daily lives.

[0504] "Category information based on user interests" refers to information that indicates the themes and genres that users are interested in, and is data used for content selection.

[0505] "Digital content" refers to information resources such as audio, video, and text that are stored and used electronically.

[0506] "Speech synthesis technology" is a technology that converts text data into speech data, and is used to generate speech that sounds like a human voice.

[0507] "Video generation technology" is a technology that expresses still images and text as videos, and is a technology for creating visually appealing videos.

[0508] A "visual display device" is a device used to visually display digital content and to provide images and videos to users.

[0509] A "portable display device" is a device for displaying content that can be carried and used by the user.

[0510] "Content viewing data" refers to statistical data about the content that users have played, including viewing time and viewing completion rate.

[0511] The system for implementing this invention mainly consists of a server, a terminal, and a user. The server is equipped with means for selecting digital content based on category information derived from the user's interests. This information is collected by referring to the user's set interest data, and the latest digital content is collected from contracted information sources via API, and the content is analyzed using a generative AI model.

[0512] Furthermore, the server uses speech synthesis technology to convert the selected content into speech that closely resembles a human voice, and then processes it into video format using video generation technology. This process utilizes technologies such as "Google Text-to-Speech" for speech synthesis and "OpenCV" for video generation. The generated media content is then transmitted to a portable display device (e.g., smart glasses), which is a visual display device, and delivered to the user for real-time viewing.

[0513] In this system, the device functions as a platform for users to enjoy highly personalized services while maintaining a portable form. Users can access the latest information and content of interest presented in audio and video formats in a way that is easy to use while on the go. Their viewing and advertising response data is then sent back to the server to help optimize future content and advertising delivery.

[0514] As an example, consider a scenario where a user selects sports news. The server filters the latest match results and player information, providing them with audio along with match highlight videos. A specific prompt might be: "Summarize the latest soccer match and generate audio commentary along with key match highlight videos. The audio should be natural and easy to listen to, and include match results and information on notable players." In this way, users can efficiently view content tailored to their interests and maximize their viewing time.

[0515] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0516] Step 1:

[0517] The server collects the latest digital content through the APIs of contracted information sources. It uses category information and access information for the information sources as input. Based on this, it executes API requests and receives the latest digital content as output.

[0518] Step 2:

[0519] The server analyzes the collected digital content using a generative AI model. The input consists of content data obtained in step 1 and category information based on user interests. The AI ​​model classifies the content by topic and evaluates its importance. The output generates a list of content highly relevant to the user.

[0520] Step 3:

[0521] The server converts selected content into audio data using speech synthesis technology. Text-based content data is used as input. The speech synthesis engine converts the text into natural-sounding speech data and outputs an audio file.

[0522] Step 4:

[0523] The server generates video content using video generation technology along with audio data. It uses audio files and associated images or video footage as input. Video editing software is used to synchronize the audio and video, and the completed video file is output.

[0524] Step 5:

[0525] The server transmits the generated audio and video content to the user's portable display device. Using the user's terminal information and video files as input, the server transmits data to the terminal via the network. Upon receiving the content on the terminal, the user can immediately view it.

[0526] Step 6:

[0527] The terminal provides the received content to the user visually and aurally. It uses audio and video data received from the server as input. Playback is performed through the terminal's built-in display and speaker, completing the presentation on the visual display device. The user then utilizes the content in this step.

[0528] Step 7:

[0529] The server collects user content viewing data. It uses data such as viewing completion rates and ad responses as input. This data is then analyzed to generate output data for optimizing future content delivery and advertising.

[0530] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0531] This invention incorporates an emotion engine into an information distribution system that provides digital content tailored to user interests, thereby realizing a personalized content experience that responds to the user's emotional state. This system mainly consists of a server, terminals, and users, and operates with integrated emotion recognition technology.

[0532] The server retrieves the latest news and articles from contracted media outlets and analyzes them using generative AI. This allows it to select highly relevant digital content, taking into account the user's profile information and emotional state.

[0533] A terminal is a device that allows users to receive information and view content. By incorporating an emotion engine that detects the user's facial expressions and voice into this terminal, the user's emotional state can be recognized in real time.

[0534] User emotional data is collected through facial expression and voice analysis. Based on this data, the server utilizes speech synthesis and video generation technologies to adjust the tone and style of the content. For example, if the user is relaxed, content with a calm tone will be provided; if they are excited, an energetic tone will be selected.

[0535] As a concrete example, suppose a sports-loving user receives the latest match results. If the emotion engine detects that the user is disappointed because their favorite team lost, the server can supplement the information with positive news or uplifting content to capture their interest.

[0536] Furthermore, the device sends this emotional data to the server, which analyzes it along with viewing data to optimize future content delivery. This enriches the experience for each individual user and contributes to the appropriate delivery of advertisements.

[0537] This invention enables the provision of information that resonates with users' emotions, resulting in more effective content experiences and advertising services.

[0538] The following describes the processing flow.

[0539] Step 1:

[0540] Users use their devices to set category information based on their interests and preferences. This sends the user's profile to the server.

[0541] Step 2:

[0542] The server collects the latest digital content using the APIs of contracted media platforms. This information is then subjected to topic analysis and keyword extraction by generative AI.

[0543] Step 3:

[0544] An emotion engine built into the device analyzes the user's facial expressions and voice to recognize the user's emotional state in real time.

[0545] Step 4:

[0546] The server receives user emotion data sent from the emotion engine and selects content according to the user's current emotional state. This includes selecting calming content when the user is relaxed and energetic content when the user is excited.

[0547] Step 5:

[0548] The server generates narration using speech synthesis technology for the selected content and creates video content using video generation technology as needed.

[0549] Step 6:

[0550] The server delivers the generated audio and video content to the user's device. It can also send push notifications at the optimal time, taking into account changes in the user's emotional state.

[0551] Step 7:

[0552] Users view the delivered content using their devices. The emotion engine continues to recognize changes in the user's emotions while they are viewing the content, and this data is sent to the server.

[0553] Step 8:

[0554] The server collects viewing and sentiment data, which is then used to improve future content delivery. This process makes the user experience more personalized.

[0555] (Example 2)

[0556] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0557] Conventional information distribution systems are insufficient in providing personalized content based on users' interests and emotional states, limiting their ability to improve user experience and advertising effectiveness. Furthermore, they are unable to properly analyze emotional states and reflect them in real time, creating a need for optimal information delivery tailored to each user's needs.

[0558] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0559] This invention includes a server that receives information based on the user's interests and emotional state and uses that information to select content; a server that converts the selected content into audio and video formats using speech synthesis and video creation technologies; and a server that analyzes the user's facial expressions and voice data to recognize their emotional state. This enables the provision of personalized content that is tailored to each user's emotions, resulting in a richer experience and more effective advertising delivery.

[0560] "User interests" refer to the content and topics that users are interested in, and serve as the criteria for selecting the information to provide based on these interests.

[0561] "Emotional state" refers to data that indicates the user's emotional condition, such as joy, sadness, excitement, or relaxation.

[0562] "Content selection" refers to the process of choosing appropriate digital information based on the user's interests and emotional state.

[0563] "Speech synthesis technology" refers to the technology that synthesizes human voices using analog or digital data and outputs them in speech format.

[0564] "Video creation technology" refers to the technology used to process visual digital data and create content in video format.

[0565] "Facial and voice data" refers to information that captures the user's facial movements and voice tone, and is used to analyze their emotional state.

[0566] A "contracted source" refers to a source of data from which a server retrieves information for content delivery, and is a medium or platform that is officially granted access rights.

[0567] A "generative AI model" is a type of artificial intelligence that refers to an algorithm capable of generating new data and content based on large amounts of existing data.

[0568] This information distribution system consists of servers, terminals, and users, and is designed to provide personalized content based on users' interests and emotions. The specific roles of each component are shown below.

[0569] The server collects the latest news and articles from contracted information sources using APIs and web scraping techniques. The collected information is analyzed using generative AI models. Specifically, the server uses this information to classify topics and perform sentiment analysis, selecting the most relevant content based on the user's profile information and emotional state. Python libraries and machine learning algorithms are used in this selection process.

[0570] The device is where users view content and is equipped with an emotion engine. This engine has the ability to detect the user's facial expressions and voice in real time and recognize their emotional state. For example, it uses a camera and microphone to capture facial expressions and voice tone, and then analyzes this data with a machine learning model.

[0571] If a user prefers sports news, the server selects the latest match results based on that interest, and the generated content is delivered to the device. If the emotion engine determines that the user is disappointed, the server improves the quality of the user experience by adding positive information to the delivery.

[0572] As a concrete example, consider the following prompt statements: "If the user is relaxed while reading this news article, create a summary in a calm tone," or "Based on the user's emotional response, suggest ways to optimize the selection of content to deliver next." This allows for the delivery of information that matches the user's emotions and interests, resulting in a better content experience.

[0573] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0574] Step 1:

[0575] The server retrieves the latest news and articles from contracted information sources using APIs and web scraping techniques. Input requires access information such as the URL of the information source and an API key. The retrieved information is stored in a database and used in the next processing step. This step specifically involves running a script that automatically retrieves information every hour.

[0576] Step 2:

[0577] The server analyzes the acquired articles and news using a generative AI model. The input is the text data of the articles acquired in step 1. The generative AI model classifies topics, performs sentiment analysis, and labels the content of the articles. The program outputs these results and selects content that matches the user's interests. Specifically, this involves data analysis using scripts that leverage natural language processing techniques.

[0578] Step 3:

[0579] The device uses an emotion engine to collect the user's facial expressions and voice in real time. Inputs include image and audio data acquired from the device's built-in camera and microphone. This data is analyzed using machine learning algorithms to recognize the user's emotional state and output an evaluation result. Specific operations include the execution of facial recognition software and voice analysis software.

[0580] Step 4:

[0581] The server creates personalized content based on the user's emotional state and interests. Its inputs include content information generated in step 2 and the emotional evaluation results obtained in step 3. Using these, it employs speech synthesis and video creation technologies to output digital content in audio and video formats preferred by the user. Specific operations include the process of adding selected music and visual effects.

[0582] Step 5:

[0583] The user views the provided personalized content on their device. The input is the audio and video content delivered from the server in step 4. The output is the user's viewing experience, and their emotional reactions based on that experience. Specific actions include using a video player and providing interactive feedback.

[0584] Step 6:

[0585] The device sends user viewing data and reactions to the server. Input includes log data of the content the user viewed and real-time sentiment data. This data is aggregated and analyzed on the server to optimize the content delivered next time. Specific actions in this step include secure data transfer and statistical analysis within the server.

[0586] (Application Example 2)

[0587] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0588] In digital content distribution, it is essential to consider not only users' interests but also their emotional states in real time to provide appropriate content and advertising experiences. However, conventional systems do not adequately consider users' emotional states, making it difficult to provide personalized experiences based on those emotions. To solve this problem, a content distribution method that is attentive to users' emotions is necessary.

[0589] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0590] This invention includes a server that receives category information based on the user's interests and emotional state, and uses that information to select digital content; a server that converts the selected digital content into audio and video formats using speech synthesis and video generation technologies, and adjusts the tone and style according to the user's emotional state; and a server that delivers the converted audio and video content to the user's terminal and receives feedback based on the emotional state. This makes it possible to provide personalized content and advertising experiences based on the user's emotions.

[0591] "User interests" refer to information that indicates the degree of interest a user has in a particular field or topic, based on their past behavioral history and preferences.

[0592] "Emotional state" refers to the psychological or physiological state that a user is currently experiencing, and is usually determined by analyzing facial expressions and voice.

[0593] "Category information" refers to information used to classify digital content according to a specific theme or topic.

[0594] "Digital content" is a general term for media that can be distributed in digital format, such as audio, video, and text.

[0595] "Speech synthesis technology" is a technology that converts text into speech data and is widely used in assistive technologies for the visually impaired and in artificial intelligence-based dialogue systems.

[0596] "Video generation technology" is a technology that converts abstract data and information into a visual video format, and is used when creating animations and computer graphics (CG).

[0597] "Adjusting tone and style" refers to the process of changing the atmosphere and presentation style according to the content, and is particularly important for providing a personalized experience that responds to the user's emotions.

[0598] "User terminal" refers to equipment or devices used by users to receive or view information, and includes smartphones, tablets, and other similar devices.

[0599] "Feedback" refers to the reactions and opinions received from users, and is information used to improve and optimize the system.

[0600] The system based on this invention consists of a terminal equipped with an emotion engine and an associated server. In order to implement the invention, it is necessary to build a system that can analyze the user's emotional state in real time. The emotion engine built into the terminal analyzes facial expressions and voice and generates the user's emotional data. In this process, emotion recognition technologies such as OpenCV and MediaPipe are utilized to determine emotional states such as relaxation and stress from the facial expression data and voice data.

[0601] The server uses a generative AI model to select digital content, taking into account emotional data and user interests. This generative AI has the ability to retrieve and analyze the latest news and articles from contracted content providers. The analyzed content is then adjusted to match the user's emotional state using speech synthesis and video generation technologies. For example, if the user is feeling tired, relaxing videos with soothing background music will be provided.

[0602] The server then delivers the adjusted content to the device and collects user feedback. This feedback data is used to optimize future content delivery. By monitoring user reactions in real time and having the generative AI model learn from this information, the quality of the content delivered improves over time.

[0603] For example, if a user opens a fitness app on a holiday and facial recognition determines that the user is in need of energy, the server can select and deliver an energetic training video to boost their spirits.

[0604] An example of a prompt message would be, "Please provide fitness content that matches the perceived emotion. If you are relaxed, please select yoga; if you are stressed, please select full-body stretches." This system operates on user devices such as smartphones, enabling flexible enhancement of the user's content experience.

[0605] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0606] Step 1:

[0607] The device collects the user's facial expressions and voice in real time using an emotion engine. The input is the user's video and audio data. This data is processed into emotional state data using emotion recognition libraries such as OpenCV and MediaPipe. The output is the user's emotional state (e.g., relaxed, stressed).

[0608] Step 2:

[0609] The server receives emotional states transmitted from the terminal and pre-stored user interest data as input. This input is fed to a generative AI model, which selects relevant digital content based on conditions indicated by prompts. News and articles obtained in real-time from contracted media outlets are used for content selection. The output is digital content that matches the user's emotional state.

[0610] Step 3:

[0611] The server converts selected digital content into audio and video formats using speech synthesis and video generation technologies. During this process, it adjusts the tone and style according to the emotional state. The input is the selected digital content. The output is emotion-sensitive audio and video data for delivery to the user.

[0612] Step 4:

[0613] The server delivers the converted audio and video content to the user's device. The device receives this data as input and plays it back as content for the user. The output obtained from this playback is feedback data such as the user's satisfaction level and emotional changes.

[0614] Step 5:

[0615] The device collects data about the user's content viewing and sends it to the server. This data includes playback time, skipped sections, and the user's emotional state. The input is the user's viewing behavior and emotional data, and the output is analytical data sent to the server.

[0616] Step 6:

[0617] The server uses feedback and collected data to optimize future content delivery. This process involves a generative AI model learning to provide even more personalized content. The input is user feedback data, and the output is the improved content delivery strategy.

[0618] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0619] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0620] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0621] [Fourth Embodiment]

[0622] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0623] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0624] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0625] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0626] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0627] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0628] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0629] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0630] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0631] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0632] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0633] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0634] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0635] This invention is a system that delivers digital content in audio and video formats based on user interests. This system mainly consists of a server, terminals, and users.

[0636] The server uses APIs to collect information daily from contracted media platforms to obtain the latest digital content. The server analyzes this collected digital content using AI to determine the topic and importance of each article. In doing so, it refers to the user's profile information, i.e., category information indicating their interests, to select the content most relevant to the user.

[0637] A terminal is a device owned by the user and used to view content. Users can set categories of interest through their terminal. This setting allows the server to filter content to suit individual users.

[0638] The server can convert the selected content into speech that sounds like a human voice using speech synthesis technology, and then further convert it into video format using video generation technology. This audio and video content is sent to the user's device when it is deemed appropriate for distribution.

[0639] For example, for users interested in sports, the latest match results and related news are provided as audio, and match highlight videos are generated. Users can then view these on their devices while commuting.

[0640] Furthermore, after delivery, the server collects user viewing data and uses it to further personalize content. This data includes playback time, completion rate, and ad response, and contributes to optimizing future content delivery and advertising.

[0641] In this way, the present invention is designed to enable users to efficiently consume meaningful content tailored to their daily interests through the system. For advertisers, it also enables targeted advertising, leading to improved profitability.

[0642] The following describes the processing flow.

[0643] Step 1:

[0644] As part of the initial setup, users use the device and set their categories of interest. This information is stored on the server as a user profile.

[0645] Step 2:

[0646] The server accesses the APIs of contracted media outlets at specified times to collect the latest news articles and digital content. This ensures that the information is always up-to-date.

[0647] Step 3:

[0648] The server uses AI to analyze the collected content. The AI ​​understands the content of the articles and extracts each topic and related keywords.

[0649] Step 4:

[0650] The server filters relevant content based on the user's profile. The filtered content is selected to best match the user's interests.

[0651] Step 5:

[0652] The server uses speech synthesis technology to convert filtered articles into speech. At this stage, it is adjusted to produce natural-sounding speech.

[0653] Step 6:

[0654] The server utilizes video generation technology as needed to create video that corresponds to the audio content. This enables the delivery of visual content.

[0655] Step 7:

[0656] The server delivers the generated audio and video content to the user's device. The delivery timing is adjusted according to the user's settings.

[0657] Step 8:

[0658] Users view content delivered to their devices. Data such as their actions during viewing and their reactions to advertisements are recorded.

[0659] Step 9:

[0660] The server collects viewing data and analyzes it for future content delivery. This allows for a more personalized content experience.

[0661] (Example 1)

[0662] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0663] In modern society, users are surrounded by a vast amount of digital information, making it difficult to efficiently acquire information that matches their interests. Furthermore, traditional information delivery systems have not fully utilized individual users' interests and behavioral history, making personalized information delivery difficult. Additionally, insufficient advertising targeting has posed a challenge in improving advertising effectiveness.

[0664] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0665] In this invention, the server includes means for receiving information based on the user's interests, analyzing the digital information using a generative AI model, and selecting highly relevant information; means for converting the selected digital information into audio and video formats using speech synthesis technology and video generation technology; and means for transmitting the converted audio and video information to the user's terminal. This enables the provision of personalized information based on the user's interests and improves the effectiveness of targeted advertising.

[0666] "Information based on user interests" refers to information that indicates a user's interests and concerns, obtained through their profile information and behavioral history.

[0667] A "generative AI model" is an algorithmic model that uses artificial intelligence technology to analyze data and make predictions or generate data.

[0668] "Digital information" refers to information that exists in electronic form and includes a variety of media such as text, images, audio, and video.

[0669] "Speech synthesis technology" is a technology that converts text data into speech data and generates speech that closely resembles a human voice.

[0670] "Video generation technology" refers to technology used to convert content such as text and still images into video format.

[0671] A "user terminal" is an electronic device owned by a user and used for transmitting or receiving information.

[0672] "Notification" refers to a means of sending messages to inform users of information or events.

[0673] "Interest-based personalization" refers to identifying and optimizing information and services according to the individual interests and preferences of users.

[0674] "Targeted advertising effectiveness" refers to the efficiency of advertising activities that target specific consumers or market segments.

[0675] This invention is an information distribution system that efficiently delivers digital information in audio and video formats based on user interests. This system primarily consists of a server, terminals, and users.

[0676] The server periodically acquires the latest digital information via the internet, based on contracts with multiple online information sources and providers. This includes database access and automated information retrieval using APIs. The acquired information is analyzed using a generative AI model. This model is trained on a large dataset and is used, for example, to understand the meaning of text, summarize it, and classify it as a topic. Specifically, the server uses prompts such as "Output a summary and topic of this article" to provide the generative AI model.

[0677] The server uses the output of this generating AI model to combine it with user profile information, i.e., data indicating their interests and preferences, and selects highly relevant information. This results in information optimized for each individual user.

[0678] The user's device is a smartphone, tablet, or other similar electronic device that functions as a receiver for information transmitted from the server. The device also provides an interface for the user to set their interest categories. This setting allows the server to select information with greater precision.

[0679] The selected information is converted into audio data on the server using speech synthesis technology. Furthermore, video data is generated by adding the relevant video material and visual representations using video generation technology. Specifically, general text-to-speech software is used for speech synthesis, and simple video editing software is used for video generation.

[0680] The audio and video information generated in this way is transmitted to the user's device. For example, a user can receive a summary of the latest news and related highlight videos during their commute. Furthermore, this system collects the user's viewing data after delivery and continues to use it for future deliveries and advertisements, thereby continuously improving its personalization capabilities.

[0681] By using such an information distribution system, users can efficiently consume important information based on their daily interests, and advertisers can develop effective advertising strategies with clearly defined targets.

[0682] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0683] Step 1:

[0684] The server retrieves the latest digital information from the internet via the API of its contracted information source. It sends HTTP requests using the API endpoint URL and authentication information as input. As output, it receives JSON data of articles and news, which it stores in its internal storage.

[0685] Step 2:

[0686] The server converts the acquired digital information into a format necessary for inputting it into the generating AI model. Specifically, it parses JSON data and extracts the article text and metadata. The stored JSON data is used as input, and prompts for the generating AI model are generated as output.

[0687] Step 3:

[0688] The server sends prompts to the generating AI model to analyze the digital information. The input is a prompt (e.g., "Please output the summary and topic of this article"). The output is a summary and topic classification results returned by the AI ​​model.

[0689] Step 4:

[0690] The server matches the output of the AI ​​model with the user's profile information to select information of interest. The inputs used are the analysis results of the AI ​​model and the user's interest data. The output is a list of information highly relevant to the user.

[0691] Step 5:

[0692] The server processes the selected information for visualization and audio production using speech synthesis and video generation technologies. The selected information is used as input, and audio and video data are created based on its title and content. The output is data converted into auditory and visual media formats.

[0693] Step 6:

[0694] The server transmits the generated audio and video data to the user's device. Visual and auditory media data are used as input, and these are delivered to the user's device as output. Delivery is performed according to the delivery schedule set by the user.

[0695] Step 7:

[0696] The device presents the received content to the user and records viewing data. Received audio and video data are used as input, and user data such as viewing start time, completion rate, and ad response are returned to the server as output.

[0697] Step 8:

[0698] The server uses the collected viewing data to refine its next delivery strategy. The collected viewing data is used as input, and the output is a new content recommendation list based on user preferences and behavior, optimizing future delivery and advertising strategies.

[0699] (Application Example 1)

[0700] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0701] In today's world, where digital content is becoming increasingly abundant and diverse, it is becoming difficult for users to efficiently acquire information that matches their interests. Furthermore, there is a growing need for easy ways to enjoy content of interest while on the go or amidst busy daily life. In addition, optimizing content delivery based on viewing data, including ad viewing, is becoming increasingly important.

[0702] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0703] In this invention, the server includes means for receiving category information based on the user's interests and using that information to select digital content; means for converting the selected digital content into audio and video formats using speech synthesis technology and video generation technology; and means for transmitting the generated content to a portable display device and providing audio and video. As a result, users can always efficiently receive content based on their interests and easily acquire information while on the go or in their busy daily lives.

[0704] "Category information based on user interests" refers to information that indicates the themes and genres that users are interested in, and is data used for content selection.

[0705] "Digital content" refers to information resources such as audio, video, and text that are stored and used electronically.

[0706] "Speech synthesis technology" is a technology that converts text data into speech data, and is used to generate speech that sounds like a human voice.

[0707] "Video generation technology" is a technology that expresses still images and text as videos, and is a technology for creating visually appealing videos.

[0708] A "visual display device" is a device used to visually display digital content and to provide images and videos to users.

[0709] A "portable display device" is a device for displaying content that can be carried and used by the user.

[0710] "Content viewing data" refers to statistical data about the content that users have played, including viewing time and viewing completion rate.

[0711] The system for implementing this invention mainly consists of a server, a terminal, and a user. The server is equipped with means for selecting digital content based on category information derived from the user's interests. This information is collected by referring to the user's set interest data, and the latest digital content is collected from contracted information sources via API, and the content is analyzed using a generative AI model.

[0712] Furthermore, the server uses speech synthesis technology to convert the selected content into speech that closely resembles a human voice, and then processes it into video format using video generation technology. This process utilizes technologies such as "Google Text-to-Speech" for speech synthesis and "OpenCV" for video generation. The generated media content is then transmitted to a portable display device (e.g., smart glasses), which is a visual display device, and delivered to the user for real-time viewing.

[0713] In this system, the device functions as a platform for users to enjoy highly personalized services while maintaining a portable form. Users can access the latest information and content of interest presented in audio and video formats in a way that is easy to use while on the go. Their viewing and advertising response data is then sent back to the server to help optimize future content and advertising delivery.

[0714] As an example, consider a scenario where a user selects sports news. The server filters the latest match results and player information, providing them with audio along with match highlight videos. A specific prompt might be: "Summarize the latest soccer match and generate audio commentary along with key match highlight videos. The audio should be natural and easy to listen to, and include match results and information on notable players." In this way, users can efficiently view content tailored to their interests and maximize their viewing time.

[0715] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0716] Step 1:

[0717] The server collects the latest digital content through the APIs of contracted information sources. It uses category information and access information for the information sources as input. Based on this, it executes API requests and receives the latest digital content as output.

[0718] Step 2:

[0719] The server analyzes the collected digital content using a generative AI model. The input consists of content data obtained in step 1 and category information based on user interests. The AI ​​model classifies the content by topic and evaluates its importance. The output generates a list of content highly relevant to the user.

[0720] Step 3:

[0721] The server converts selected content into audio data using speech synthesis technology. Text-based content data is used as input. The speech synthesis engine converts the text into natural-sounding speech data and outputs an audio file.

[0722] Step 4:

[0723] The server generates video content using video generation technology along with audio data. It uses audio files and associated images or video footage as input. Video editing software is used to synchronize the audio and video, and the completed video file is output.

[0724] Step 5:

[0725] The server transmits the generated audio and video content to the user's portable display device. Using the user's terminal information and video files as input, the server transmits data to the terminal via the network. Upon receiving the content on the terminal, the user can immediately view it.

[0726] Step 6:

[0727] The terminal provides the received content to the user visually and aurally. It uses audio and video data received from the server as input. Playback is performed through the terminal's built-in display and speaker, completing the presentation on the visual display device. The user then utilizes the content in this step.

[0728] Step 7:

[0729] The server collects user content viewing data. It uses data such as viewing completion rates and ad responses as input. This data is then analyzed to generate output data for optimizing future content delivery and advertising.

[0730] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0731] This invention incorporates an emotion engine into an information distribution system that provides digital content tailored to user interests, thereby realizing a personalized content experience that responds to the user's emotional state. This system mainly consists of a server, terminals, and users, and operates with integrated emotion recognition technology.

[0732] The server retrieves the latest news and articles from contracted media outlets and analyzes them using generative AI. This allows it to select highly relevant digital content, taking into account the user's profile information and emotional state.

[0733] A terminal is a device that allows users to receive information and view content. By incorporating an emotion engine that detects the user's facial expressions and voice into this terminal, the user's emotional state can be recognized in real time.

[0734] User emotional data is collected through facial expression and voice analysis. Based on this data, the server utilizes speech synthesis and video generation technologies to adjust the tone and style of the content. For example, if the user is relaxed, content with a calm tone will be provided; if they are excited, an energetic tone will be selected.

[0735] As a concrete example, suppose a sports-loving user receives the latest match results. If the emotion engine detects that the user is disappointed because their favorite team lost, the server can supplement the information with positive news or uplifting content to capture their interest.

[0736] Furthermore, the device sends this emotional data to the server, which analyzes it along with viewing data to optimize future content delivery. This enriches the experience for each individual user and contributes to the appropriate delivery of advertisements.

[0737] This invention enables the provision of information that resonates with users' emotions, resulting in more effective content experiences and advertising services.

[0738] The following describes the processing flow.

[0739] Step 1:

[0740] Users use their devices to set category information based on their interests and preferences. This sends the user's profile to the server.

[0741] Step 2:

[0742] The server collects the latest digital content using the APIs of contracted media platforms. This information is then subjected to topic analysis and keyword extraction by generative AI.

[0743] Step 3:

[0744] An emotion engine built into the device analyzes the user's facial expressions and voice to recognize the user's emotional state in real time.

[0745] Step 4:

[0746] The server receives user emotion data sent from the emotion engine and selects content according to the user's current emotional state. This includes selecting calming content when the user is relaxed and energetic content when the user is excited.

[0747] Step 5:

[0748] The server generates narration using speech synthesis technology for the selected content and creates video content using video generation technology as needed.

[0749] Step 6:

[0750] The server delivers the generated audio and video content to the user's device. It can also send push notifications at the optimal time, taking into account changes in the user's emotional state.

[0751] Step 7:

[0752] Users view the delivered content using their devices. The emotion engine continues to recognize changes in the user's emotions while they are viewing the content, and this data is sent to the server.

[0753] Step 8:

[0754] The server collects viewing and sentiment data, which is then used to improve future content delivery. This process makes the user experience more personalized.

[0755] (Example 2)

[0756] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0757] Conventional information distribution systems are insufficient in providing personalized content based on users' interests and emotional states, limiting their ability to improve user experience and advertising effectiveness. Furthermore, they are unable to properly analyze emotional states and reflect them in real time, creating a need for optimal information delivery tailored to each user's needs.

[0758] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0759] This invention includes a server that receives information based on the user's interests and emotional state and uses that information to select content; a server that converts the selected content into audio and video formats using speech synthesis and video creation technologies; and a server that analyzes the user's facial expressions and voice data to recognize their emotional state. This enables the provision of personalized content that is tailored to each user's emotions, resulting in a richer experience and more effective advertising delivery.

[0760] "User interests" refer to the content and topics that users are interested in, and serve as the criteria for selecting the information to provide based on these interests.

[0761] "Emotional state" refers to data that indicates the user's emotional condition, such as joy, sadness, excitement, or relaxation.

[0762] "Content selection" refers to the process of choosing appropriate digital information based on the user's interests and emotional state.

[0763] "Speech synthesis technology" refers to the technology that synthesizes human voices using analog or digital data and outputs them in speech format.

[0764] "Video creation technology" refers to the technology used to process visual digital data and create content in video format.

[0765] "Facial and voice data" refers to information that captures the user's facial movements and voice tone, and is used to analyze their emotional state.

[0766] A "contracted source" refers to a source of data from which a server retrieves information for content delivery, and is a medium or platform that is officially granted access rights.

[0767] A "generative AI model" is a type of artificial intelligence that refers to an algorithm capable of generating new data and content based on large amounts of existing data.

[0768] This information distribution system consists of servers, terminals, and users, and is designed to provide personalized content based on users' interests and emotions. The specific roles of each component are shown below.

[0769] The server collects the latest news and articles from contracted information sources using APIs and web scraping techniques. The collected information is analyzed using generative AI models. Specifically, the server uses this information to classify topics and perform sentiment analysis, selecting the most relevant content based on the user's profile information and emotional state. Python libraries and machine learning algorithms are used in this selection process.

[0770] The device is where users view content and is equipped with an emotion engine. This engine has the ability to detect the user's facial expressions and voice in real time and recognize their emotional state. For example, it uses a camera and microphone to capture facial expressions and voice tone, and then analyzes this data with a machine learning model.

[0771] If a user prefers sports news, the server selects the latest match results based on that interest, and the generated content is delivered to the device. If the emotion engine determines that the user is disappointed, the server improves the quality of the user experience by adding positive information to the delivery.

[0772] As a concrete example, consider the following prompt statements: "If the user is relaxed while reading this news article, create a summary in a calm tone," or "Based on the user's emotional response, suggest ways to optimize the selection of content to deliver next." This allows for the delivery of information that matches the user's emotions and interests, resulting in a better content experience.

[0773] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0774] Step 1:

[0775] The server retrieves the latest news and articles from contracted information sources using APIs and web scraping techniques. Input requires access information such as the URL of the information source and an API key. The retrieved information is stored in a database and used in the next processing step. This step specifically involves running a script that automatically retrieves information every hour.

[0776] Step 2:

[0777] The server analyzes the acquired articles and news using a generative AI model. The input is the text data of the articles acquired in step 1. The generative AI model classifies topics, performs sentiment analysis, and labels the content of the articles. The program outputs these results and selects content that matches the user's interests. Specifically, this involves data analysis using scripts that leverage natural language processing techniques.

[0778] Step 3:

[0779] The device uses an emotion engine to collect the user's facial expressions and voice in real time. Inputs include image and audio data acquired from the device's built-in camera and microphone. This data is analyzed using machine learning algorithms to recognize the user's emotional state and output an evaluation result. Specific operations include the execution of facial recognition software and voice analysis software.

[0780] Step 4:

[0781] The server creates personalized content based on the user's emotional state and interests. Its inputs include content information generated in step 2 and the emotional evaluation results obtained in step 3. Using these, it employs speech synthesis and video creation technologies to output digital content in audio and video formats preferred by the user. Specific operations include the process of adding selected music and visual effects.

[0782] Step 5:

[0783] The user views the provided personalized content on their device. The input is the audio and video content delivered from the server in step 4. The output is the user's viewing experience, and their emotional reactions based on that experience. Specific actions include using a video player and providing interactive feedback.

[0784] Step 6:

[0785] The device sends user viewing data and reactions to the server. Input includes log data of the content the user viewed and real-time sentiment data. This data is aggregated and analyzed on the server to optimize the content delivered next time. Specific actions in this step include secure data transfer and statistical analysis within the server.

[0786] (Application Example 2)

[0787] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0788] In digital content distribution, it is essential to consider not only users' interests but also their emotional states in real time to provide appropriate content and advertising experiences. However, conventional systems do not adequately consider users' emotional states, making it difficult to provide personalized experiences based on those emotions. To solve this problem, a content distribution method that is attentive to users' emotions is necessary.

[0789] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0790] This invention includes a server that receives category information based on the user's interests and emotional state, and uses that information to select digital content; a server that converts the selected digital content into audio and video formats using speech synthesis and video generation technologies, and adjusts the tone and style according to the user's emotional state; and a server that delivers the converted audio and video content to the user's terminal and receives feedback based on the emotional state. This makes it possible to provide personalized content and advertising experiences based on the user's emotions.

[0791] "User interests" refer to information that indicates the degree of interest a user has in a particular field or topic, based on their past behavioral history and preferences.

[0792] "Emotional state" refers to the psychological or physiological state that a user is currently experiencing, and is usually determined by analyzing facial expressions and voice.

[0793] "Category information" refers to information used to classify digital content according to a specific theme or topic.

[0794] "Digital content" is a general term for media that can be distributed in digital format, such as audio, video, and text.

[0795] "Speech synthesis technology" is a technology that converts text into speech data and is widely used in assistive technologies for the visually impaired and in artificial intelligence-based dialogue systems.

[0796] "Video generation technology" is a technology that converts abstract data and information into a visual video format, and is used when creating animations and computer graphics (CG).

[0797] "Adjusting tone and style" refers to the process of changing the atmosphere and presentation style according to the content, and is particularly important for providing a personalized experience that responds to the user's emotions.

[0798] "User terminal" refers to equipment or devices used by users to receive or view information, and includes smartphones, tablets, and other similar devices.

[0799] "Feedback" refers to the reactions and opinions received from users, and is information used to improve and optimize the system.

[0800] The system based on this invention consists of a terminal equipped with an emotion engine and an associated server. In order to implement the invention, it is necessary to build a system that can analyze the user's emotional state in real time. The emotion engine built into the terminal analyzes facial expressions and voice and generates the user's emotional data. In this process, emotion recognition technologies such as OpenCV and MediaPipe are utilized to determine emotional states such as relaxation and stress from the facial expression data and voice data.

[0801] The server uses a generative AI model to select digital content, taking into account emotional data and user interests. This generative AI has the ability to retrieve and analyze the latest news and articles from contracted content providers. The analyzed content is then adjusted to match the user's emotional state using speech synthesis and video generation technologies. For example, if the user is feeling tired, relaxing videos with soothing background music will be provided.

[0802] The server then delivers the adjusted content to the device and collects user feedback. This feedback data is used to optimize future content delivery. By monitoring user reactions in real time and having the generative AI model learn from this information, the quality of the content delivered improves over time.

[0803] For example, if a user opens a fitness app on a holiday and facial recognition determines that the user is in need of energy, the server can select and deliver an energetic training video to boost their spirits.

[0804] An example of a prompt message would be, "Please provide fitness content that matches the perceived emotion. If you are relaxed, please select yoga; if you are stressed, please select full-body stretches." This system operates on user devices such as smartphones, enabling flexible enhancement of the user's content experience.

[0805] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0806] Step 1:

[0807] The device collects the user's facial expressions and voice in real time using an emotion engine. The input is the user's video and audio data. This data is processed into emotional state data using emotion recognition libraries such as OpenCV and MediaPipe. The output is the user's emotional state (e.g., relaxed, stressed).

[0808] Step 2:

[0809] The server receives emotional states transmitted from the terminal and pre-stored user interest data as input. This input is fed to a generative AI model, which selects relevant digital content based on conditions indicated by prompts. News and articles obtained in real-time from contracted media outlets are used for content selection. The output is digital content that matches the user's emotional state.

[0810] Step 3:

[0811] The server converts selected digital content into audio and video formats using speech synthesis and video generation technologies. During this process, it adjusts the tone and style according to the emotional state. The input is the selected digital content. The output is emotion-sensitive audio and video data for delivery to the user.

[0812] Step 4:

[0813] The server delivers the converted audio and video content to the user's device. The device receives this data as input and plays it back as content for the user. The output obtained from this playback is feedback data such as the user's satisfaction level and emotional changes.

[0814] Step 5:

[0815] The device collects data about the user's content viewing and sends it to the server. This data includes playback time, skipped sections, and the user's emotional state. The input is the user's viewing behavior and emotional data, and the output is analytical data sent to the server.

[0816] Step 6:

[0817] The server uses feedback and collected data to optimize future content delivery. This process involves a generative AI model learning to provide even more personalized content. The input is user feedback data, and the output is the improved content delivery strategy.

[0818] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0819] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0820] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0821] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0822] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0823] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0824] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0825] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0826] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0827] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0828] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0829] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0830] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0831] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0832] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0833] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0834] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0835] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0836] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0837] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0838] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[0839] The following is further disclosed regarding the embodiments described above.

[0840] (Claim 1)

[0841] A means of receiving category information based on user interests and using that information to select digital content,

[0842] A means for converting selected digital content into audio and video formats using speech synthesis technology and video generation technology,

[0843] A means for delivering converted audio and video content to the user's terminal,

[0844] A means of collecting user content viewing and advertising viewing data and optimizing future content delivery based on this data,

[0845] An information distribution system that includes this.

[0846] (Claim 2)

[0847] The system according to claim 1, which automatically retrieves the latest articles and news from contracted media and analyzes them.

[0848] (Claim 3)

[0849] The system according to claim 1, which sends a push notification to the user's terminal and adjusts the timing of content playback.

[0850] "Example 1"

[0851] (Claim 1)

[0852] A means of receiving information based on user interests, analyzing digital information using a generative AI model, and selecting highly relevant information,

[0853] A means for converting selected digital information into audio and video formats using speech synthesis technology and video generation technology,

[0854] A means for transmitting the converted audio and video information to the user's terminal,

[0855] A means of collecting user information viewing and advertising viewing data and optimizing future information delivery based on this data,

[0856] A system that includes this.

[0857] (Claim 2)

[0858] The system according to claim 1, which automatically retrieves the latest articles and news from contracted information sources and analyzes them using a generative AI model.

[0859] (Claim 3)

[0860] The system according to claim 1, which sends a notification to the user's terminal and adjusts the timing of information playback.

[0861] "Application Example 1"

[0862] (Claim 1)

[0863] A means of receiving category information based on user interests and using that information to select digital content,

[0864] A means for converting selected digital content into audio and video formats using speech synthesis technology and video generation technology,

[0865] Means for distributing converted audio and video content via a visual display device,

[0866] A means of collecting user content viewing and advertising viewing data and optimizing future content delivery based on this data,

[0867] A means for transmitting the generated content to a portable display device and providing audio and video,

[0868] A system that includes this.

[0869] (Claim 2)

[0870] The system according to claim 1, which automatically retrieves the latest articles and news from contracted sources and analyzes them.

[0871] (Claim 3)

[0872] The system according to claim 1, which sends a notification to a visual display device and adjusts the playback timing of the content.

[0873] "Example 2 of combining an emotion engine"

[0874] (Claim 1)

[0875] A means of receiving information based on the user's interests and emotional state, and using that information to select content,

[0876] A means for converting selected content into audio and video formats using speech synthesis and video creation technologies,

[0877] A means for delivering converted audio and video content to the user's terminal,

[0878] A means of collecting user content viewing and advertising viewing data and optimizing future content delivery based on this data,

[0879] A means of analyzing the user's facial expressions and voice data to recognize their emotional state,

[0880] A system that includes this.

[0881] (Claim 2)

[0882] The system according to claim 1, which automatically acquires the latest data from a contracted information source and analyzes it using a generative AI model.

[0883] (Claim 3)

[0884] The system according to claim 1, which adjusts the tone and style of content based on the emotional state of the user.

[0885] "Application example 2 when combining with an emotional engine"

[0886] (Claim 1)

[0887] A means of receiving category information based on users' interests and emotional states, and using that information to select digital content,

[0888] A means for converting selected digital content into audio and video formats using speech synthesis and video generation technologies, and for adjusting the tone and style according to the user's emotional state,

[0889] A means for delivering converted audio and video content to the user's terminal and receiving feedback based on emotional state,

[0890] In addition to user content viewing and advertising viewing data, a means of collecting sentiment data and optimizing future content delivery based on this data,

[0891] A system that includes this.

[0892] (Claim 2)

[0893] The system according to claim 1, which automatically retrieves the latest articles and news from contracted media and analyzes them while taking into account the user's emotional state.

[0894] (Claim 3)

[0895] The system according to claim 1, which sends a push notification to the user's terminal and adjusts the playback timing and tone of the content to match the user's emotional state. [Explanation of Symbols]

[0896] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. A means of receiving category information based on user interests and using that information to select digital content, A means for converting selected digital content into audio and video formats using speech synthesis technology and video generation technology, A means for delivering converted audio and video content to the user's terminal, A means of collecting user content viewing and advertising viewing data and optimizing future content delivery based on this data, An information distribution system that includes this.

2. The system according to claim 1, which automatically retrieves the latest articles and news from contracted media and analyzes them.

3. The system according to claim 1, which sends a push notification to the user's terminal and adjusts the timing of content playback.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A