System
The system addresses the lack of personalization in audio guides by generating detailed and emotionally nuanced audio descriptions based on user preferences and feedback, enhancing the movie-watching experience for visually impaired and movie enthusiasts.
Patent Information
- Application Number
- JP2024117347
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-22
- Publication Date
- 2026-02-03
AI Technical Summary
Conventional audio guides for visually impaired individuals and movie enthusiasts lack personalization, failing to provide detailed scene descriptions and emotional insights, limiting the movie-watching experience.
A system that collects user preferences and viewing history, analyzes movie data using computer vision and natural language processing, generates personalized audio guides with a speech synthesis engine, and updates profiles based on user feedback to enhance the movie-watching experience.
Provides a tailored and enriched movie-watching experience for diverse audiences, including visually impaired users, by delivering detailed scene descriptions and emotional insights.
Smart Images

Figure 2026016257000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] Conventional audio guides primarily convey on-screen information to the visually impaired and do not include other elements that can enrich the movie-watching experience. For example, viewers who cannot fast-forward through silent scenes or understand the emotions behind the dialogue find it difficult to fully enjoy the movie. Furthermore, as film and TV discussions become more popular, more viewers are seeking deeper insights into the story, but current audio guides are unable to provide the necessary information. There is a need to solve these issues and make the movie-watching experience more enjoyable and with greater added value. [Means for solving the problem]
[0005] The present invention provides a system that includes a means for collecting user preferences and viewing history and creating a profile, and a means for acquiring and analyzing basic movie information, reviews, interviews, scripts, and subtitles. The system also includes a means for generating a personalized audio guide based on the user profile, a means for generating the generated audio guide using a speech synthesis engine, and a means for delivering the generated audio guide to the user's device. The system also includes a means for analyzing the collected feedback and updating the user profile, and a means for analyzing visual information using computer vision technology to identify important scenes and character appearances. Using these means, the movie-watching experience can be personalized according to the preferences of each individual user, providing a high-value viewing experience not only for the visually impaired but also for a diverse audience.
[0006] A "user" is an individual who uses this system to watch movies.
[0007] A "profile" is a data set that includes a user's preferences, viewing history, and other attribute information.
[0008] "Basic movie information" refers to basic data such as the movie title, genre, cast, director, and running time.
[0009] A "review" is a piece of writing or commentary that includes an evaluation or opinion about a released movie.
[0010] An "interview article" is an article that includes information obtained from interviews with the film's cast, director, etc.
[0011] A "script" is a movie script or screenplay, a document that describes the lines of characters and details of scenes.
[0012] "Subtitles" are text that displays dialogue and explanations from a movie, usually at the bottom of the screen.
[0013] "Personalization" refers to customizing content and functionality according to a user's individual preferences and needs.
[0014] An "audio description" is an audio explanation or commentary provided to complement the visual and auditory senses when watching a movie.
[0015] A "speech synthesis engine" is software or hardware for converting text data into voice data.
[0016] A "terminal" is a device (e.g., smartphone, tablet, PC, etc.) that a user uses to watch a movie.
[0017] "Feedback" refers to information such as evaluations, opinions, and impressions that users provide to the system.
[0018] "Computer vision technology" is a technology in which a computer analyzes image and video data to recognize objects and scenes.
[0019] A "scene" is a part of a film that represents a series of related events or situations.
[0020] A "character" is a person or role that appears in a film. [Brief explanation of the drawings]
[0021] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5]FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0022] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0023] First, the terms used in the following description will be explained.
[0024] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0025] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0026] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0027] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0028] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0029] [First embodiment]
[0030] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0031] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0032] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0033] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0034] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0035] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0036] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0037] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0038] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0039] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0040] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0041] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0042] This invention is a system for generating a personalized movie audio guide that is tailored to the user's situation and preferences. The program processing flow of this system will be explained below with specific examples.
[0043] Program processing overview
[0044] 1. Collecting User Information
[0045] When a user first uses the system, the server asks the user to register an account. The user enters basic information such as name, age, language preference, viewing history, movie preferences, etc. The server then creates a user profile and stores it in a database.
[0046] Examples:
[0047] User A creates an account and registers in his profile information that he is visually impaired, likes mystery movies, and seeks deep insights.
[0048] 2. Obtaining movie information
[0049] The server retrieves basic information about the movie selected by the user from online databases and APIs, as well as public reviews and interviews with the cast and director. It also retrieves the movie's script and subtitle files and prepares them as data for analysis.
[0050] 3. Audio guide generation
[0051] The server analyzes the collected movie data, recognizing important scenes and character appearances through scene analysis, and extracts important dialogue and emotional nuances using NLP technology.
[0052] It then generates personalized audio description text based on the user profile, providing detailed image descriptions for the visually impaired and in-depth insights and background information for movie buffs.
[0053] Examples:
[0054] For User A, a detailed audio guide is generated for each scene, and text is created that provides detailed explanations of the main cast's facial expressions and emotional changes, as well as the foreshadowing of the mystery.
[0055] The server then passes this text to a speech synthesis engine to generate a voice description that sounds more human, adding natural intonation and accents to make the description easier to understand.
[0056] 4. Audio guide distribution
[0057] The server provides the generated audio guide to the user's terminal in streaming or download format.
[0058] The terminal can play the provided audio description as secondary audio when the user watches the movie.
[0059] 5. Collecting User Feedback
[0060] The server collects feedback from users after they use the audio guide. Users can provide feedback on the quality and content of the guide, as well as areas for improvement. The server then adds this information to a profile and uses it to generate future guides.
[0061] Examples:
[0062] User A provides feedback stating that the emotional commentary section was very helpful and that he would like more specific scene descriptions. This feedback is reflected in the generation of guides for other movies he watches later.
[0063] Overall flow and summary
[0064] This system comprehensively handles everything from user registration to collecting and analyzing movie information, generating and distributing personalized audio guides, and collecting feedback. This allows for a movie-watching experience tailored to each user's needs, providing a high-value-added viewing experience not only for the visually impaired but also for a diverse audience. The introduction of this system is expected to open up new ways to enjoy movie-watching and provide a richer entertainment experience for many people.
[0065] The processing flow will be explained below.
[0066] Step 1: Registering a user
[0067] Users access the system and enter basic information such as name, age, language preference, viewing history, and movie preferences to create an account.
[0068] The server receives the entered information and creates and stores a user profile in a database.
[0069] Step 2: Update your profile
[0070] Users can update their profile information as their viewing history or preferences change.
[0071] The server stores the updated information in a database to keep the profile up to date.
[0072] Step 3: Get movie information
[0073] The user selects the movie they want to watch.
[0074] The server retrieves basic information such as the movie title, genre, cast, director, and running time from online databases and APIs.
[0075] The server also collects reviews, interviews, scripts, and subtitle files about the film.
[0076] Step 4: Analyze the scene
[0077] The server uses computer vision technology to analyze the movie's video data and identify key scenes and scenes in which characters appear.
[0078] The server uses NLP technology to analyze scripts and subtitle files to extract key dialogue and emotional nuances.
[0079] Step 5: Generate a personalized guide
[0080] The server decides how much detail to include in the audio description based on the user profile: detailed visual descriptions for the visually impaired, for example, or in-depth insights and background information for movie buffs.
[0081] The server creates the text for the audio guide based on the analysis results.
[0082] Step 6: Text-to-Speech
[0083] The server passes the generated audio guide text to a speech synthesis engine, which generates a voice that sounds like a human voice, with natural intonation and accent.
[0084] Step 7: Distributing the audio guide
[0085] The server provides the generated audio guide to the user's terminal in streaming or download format.
[0086] The terminal stores the provided audio guide and keeps it available for the user to use while watching the movie.
[0087] Step 8: Gather feedback
[0088] After watching a movie, users provide feedback on the quality and content of the audio description.
[0089] The server collects the provided feedback and stores it in a database, which is used to generate future guides.
[0090] Step 9: Improve your profile
[0091] The server analyzes the collected feedback and updates the user profile. This updated information is reflected in the next guide generation, providing an audio guide that is more tailored to the user's preferences.
[0092] Through the specific processing steps described above, the system provides a personalized movie-watching experience for each user, and meets the needs of visually impaired people and diverse audiences.
[0093] Example 1
[0094] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0095] Conventional movie audio guide systems have difficulty providing a fully personalized guide for the visually impaired or users who prefer specific movie genres. This limits the movie-watching experience, especially for users who cannot visually confirm the details of each movie scene or the emotional expressions of characters. Furthermore, it is difficult to effectively reflect user feedback and improve the guide, and further technical solutions are needed to increase user satisfaction.
[0096] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0097] In this invention, the server includes means for collecting user preferences and viewing history and creating a profile, means for acquiring and analyzing basic movie information, reviews, interviews, scripts, and subtitles, means for generating a personalized audio guide based on the user profile, means for analyzing the collected movie data and extracting important dialogue and emotional nuances using scene analysis and natural language processing techniques, means for generating the generated audio guide using a speech synthesis engine, means for delivering the generated audio guide to the user's terminal, and means for collecting user feedback and reflecting it in the profile. This allows for the provision of a detailed and personalized movie guide tailored to the user's needs, enabling a variety of users, including visually impaired people, to enjoy a richer movie-watching experience.
[0098] A "user profile" is a collection of data created based on a user's preferences, viewing history, basic information, and so on.
[0099] "Basic movie information" is basic information about a movie, such as the movie title, director, cast, and release date.
[0100] A "review" is text data that describes the ratings, impressions, and opinions posted by people who have watched a movie.
[0101] An "interview article" is text data that describes the contents of interviews with cast members, directors, and others involved in the film.
[0102] A "script" is text data that describes a movie script or screenplay.
[0103] "Subtitles" are text data that contains movie lines and explanations in written form.
[0104] "Scene analysis" is the process of analyzing each scene in a film to identify important scenes and characters.
[0105] "Natural language processing" is a technology that allows computers to understand, analyze, and generate human language.
[0106] An "audio description" is a guide that supplements visual information and provides audio information about the contents of a film.
[0107] A "speech synthesis engine" is a software system that generates natural-sounding speech based on text data.
[0108] "Feedback" refers to information such as evaluations, opinions, and impressions collected from users.
[0109] "Personalization" refers to a state that is customized based on the characteristics and preferences of an individual user.
[0110] The present invention is a system for generating a personalized movie audio guide that is adapted to the user's situation and preferences. Specific embodiments of this system will be described below.
[0111] System configuration
[0112] The system mainly consists of a server and a user device. The server creates user profiles, collects and analyzes movie data, generates audio guides, and collects feedback. The user device selects movies, plays audio guides, and provides feedback. Natural language processing technology and speech synthesis engines are used as software. Specific examples include AWS's Amazon Polly and Google Cloud Text-to-Speech.
[0113] Collecting user information
[0114] When a user first uses the system, the server asks the user to register an account. The user enters basic information such as name, age, language preference, viewing history, movie preferences, etc. The server then creates a user profile and stores it in a database.
[0115] Examples:
[0116] User A creates an account and registers in his profile information that he is visually impaired, likes mystery movies, and seeks deep insights.
[0117] Get movie information
[0118] The server retrieves basic information about the movie selected by the user from online databases and APIs, as well as public reviews and interviews with the cast and director. It also retrieves the movie's script and subtitle files and prepares them as data for analysis.
[0119] Examples:
[0120] When User A selects the movie "Sherlock Holmes," the server collects public information, reviews, scripts, etc. about the movie.
[0121] Audio guide generation
[0122] The server analyzes the collected movie data, recognizing important scenes and scenes in which characters appear through scene analysis, and extracting important lines and emotional nuances using natural language processing technology.
[0123] It then generates personalized audio description text based on the user profile, providing detailed visual descriptions for the visually impaired and in-depth insights and context for movie buffs. The generated text is then passed to a speech synthesis engine (e.g., AWS's Amazon Polly, Google Cloud Text-to-Speech) to generate a human-like audio description.
[0124] Examples:
[0125] A detailed audio guide for each scene is generated for User A, and text is created that provides detailed explanations of the facial expressions and emotional changes of the main cast, as well as hints at the mystery. An example of a prompt would be, "User A is visually impaired and loves mystery movies. Please generate a detailed audio guide for each scene of the movie 'Sherlock Holmes.' Please provide detailed explanations of the facial expressions and emotional changes of the main cast, as well as hints at the mystery."
[0126] Audio guide distribution
[0127] The server provides the generated audio guide to the user's terminal in a streaming or download format, and the terminal can play the provided audio guide as a secondary audio when the user watches the movie.
[0128] Collecting user feedback
[0129] The server collects feedback from users after they use the audio guide. Users can provide feedback on the quality and content of the guide, as well as areas for improvement. The server then adds this information to a profile and uses it to generate future guides.
[0130] Examples:
[0131] User A provides feedback stating that the emotional commentary section was very helpful and that he would like more specific scene descriptions. This feedback is reflected in the generation of guides for other movies he watches later.
[0132] In this way, the system can comprehensively handle everything from user registration to collecting and analyzing movie information, generating and distributing personalized audio guides, and collecting feedback, making it possible to provide a high-value movie-watching experience for a diverse audience, including the visually impaired.
[0133] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0134] Step 1:
[0135] Account Registration
[0136] When a user first accesses the system, the server displays an account registration form. The user enters basic information such as name, age, language preference, viewing history, and movie preferences, and then presses the submit button. The server receives the input data and creates a user profile. The created user profile is stored in a database.
[0137] Input: Basic information entered by the user (name, age, language preference, viewing history, movie preferences)
[0138] Output: User profile stored in the database
[0139] Specific behavior:
[0140] The server displays a registration form through a user interface. The user fills in the form and presses the submit button. The server takes the input data, generates a user profile, and stores it in a database.
[0141] Step 2:
[0142] Movie Selection
[0143] The user selects a movie, enters the movie title in the search box on the user terminal, and presses the search button, and the selected movie title is sent to the server.
[0144] Input: The movie title selected by the user
[0145] Output: Movie title sent to the server
[0146] Specific behavior:
[0147] The user terminal displays a movie title entry form, and when the user enters a title and presses the search button, the selection is sent to the server.
[0148] Step 3:
[0149] Movie data collection
[0150] Based on the received movie title, the server uses online databases and APIs to collect related data such as basic movie information, reviews, interviews, scripts, subtitles, etc. The collected data is then stored for analysis.
[0151] Input: Movie title
[0152] Output: Data stored on the server, including basic movie information, reviews, interviews, scripts, subtitles, etc.
[0153] Specific behavior:
[0154] The server calls a specific API, sends the movie title, and retrieves the relevant data, which is then stored in a database for analysis.
[0155] Step 4:
[0156] Data analysis and audio guide generation
[0157] The server analyzes the stored movie data, performs scene analysis to identify key scenes and character appearances, and uses natural language processing techniques to extract key lines and emotional nuances. Based on the extracted data, it generates personalized audio description text that matches the user profile. The generated text is then passed to a speech synthesis engine.
[0158] Input: saved movie data, user profile
[0159] Output: Personalized audio guide text and audio files
[0160] Specific behavior:
[0161] The server runs an algorithm that analyzes the movie scenes to determine their importance. It then uses natural language processing techniques to analyze the dialogue and extract key elements. It then inputs the prompt into a generative AI model to generate personalized text in a specific format. Finally, it passes the text to a speech synthesis engine to generate an audio file.
[0162] Step 5:
[0163] Audio guide distribution
[0164] The server provides the generated audio guide to the user's device in a streaming or download format, and the device can play the provided audio guide as a secondary audio when the user watches the movie.
[0165] Input: Personalized audio guide audio file
[0166] Output: Audio guide delivered to the user's device
[0167] Specific behavior:
[0168] The server uploads the generated audio guide to cloud storage and generates an access URL, which the user's device uses to play or download the audio guide.
[0169] Step 6:
[0170] Gathering feedback
[0171] After the user has finished watching, the server provides a feedback form. The user can enter their evaluation of the quality and content of the guide and suggestions for improvement in the feedback form and submit it. The server then reflects the collected feedback in the profile and uses it for generating audio guides from the next time onwards.
[0172] Input: User feedback
[0173] Output: Updated user profile
[0174] Specific behavior:
[0175] After the user has finished watching, the server displays a feedback form. The user enters their feedback in the form and presses the submit button. The server receives the feedback data and updates the user profile.
[0176] (Application example 1)
[0177] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0178] Conventional movie audio guide systems have not adequately provided personalized guides based on the viewer's individual preferences and viewing history. As a result, it has been difficult to provide an optimal viewing experience for visually impaired people or users who prefer specific movie genres. Furthermore, continuous improvement is required due to insufficient improvement in the quality of generated audio guides and profile updates based on user feedback.
[0179] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0180] In this invention, the server includes means for collecting user preferences and viewing history and creating a profile, means for acquiring and analyzing basic movie information, reviews, interviews, scripts, and subtitles, means for generating a personalized audio guide based on the user profile, means for generating the generated audio guide using a speech synthesis engine, means for delivering the generated audio guide to the user's terminal, means for generating text for the personalized audio guide based on the user profile and movie data using a generative AI model, means for calling the generative AI model using a prompt sentence in generating the audio guide, and means for converting the generated audio guide into a format suitable for smart devices and delivering it, thereby enabling the provision of an audio guide optimized for each user.
[0181] "User preferences" refers to information that indicates the user's tastes in movie genres and specific movies.
[0182] "Viewing history" refers to data that records the history of movies and dramas that a user has watched.
[0183] A "profile" refers to a collection of information that compiles a user's personal information, viewing history, preferences, etc.
[0184] "Basic movie information" refers to basic data such as the movie title, director, cast, release year, and genre.
[0185] A "review" is a text that shows criticism or evaluation of a movie, and is a source of information that allows users to learn the opinions of other viewers and experts.
[0186] An "interview article" refers to a written record of an interview conducted with a film's cast, director, or production staff.
[0187] A "script" refers to a film script or screenplay, a document that contains the actual dialogue and scene details of a film.
[0188] "Subtitles" refers to text that visually displays dialogue or narration in a film.
[0189] "Personalized audio description" refers to audio description that is customized based on a user's individual profile.
[0190] A "speech synthesis engine" is a program for converting text into voice data.
[0191] "Generative AI Model" refers to a model that uses artificial intelligence technology to generate personalized guides based on a user's profile and movie data.
[0192] A "prompt" is an instruction or question input to a generative AI model, and is text used to control the model's output.
[0193] "Smart devices" refers to electronic devices with advanced functions such as smartphones and smart glasses.
[0194] This invention is a system that generates and distributes personalized audio guides that are tailored to the user's circumstances and preferences. The program processing flow of this system is explained below with specific examples.
[0195] 1. Collection of User Information
[0196] When a user first uses the system, the server asks the user to register an account. The user enters basic information such as name, age, language preference, viewing history, movie preferences, etc. The server then creates a user profile and stores it in a database.
[0197] 2. Obtaining movie information
[0198] The server retrieves basic information about the movie selected by the user from online databases and APIs, including the movie title, director, cast, reviews, interviews, script, and subtitle files, and prepares this data for analysis.
[0199] 3. Audio guide generation
[0200] The server analyzes the collected movie data and uses NLP techniques to extract key scenes and emotional nuances. It then generates personalized audio description text based on the user profile. A generative AI model is used to generate the text, and a prompt is input to the model. For example, a prompt might be in the form: "User profile: {Name: User A, Age: 30, Language: Japanese, Favorite movie genre: Mystery} Movie data: {Title: Inception, Director: Christopher Nolan, Cast: {Leonardo DiCaprio: Dom Cobb}, Review: An extremely complex and intricately structured story...} Please generate a personalized movie description."
[0201] After the audio description text is generated, the server passes it to a speech synthesis engine to generate a human-like audio description, adding natural intonation and accents to make it easier to listen to.
[0202] 4. Audio guide distribution
[0203] The generated audio guide is provided to the user's smart device, such as a smartphone or smart glasses, via streaming or download. The server converts the audio guide into a format optimized for the user's device and delivers it appropriately.
[0204] 5. Collecting User Feedback
[0205] Users can provide feedback after using the audio guide. The server collects and analyzes this feedback. The collected feedback is reflected in the user profile and is used to generate audio guides in the future.
[0206] This system makes it possible to provide a personalized movie audio guide tailored to the user's situation and preferences, providing a richer movie-watching experience.
[0207] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0208] Step 1:
[0209] The server requires a user to register an account when using the system for the first time. The user enters basic information such as name, age, language setting, viewing history, and movie preferences. Based on this information, the server creates a user profile and stores it in a database. Input data: name, age, language setting, viewing history, and movie preferences. Output data: user profile.
[0210] Step 2:
[0211] When a user selects a movie to watch, the server retrieves basic information about the movie, including the movie title, director, cast, reviews, interviews, script, and subtitle files. Data is collected from online databases and APIs. Input data: movie title. Output data: movie basic information, reviews, interviews, script, and subtitle files.
[0212] Step 3:
[0213] The server analyzes the collected movie data using NLP technology, extracting key scenes and emotional nuances. Input data: Basic movie information, reviews, interviews, scripts, and subtitle files. Output data: Analyzed movie data.
[0214] Step 4:
[0215] Using the generative AI model, the server generates personalized audio guide text based on the user profile and analyzed movie data. A prompt is input to the generative AI model. For example, a prompt in the following format might be used: "User profile: {Name: User A, Age: 30, Language: Japanese, Favorite movie genre: Mystery} Movie data: {Title: Inception, Director: Generic name, Cast: {Starring actor: Lead role}, Review: Extremely complex and intricately structured story} Please generate a personalized movie audio guide." Input data: User profile, analyzed movie data. Output data: Audio guide text.
[0216] Step 5:
[0217] The server passes the generated audio description text to a speech synthesis engine to generate an audio description that sounds like a human voice, with natural intonation and accent. Input data: Audio description text. Output data: Audio data.
[0218] Step 6:
[0219] The server provides the generated audio guide to the user's smart device, such as a smartphone or smart glasses, in streaming or download format. The audio guide is converted into a format optimized for the user's device. Input data: Audio data. Output data: Audio data in a format suitable for the smart device.
[0220] Step 7:
[0221] Users can provide feedback after using the audio guide. The server collects and analyzes this feedback. The collected feedback is reflected in the user profile and is used to generate audio guides from next time onwards. Input data: User feedback. Output data: Updated user profile.
[0222] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0223] This invention is a system that generates a personalized movie audio guide tailored to the user's situation and preferences, and further combines an emotion engine to recognize the user's emotions and dynamically adjust the content of the audio guide. The program processing flow of this system is explained below with concrete examples.
[0224] Program processing overview
[0225] 1. Collecting User Information
[0226] When a user first uses the system, the server asks the user to register an account. The user enters basic information such as name, age, language preference, viewing history, and movie preferences, and also asks for consent to emotion recognition. Based on this information, the server creates a user profile and stores it in a database.
[0227] Examples:
[0228] User A creates an account and registers in his profile information that he is visually impaired, likes mystery movies, and seeks deep insights. He also agrees to emotion recognition.
[0229] 2. Obtaining movie information
[0230] The user selects a movie they want to watch. The server retrieves basic information about the movie, such as the title, genre, cast, director, and running time, from online databases and APIs. It also collects public reviews, interviews with the cast and director, scripts, and subtitle files.
[0231] 3. Scene Analysis
[0232] The server uses computer vision technology to analyze the film's video data to identify key scenes and character appearances, and NLP technology to analyze the script and subtitle files to extract key dialogue and emotional nuances.
[0233] 4. User Emotion Recognition by Emotion Engine
[0234] The server uses an emotion engine to recognize emotions from the user's facial expressions and voice in real time. While the user is watching a movie, the device uses a camera and microphone to capture the user's facial expressions and voice and transmits the data to the server.
[0235] Examples:
[0236] While User A is watching a movie, the emotion engine recognizes that he is crying at a touching scene. This information is sent to the server.
[0237] 5. Generate and refine personalized guides
[0238] The server generates personalized audio guide text based on the user profile and the recognition results of the emotion engine. Furthermore, it dynamically adjusts the content of the audio guide according to the recognized emotion. For example, if the user is emotional, it adds commentary that reflects that emotion.
[0239] Examples:
[0240] If user A is in tears during a touching scene, the audio guide for that scene will add commentary that delves deeper into the character's emotions.
[0241] 6. Speech Synthesis
[0242] The server then passes the generated audio guide text to a speech synthesis engine, which generates a voice that sounds like a human voice, with natural intonation and accent.
[0243] 7. Audio guide distribution
[0244] The server provides the generated audio guide to the user's terminal in streaming or download format, and the terminal stores the provided audio guide so that the user can use it while watching the movie.
[0245] 8. Gathering Feedback
[0246] After watching a movie, users provide feedback on the quality and content of the audio guide. The server collects the feedback and stores it in a database. The feedback is used to generate future guides.
[0247] 9. Profile Improvements
[0248] The server analyzes the collected feedback and updates the user profile. This updated information is reflected in the next guide generation, providing an audio guide that is more tailored to the user's preferences.
[0249] Overall flow and summary
[0250] This system comprehensively handles everything from user registration to collecting and analyzing movie information, real-time emotion recognition using an emotion engine, generating and adjusting personalized audio guides, delivering them, and collecting feedback. This provides a movie-watching experience tailored to each user's needs and emotions, and addresses the needs of visually impaired people and a diverse audience. The introduction of this system is expected to open up new ways to enjoy movie-going and provide a richer entertainment experience for many people.
[0251] The processing flow will be explained below.
[0252] Step 1: Registering a user
[0253] Users access the system and enter their name, age, language preference, viewing history, favorite movie genre, preferred level of visual and auditory assistance, and consent to emotion recognition on the account creation screen.
[0254] The server receives the information entered by the user and stores it in a database as a new user profile.
[0255] Step 2: Update your profile
[0256] Users can update their profile to reflect changes to their viewing history or new movie preferences by entering and saving changes on the profile page.
[0257] The server updates and saves the changes made by the user to the profile information in the database.
[0258] Step 3: Get movie information
[0259] The user selects the movie they want to watch from within the system.
[0260] The server retrieves basic information about the selected movie, such as title, genre, cast, director, and running time, from online databases or APIs.
[0261] The server also collects reviews of the film, interviews with the cast and director, scripts, and subtitle files.
[0262] Step 4: Analyze the scene
[0263] The server uses computer vision technology to analyze the movie's video data and identify key scenes and scenes in which characters appear.
[0264] The server uses NLP technology to analyze scripts and subtitle files to extract key dialogue and emotional nuances.
[0265] Step 5: Recognizing user emotions with the emotion engine
[0266] The device uses a built-in camera and microphone to capture the user's facial expressions and voice in real time while the movie is playing.
[0267] The terminal sends the captured data to an emotion engine to recognize the user's emotional state.
[0268] The emotion engine (server) analyzes the received data, identifies the user's emotion (such as joy, sadness, surprise, etc.), and generates an emotion status.
[0269] Step 6: Generate and refine your personalized guide
[0270] The server generates personalized audio guide text based on the user profile and the recognition results of the emotion engine.
[0271] The server dynamically adjusts the content of the audio guide depending on the recognized emotion, for example, by adding explanations to calm the user or by emphasizing emotional background information for moving scenes.
[0272] Examples:
[0273] If user A is in tears during a moving scene in a movie, the server can incorporate additional emotional context into the audio description to enhance the viewing experience of that scene.
[0274] Step 7: Text-to-Speech
[0275] The server then passes the generated audio guide text to a speech synthesis engine, which generates a voice that sounds similar to a natural human voice, with natural intonation and accent.
[0276] Step 8: Distributing the audio guide
[0277] The server provides the generated audio guide to the user's terminal in streaming or download format.
[0278] The terminal stores the provided audio description and prepares it for playback as secondary audio during movie playback.
[0279] Step 9: Gather feedback
[0280] After watching a movie, users can enter their feedback on the quality and content of the audio guide on a feedback page within the system.
[0281] The server receives the provided feedback and stores it in a database, which is used to generate subsequent guides.
[0282] Step 10: Refine your profile
[0283] The server analyzes the collected feedback and updates the user profile. This updated information is reflected in the next guide generation, providing an audio guide that is more tailored to the user's preferences.
[0284] Overall flow and summary
[0285] This process comprehensively covers user registration, movie information collection, real-time emotion recognition using an emotion engine, personalized and tailored audio guide generation and delivery, and feedback collection. By dynamically reflecting user emotions, the system can promote a more personalized movie-watching experience and cater to a diverse audience.
[0286] Example 2
[0287] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0288] There is a demand for a variety of viewers, including the visually impaired, to enrich their movie experience and provide personalized guidance that reflects the user's individual emotions and preferences. Conventional systems have difficulty providing guidance that reflects the user's emotions in real time, resulting in only a uniform guidance that makes it difficult to increase user satisfaction.
[0289] The specification process by the specification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for collecting a user's preferences and viewing history and creating a profile; means for acquiring and analyzing basic movie information, reviews, interviews, scripts, and subtitles; means for analyzing visual information using computer vision technology and identifying important scenes and scenes in which characters appear; means for capturing the user's facial expressions and voice through a terminal while watching a movie and recognizing the user's emotions using an emotion engine; means for generating and dynamically adjusting a personalized audio guide based on the user profile and the recognized emotions; means for generating the generated audio guide using a speech synthesis engine; and means for delivering the generated audio guide to the user's terminal. This makes it possible to provide a personalized movie experience tailored to the user's emotions and preferences.
[0290] "User preferences" refer to the tastes and preferences that a user has for a particular genre or content.
[0291] A "viewing history" is a record of movies and programs a user has watched.
[0292] A "profile" is a collection of personal information such as a user's name, age, language preference, viewing history, and movie preferences.
[0293] "Basic movie information" refers to information such as the movie title, genre, cast, director, and screening time.
[0294] A "review" is a piece of writing that expresses an opinion or rating about a movie.
[0295] An "interview article" is a written version of interviews with cast members, directors, and others involved in the film.
[0296] A "script" is a screenplay that contains the dialogue and direction of a movie.
[0297] "Subtitles" are the written versions of the audio in a movie that are displayed on the screen.
[0298] "Computer vision technology" is a technology that allows computers to analyze the content of images and videos.
[0299] An "emotion engine" is a software technology that recognizes emotions by analyzing a user's facial expressions and voice.
[0300] A "speech synthesis engine" is a technology that generates speech by adding natural intonation and accent based on text information.
[0301] A "terminal" is a device that a user uses to watch a movie, including a smartphone, tablet, or PC.
[0302] The present invention provides a system for generating a personalized movie audio guide based on a user's preferences, viewing history, and emotions, thereby improving the user's movie-watching experience. Specific embodiments of the present invention will be described below.
[0303] Hardware and software used
[0304] server:
[0305] Database: Used to store user profiles, movie information, feedback, etc.
[0306] Computer vision technology: Used to analyze film footage and identify key scenes and character appearances.
[0307] Emotion engine: Used to analyze the user's facial expressions and voice to recognize emotions in real time.
[0308] Speech synthesis engine: Used to convert text to speech for personalized audio descriptions.
[0309] Device:
[0310] Camera: Used to capture the user's facial expressions.
[0311] Microphone: Used to capture the user's voice.
[0312] Streaming technology: Used to play audio descriptions provided by the server.
[0313] Explanation of program processing
[0314] Collecting user information
[0315] When a user first uses the system, the server displays an account registration page where the user enters information such as name, age, language preference, viewing history, and preferred genres, and also selects consent to emotion recognition. The server then creates a user profile based on this information and stores it in a database.
[0316] Examples:
[0317] When User A enters information into the registration form, the server receives it, generates a profile and stores it in the database.
[0318] Get movie information
[0319] When a user selects a movie they want to watch, the server retrieves basic information about the movie (title, genre, cast, director, and running time) from an online database, as well as public reviews, interviews with the cast and director, scripts, and subtitle files.
[0320] Examples:
[0321] When User A selects "Mystery Movie X," the server collects basic information about the movie and any associated scripts and subtitle files.
[0322] Scene Analysis
[0323] The server uses computer vision technology to analyze the film's video data to identify key scenes and character appearances, and NLP technology to analyze the script and subtitle files to extract key dialogue and emotional nuances.
[0324] Examples:
[0325] The server breaks down the data for Movie X frame by frame to identify key scenes and characters, while also analyzing the script to extract key lines and moving moments.
[0326] Recognizing user emotions with an emotion engine
[0327] The server uses an emotion engine to recognize emotions in real time from the user's facial expressions and voice. The device uses a camera and microphone to capture the user's facial expressions and voice while watching a movie, and sends the data to the server.
[0328] Examples:
[0329] When user A sheds tears during a touching scene, the device's camera captures this and sends the data to the server.
[0330] Generate and refine personalized guides
[0331] The server generates personalized audio guide text based on the user profile and the recognition results of the emotion engine, and dynamically adjusts the content of the audio guide based on the recognized emotions.
[0332] Examples:
[0333] The server recognizes that user A is moved and adds a specific example to the audio guide for that scene, such as "Character Y's deep emotions are expressed."
[0334] Speech synthesis
[0335] The server passes the generated audio guide text to a speech synthesis engine, which generates natural-sounding speech.
[0336] Examples:
[0337] The server passes the text "Character Y's deep emotions are expressed" to a speech synthesis engine, which generates realistic speech.
[0338] Audio guide distribution
[0339] The server provides the generated audio guide to the user's device in streaming or download format, and the device prepares to play the provided audio guide.
[0340] Examples:
[0341] The server sends the generated audio guide to the terminal, which prepares it for playback.
[0342] Gathering feedback
[0343] After watching a movie, users provide feedback on the quality and content of the audio guide, which the server collects and stores in a database.
[0344] Examples:
[0345] User A provides feedback saying, "The audio guide was easy to understand," and the server receives this and stores it in the database.
[0346] Profile Improvements
[0347] The server analyzes the feedback and updates the user profile, which is then reflected in the next guide generation.
[0348] Examples:
[0349] The server analyzes User A's feedback, adds "I prefer detailed explanations" to his profile, and reflects this in the next audio guide.
[0350] Prompt Sentence Examples
[0351] "Imagine a situation where a user is watching a mystery movie and a touching scene occurs. This makes the user's face begin to shed tears. We would like an explanation of how the emotion engine should recognize this situation and what kind of voice guidance should be generated."
[0352] Examples of prompts:
[0353] "If a user is crying during an emotional scene in a mystery movie, show how an emotion engine can recognize that emotion and add commentary that aligns with the user's emotions."
[0354] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0355] Step 1:
[0356] The server displays an account registration page when a user uses the system for the first time.
[0357] Input: The user enters basic information such as name, age, language preference, viewing history, and movie preferences, and selects consent for emotion recognition.
[0358] Data processing: The server receives this information and creates a user profile.
[0359] Output: The user profile is saved in the database.
[0360] Specific operation: When User A enters information into the registration form and clicks the submit button, the server receives it, generates a profile and saves it in the database.
[0361] Step 2:
[0362] The user selects the movie they want to watch.
[0363] Input: The user selects the movie they want to watch from the movie list.
[0364] Data processing: The server retrieves basic movie information (title, genre, cast, director, running time) from an online database, and collects public reviews, interviews with the cast and director, scripts, and subtitle files.
[0365] Output: The necessary movie information is stored on the server.
[0366] Specific operation: When user A selects the movie "Mystery Movie X," the server collects its basic information and related script and subtitle files.
[0367] Step 3:
[0368] The server analyzes the movie's video data using computer vision technology.
[0369] Input: Collected movie footage data
[0370] Data Processing: Video data is broken down frame by frame to identify key scenes and character appearances, and NLP techniques are used to analyze scripts and subtitle files to extract key dialogue and emotional nuances.
[0371] Output: A dataset containing identified scene and character information, as well as extracted dialogue and emotional nuances.
[0372] What it does: The server breaks down the data of Movie X frame by frame to identify important scenes and characters, while also analyzing the script to extract key lines and moving moments.
[0373] Step 4:
[0374] The device uses a camera and microphone to capture the user's facial expressions and voice while watching a movie, and sends the data to a server.
[0375] Input: facial expressions and voice data of the user while watching a movie
[0376] Data processing: Using an emotion engine, emotions are recognized in real time from the user's facial expressions and voice.
[0377] Output: Data about the user's emotional state.
[0378] Specific operation: When user A sheds tears during a touching scene, the device camera captures this and sends the data to the server.
[0379] Step 5:
[0380] The server generates personalized audio guide text based on the user profile and the recognition results of the emotion engine.
[0381] Input: User profile and real-time emotional data
[0382] Data processing: Generate text that is attuned to the user's preferences and emotions, and dynamically adjust content based on recognized emotions.
[0383] Output: Personalized audio guide text.
[0384] Specific operation: The server recognizes that user A is moved and adds a specific example to the audio description of that scene, such as "Character Y's deep emotions are expressed."
[0385] Step 6:
[0386] The server passes the data to a speech synthesis engine to generate speech from the text.
[0387] Input: The generated audio description text
[0388] Data processing: A speech synthesis engine converts text into speech and adds natural intonation and accent.
[0389] Output: Audio data that sounds similar to a human voice.
[0390] Specific operation: The server passes the text "Character Y's deep emotions are expressed" to the speech synthesis engine, which generates realistic speech.
[0391] Step 7:
[0392] The server provides the generated audio guide to the user's terminal in streaming or download format.
[0393] Input: Generated audio data
[0394] Data processing: The generated audio is converted into an appropriate format and sent to the user's device.
[0395] Output: Audio guide stored on the user's device.
[0396] Specific operation: The server sends the generated audio guide to the terminal, which prepares it for playback.
[0397] Step 8:
[0398] Users provide feedback after watching a movie.
[0399] Input: Feedback information about the quality and content of the audio description
[0400] Data processing: The server stores the collected feedback in a database.
[0401] Output: A database containing the feedback.
[0402] Specific operation: User A provides feedback such as "The audio guide was easy to understand," and the server receives this and stores it in the database.
[0403] Step 9:
[0404] The server analyzes the feedback and updates the user profile.
[0405] Input: Saved Feedback
[0406] Data processing: Analyze feedback and add user preferences and trends to your profile.
[0407] Output: The updated user profile.
[0408] Specific operation: The server analyzes user A's feedback, adds "I prefer detailed explanations" to his profile, and reflects this in the next audio guide.
[0409] (Application example 2)
[0410] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0411] Conventional movie viewing systems have limited means to accommodate visually impaired users and users who want to understand specific scenes or information in depth. In particular, the lack of a function to provide personalized audio guides that are sensitive to the user's emotions results in a poor movie-viewing experience. Furthermore, there is a lack of a mechanism to learn from user feedback and improve the guide content for future viewings. This makes it difficult to improve user satisfaction.
[0412] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0413] In this invention, the server includes means for collecting user preferences and viewing history and creating a profile, means for acquiring and analyzing basic movie information, ratings, related party information, dialogue, and commentary, means for generating a personalized audio guide based on the user profile, means for recognizing the user's emotions and dynamically adjusting the content of the audio guide, means for generating the generated audio guide using a speech synthesis engine, means for delivering the generated audio guide to the user's device, means for analyzing the collected feedback and updating the user profile, means for analyzing visual information using image analysis technology and identifying important scenes and scenes with characters, and means for generating prompt sentences using a generative AI model. This allows visually impaired people and emotionally sensitive users to better understand and enjoy movies.
[0414] "User" refers to a person who uses the system to watch a movie.
[0415] "Preferences" refer to personal preferences that a user has for particular movie genres or content.
[0416] "Viewing history" refers to historical information about movies and video content that a user has viewed in the past.
[0417] A "profile" refers to a data set that compiles information such as a user's preferences and viewing history.
[0418] "Basic movie information" refers to basic data about a movie, such as the title, genre, cast, director, and running time.
[0419] "Rating" refers to the ratings and reviews of movies by users and experts.
[0420] "Related Person Information" refers to detailed information about people such as the film's cast and director.
[0421] "Dialogue" refers to the words spoken by characters in a film.
[0422] "Explanatory text" refers to text that explains the content and background of the film.
[0423] "Emotion" refers to the psychological state a user experiences while watching a movie.
[0424] "Audio guide" refers to audio guidance or explanations provided to the user.
[0425] "Dynamic adjustment" refers to changing content in response to real-time conditions.
[0426] "Speech synthesis engine" refers to software or hardware for converting text data into speech data.
[0427] "Terminal" refers to the electronic device that a user uses to watch movies and use the system.
[0428] "Feedback" refers to opinions and evaluations from users.
[0429] "Image analysis technology" refers to technology for processing visual information and extracting specific patterns and features.
[0430] A "scene" refers to a specific situation or situation within a film.
[0431] A "character scene" refers to a specific moment in a film in which a character appears.
[0432] "Generative AI models" refer to artificial intelligence techniques for performing tasks such as text generation.
[0433] A "prompt sentence" refers to an input sentence that causes a generative AI model to generate text.
[0434] The present invention is a system for improving a user's viewing experience, which includes the following means: A server collects a user's preferences and viewing history and creates a profile, which includes information provided by the user such as name, age, language, viewing history, and movie preferences. The profile is stored in a database.
[0435] The server then retrieves and analyzes information such as the movie title, genre, cast, director, running time, ratings, people involved, dialogue, and description from online databases and APIs, and is then ready to provide detailed movie information to the user.
[0436] Based on the user profile, the server generates a personalized audio guide. To do so, it uses NLP technology to analyze the movie script and subtitles and extract important lines and scenes. It also recognizes the user's emotions in real time and dynamically adjusts the content of the audio guide based on that information. Emotion recognition is performed using, for example, Microsoft Azure Cognitive Services and Google Cloud's Natural Language API. This allows the server to provide commentary that matches the emotions the user feels while watching the movie.
[0437] The generated audio description is converted into speech with natural intonation and accent by a speech synthesis engine (e.g., Amazon Polly or Google Text-to-Speech), and is delivered to the user's device in streaming or download format.
[0438] Users provide feedback after watching a movie, which is collected and analyzed by the server, and their profile is updated to reflect this in the generation of personalized audio guides for future viewings.
[0439] Computer vision techniques (e.g., OpenCV) are used to analyze visual information and identify important scenes and character appearances in the film, which is also incorporated into the personalized audio description.
[0440] The generative AI model is used to generate prompts. An example of a prompt is, "Based on user profile: {'name': 'Mr. A', 'age': 30, 'language': 'jp', 'preferences': ['mystery']}, movie data: {detailed movie information}, emotion: {detailed emotion information}, please generate an in-depth explanation of the key scenes in this movie."
[0441] This allows visually impaired people and emotionally sensitive users to understand and enjoy movies more deeply. This system is ideal for content distribution services and is expected to significantly improve the movie-watching experience.
[0442] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0443] Step 1:
[0444] The server collects the user's preferences and viewing history and creates a profile. The input is information provided by the user about their name, age, language, viewing history, and movie preferences. The information is stored in a database and a user profile is generated. Specifically, the user enters information into an account creation form, which is received by the server.
[0445] Step 2:
[0446] The server retrieves basic movie information (title, genre, cast, director, running time), ratings, related party information, dialogue, and description from an online database or API. The input is the movie title and ID, and the output is the retrieved movie details. The server sends the API request and stores the retrieved data in a database for analysis.
[0447] Step 3:
[0448] The server uses computer vision technology to analyze the visual information of a movie and identify important scenes and scenes featuring characters. The input is the movie's video data, and the output is metadata for the identified scenes. Specifically, OpenCV is used to analyze the video data frame by frame and extract characteristic scenes.
[0449] Step 4:
[0450] The server uses NLP technology to analyze movie scripts and subtitles and extract important lines and scenes. The input is the text data of the script or subtitles, and the output is a list of important lines and scenes. Specifically, it uses NLP libraries such as spaCy to perform text analysis and extract important lines.
[0451] Step 5:
[0452] The device uses a camera and microphone to capture the user's facial expressions and voice and sends the data to a server. An emotion recognition engine analyzes this data. The input is the user's video and audio data, and the output is recognized emotional information. Specifically, the device captures the user's facial expressions and voice at regular intervals and sends them to the server in real time.
[0453] Step 6:
[0454] The server generates a personalized audio guide based on the user profile and real-time emotional data. A prompt is input into the generative AI model to generate the audio guide text. The input is profile information, emotional data, and the prompt, and the output is the audio guide text. An example of a specific prompt is, "Based on user profile: {'name': 'Mr. A', 'age': 30, 'language': 'jp', 'preferences': ['Mystery']}, movie data: {detailed movie information}, emotion: {detailed emotion information}, please generate an in-depth commentary on the important scenes in this movie."
[0455] Step 7:
[0456] The server passes the personalized audio guide text to a speech synthesis engine to generate a voice that sounds similar to a human voice. The input is the audio guide text data, and the output is the audio data. Specifically, the server converts the text to speech using a speech synthesis engine such as Google Text-to-Speech (gTTS).
[0457] Step 8:
[0458] The server distributes the generated audio guide to the user's device. The input is the generated audio data, and the output is the audio guide stored on the user's device. Specifically, the audio guide is provided in streaming format or file download format.
[0459] Step 9:
[0460] After watching a movie, users provide feedback on the quality and content of the audio description. The server collects this feedback and stores it in a database. The input is the user's feedback data, and the output is an updated user profile. For example, opinions can be collected through a web form or an in-app feedback feature.
[0461] Step 10:
[0462] The server improves the user profile based on the feedback and analysis data, and reflects this in future personalized audio guides. The input is the collected feedback and viewing data, and the output is an updated user profile. Specifically, the server analyzes the feedback data and adds new user preference information to the profile.
[0463] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0464] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0465] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0466] [Second embodiment]
[0467] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0468] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0469] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0470] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0471] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0472] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0473] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0474] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0475] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0476] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0477] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0478] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0479] This invention is a system for generating a personalized movie audio guide that is tailored to the user's situation and preferences. The program processing flow of this system will be explained below with specific examples.
[0480] Program processing overview
[0481] 1. Collecting User Information
[0482] When a user first uses the system, the server asks the user to register an account. The user enters basic information such as name, age, language preference, viewing history, movie preferences, etc. The server then creates a user profile and stores it in a database.
[0483] Examples:
[0484] User A creates an account and registers in his profile information that he is visually impaired, likes mystery movies, and seeks deep insights.
[0485] 2. Obtaining movie information
[0486] The server retrieves basic information about the movie selected by the user from online databases and APIs, as well as public reviews and interviews with the cast and director. It also retrieves the movie's script and subtitle files and prepares them as data for analysis.
[0487] 3. Audio guide generation
[0488] The server analyzes the collected movie data, recognizing important scenes and character appearances through scene analysis, and extracts important dialogue and emotional nuances using NLP technology.
[0489] It then generates personalized audio description text based on the user profile, providing detailed image descriptions for the visually impaired and in-depth insights and background information for movie buffs.
[0490] Examples:
[0491] For User A, a detailed audio guide is generated for each scene, and text is created that provides detailed explanations of the main cast's facial expressions and emotional changes, as well as the foreshadowing of the mystery.
[0492] The server then passes this text to a speech synthesis engine to generate a voice description that sounds more human, adding natural intonation and accents to make the description easier to understand.
[0493] 4. Audio guide distribution
[0494] The server provides the generated audio guide to the user's terminal in streaming or download format.
[0495] The terminal can play the provided audio description as secondary audio when the user watches the movie.
[0496] 5. Collecting User Feedback
[0497] The server collects feedback from users after they use the audio guide. Users can provide feedback on the quality and content of the guide, as well as areas for improvement. The server then adds this information to a profile and uses it to generate future guides.
[0498] Examples:
[0499] User A provides feedback stating that the emotional commentary section was very helpful and that he would like more specific scene descriptions. This feedback is reflected in the generation of guides for other movies he watches later.
[0500] Overall flow and summary
[0501] This system comprehensively handles everything from user registration to collecting and analyzing movie information, generating and distributing personalized audio guides, and collecting feedback. This allows for a movie-watching experience tailored to each user's needs, providing a high-value-added viewing experience not only for the visually impaired but also for a diverse audience. The introduction of this system is expected to open up new ways to enjoy movie-watching and provide a richer entertainment experience for many people.
[0502] The processing flow will be explained below.
[0503] Step 1: Registering a user
[0504] Users access the system and enter basic information such as name, age, language preference, viewing history, and movie preferences to create an account.
[0505] The server receives the entered information and creates and stores a user profile in a database.
[0506] Step 2: Update your profile
[0507] Users can update their profile information as their viewing history or preferences change.
[0508] The server stores the updated information in a database to keep the profile up to date.
[0509] Step 3: Get movie information
[0510] The user selects the movie they want to watch.
[0511] The server retrieves basic information such as the movie title, genre, cast, director, and running time from online databases and APIs.
[0512] The server also collects reviews, interviews, scripts, and subtitle files about the film.
[0513] Step 4: Analyze the scene
[0514] The server uses computer vision technology to analyze the movie's video data and identify key scenes and scenes in which characters appear.
[0515] The server uses NLP technology to analyze scripts and subtitle files to extract key dialogue and emotional nuances.
[0516] Step 5: Generate a personalized guide
[0517] The server decides how much detail to include in the audio description based on the user profile: detailed visual descriptions for the visually impaired, for example, or in-depth insights and background information for movie buffs.
[0518] The server creates the text for the audio guide based on the analysis results.
[0519] Step 6: Text-to-Speech
[0520] The server passes the generated audio guide text to a speech synthesis engine, which generates a voice that sounds like a human voice, with natural intonation and accent.
[0521] Step 7: Distributing the audio guide
[0522] The server provides the generated audio guide to the user's terminal in streaming or download format.
[0523] The terminal stores the provided audio guide and keeps it available for the user to use while watching the movie.
[0524] Step 8: Gather feedback
[0525] After watching a movie, users provide feedback on the quality and content of the audio description.
[0526] The server collects the provided feedback and stores it in a database, which is used to generate future guides.
[0527] Step 9: Improve your profile
[0528] The server analyzes the collected feedback and updates the user profile. This updated information is reflected in the next guide generation, providing an audio guide that is more tailored to the user's preferences.
[0529] Through the specific processing steps described above, the system provides a personalized movie-watching experience for each user, and meets the needs of visually impaired people and diverse audiences.
[0530] Example 1
[0531] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0532] Conventional movie audio guide systems have difficulty providing a fully personalized guide for the visually impaired or users who prefer specific movie genres. This limits the movie-watching experience, especially for users who cannot visually confirm the details of each movie scene or the emotional expressions of characters. Furthermore, it is difficult to effectively reflect user feedback and improve the guide, and further technical solutions are needed to increase user satisfaction.
[0533] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0534] In this invention, the server includes means for collecting user preferences and viewing history and creating a profile, means for acquiring and analyzing basic movie information, reviews, interviews, scripts, and subtitles, means for generating a personalized audio guide based on the user profile, means for analyzing the collected movie data and extracting important dialogue and emotional nuances using scene analysis and natural language processing techniques, means for generating the generated audio guide using a speech synthesis engine, means for delivering the generated audio guide to the user's terminal, and means for collecting user feedback and reflecting it in the profile. This allows for the provision of a detailed and personalized movie guide tailored to the user's needs, enabling a variety of users, including visually impaired people, to enjoy a richer movie-watching experience.
[0535] A "user profile" is a collection of data created based on a user's preferences, viewing history, basic information, and so on.
[0536] "Basic movie information" is basic information about a movie, such as the movie title, director, cast, and release date.
[0537] A "review" is text data that describes the ratings, impressions, and opinions posted by people who have watched a movie.
[0538] An "interview article" is text data that describes the contents of interviews with cast members, directors, and others involved in the film.
[0539] A "script" is text data that describes a movie script or screenplay.
[0540] "Subtitles" are text data that contains movie lines and explanations in written form.
[0541] "Scene analysis" is the process of analyzing each scene in a film to identify important scenes and characters.
[0542] "Natural language processing" is a technology that allows computers to understand, analyze, and generate human language.
[0543] An "audio description" is a guide that supplements visual information and provides audio information about the contents of a film.
[0544] A "speech synthesis engine" is a software system that generates natural-sounding speech based on text data.
[0545] "Feedback" refers to information such as evaluations, opinions, and impressions collected from users.
[0546] "Personalization" refers to a state that is customized based on the characteristics and preferences of an individual user.
[0547] The present invention is a system for generating a personalized movie audio guide that is adapted to the user's situation and preferences. Specific embodiments of this system will be described below.
[0548] System configuration
[0549] The system mainly consists of a server and a user device. The server creates user profiles, collects and analyzes movie data, generates audio guides, and collects feedback. The user device selects movies, plays audio guides, and provides feedback. Natural language processing technology and speech synthesis engines are used as software. Specific examples include AWS's Amazon Polly and Google Cloud Text-to-Speech.
[0550] Collecting user information
[0551] When a user first uses the system, the server asks the user to register an account. The user enters basic information such as name, age, language preference, viewing history, movie preferences, etc. The server then creates a user profile and stores it in a database.
[0552] Examples:
[0553] User A creates an account and registers in his profile information that he is visually impaired, likes mystery movies, and seeks deep insights.
[0554] Get movie information
[0555] The server retrieves basic information about the movie selected by the user from online databases and APIs, as well as public reviews and interviews with the cast and director. It also retrieves the movie's script and subtitle files and prepares them as data for analysis.
[0556] Examples:
[0557] When User A selects the movie "Sherlock Holmes," the server collects public information, reviews, scripts, etc. about the movie.
[0558] Audio guide generation
[0559] The server analyzes the collected movie data, recognizing important scenes and scenes in which characters appear through scene analysis, and extracting important lines and emotional nuances using natural language processing technology.
[0560] It then generates personalized audio description text based on the user profile, providing detailed visual descriptions for the visually impaired and in-depth insights and context for movie buffs. The generated text is then passed to a speech synthesis engine (e.g., AWS's Amazon Polly, Google Cloud Text-to-Speech) to generate a human-like audio description.
[0561] Examples:
[0562] A detailed audio guide for each scene is generated for User A, and text is created that provides detailed explanations of the facial expressions and emotional changes of the main cast, as well as hints at the mystery. An example of a prompt would be, "User A is visually impaired and loves mystery movies. Please generate a detailed audio guide for each scene of the movie 'Sherlock Holmes.' Please provide detailed explanations of the facial expressions and emotional changes of the main cast, as well as hints at the mystery."
[0563] Audio guide distribution
[0564] The server provides the generated audio guide to the user's terminal in a streaming or download format, and the terminal can play the provided audio guide as a secondary audio when the user watches the movie.
[0565] Collecting user feedback
[0566] The server collects feedback from users after they use the audio guide. Users can provide feedback on the quality and content of the guide, as well as areas for improvement. The server then adds this information to a profile and uses it to generate future guides.
[0567] Examples:
[0568] User A provides feedback stating that the emotional commentary section was very helpful and that he would like more specific scene descriptions. This feedback is reflected in the generation of guides for other movies he watches later.
[0569] In this way, the system can comprehensively handle everything from user registration to collecting and analyzing movie information, generating and distributing personalized audio guides, and collecting feedback, making it possible to provide a high-value movie-watching experience for a diverse audience, including the visually impaired.
[0570] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0571] Step 1:
[0572] Account Registration
[0573] When a user first accesses the system, the server displays an account registration form. The user enters basic information such as name, age, language preference, viewing history, and movie preferences, and then presses the submit button. The server receives the input data and creates a user profile. The created user profile is stored in a database.
[0574] Input: Basic information entered by the user (name, age, language preference, viewing history, movie preferences)
[0575] Output: User profile stored in the database
[0576] Specific behavior:
[0577] The server displays a registration form through a user interface. The user fills in the form and presses the submit button. The server takes the input data, generates a user profile, and stores it in a database.
[0578] Step 2:
[0579] Movie Selection
[0580] The user selects a movie, enters the movie title in the search box on the user terminal, and presses the search button, and the selected movie title is sent to the server.
[0581] Input: The movie title selected by the user
[0582] Output: Movie title sent to the server
[0583] Specific behavior:
[0584] The user terminal displays a movie title entry form, and when the user enters a title and presses the search button, the selection is sent to the server.
[0585] Step 3:
[0586] Movie data collection
[0587] Based on the received movie title, the server uses online databases and APIs to collect related data such as basic movie information, reviews, interviews, scripts, subtitles, etc. The collected data is then stored for analysis.
[0588] Input: Movie title
[0589] Output: Data stored on the server, including basic movie information, reviews, interviews, scripts, subtitles, etc.
[0590] Specific behavior:
[0591] The server calls a specific API, sends the movie title, and retrieves the relevant data, which is then stored in a database for analysis.
[0592] Step 4:
[0593] Data analysis and audio guide generation
[0594] The server analyzes the stored movie data, performs scene analysis to identify key scenes and character appearances, and uses natural language processing techniques to extract key lines and emotional nuances. Based on the extracted data, it generates personalized audio description text that matches the user profile. The generated text is then passed to a speech synthesis engine.
[0595] Input: saved movie data, user profile
[0596] Output: Personalized audio guide text and audio files
[0597] Specific behavior:
[0598] The server runs an algorithm that analyzes the movie scenes to determine their importance. It then uses natural language processing techniques to analyze the dialogue and extract key elements. It then inputs the prompt into a generative AI model to generate personalized text in a specific format. Finally, it passes the text to a speech synthesis engine to generate an audio file.
[0599] Step 5:
[0600] Audio guide distribution
[0601] The server provides the generated audio guide to the user's device in a streaming or download format, and the device can play the provided audio guide as a secondary audio when the user watches the movie.
[0602] Input: Personalized audio guide audio file
[0603] Output: Audio guide delivered to the user's device
[0604] Specific behavior:
[0605] The server uploads the generated audio guide to cloud storage and generates an access URL, which the user's device uses to play or download the audio guide.
[0606] Step 6:
[0607] Gathering feedback
[0608] After the user has finished watching, the server provides a feedback form. The user can enter their evaluation of the quality and content of the guide and suggestions for improvement in the feedback form and submit it. The server then reflects the collected feedback in the profile and uses it for generating audio guides from the next time onwards.
[0609] Input: User feedback
[0610] Output: Updated user profile
[0611] Specific behavior:
[0612] After the user has finished watching, the server displays a feedback form. The user enters their feedback in the form and presses the submit button. The server receives the feedback data and updates the user profile.
[0613] (Application example 1)
[0614] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0615] Conventional movie audio guide systems have not adequately provided personalized guides based on the viewer's individual preferences and viewing history. As a result, it has been difficult to provide an optimal viewing experience for visually impaired people or users who prefer specific movie genres. Furthermore, continuous improvement is required due to insufficient improvement in the quality of generated audio guides and profile updates based on user feedback.
[0616] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0617] In this invention, the server includes means for collecting user preferences and viewing history and creating a profile, means for acquiring and analyzing basic movie information, reviews, interviews, scripts, and subtitles, means for generating a personalized audio guide based on the user profile, means for generating the generated audio guide using a speech synthesis engine, means for delivering the generated audio guide to the user's terminal, means for generating text for the personalized audio guide based on the user profile and movie data using a generative AI model, means for calling the generative AI model using a prompt sentence in generating the audio guide, and means for converting the generated audio guide into a format suitable for smart devices and delivering it, thereby enabling the provision of an audio guide optimized for each user.
[0618] "User preferences" refers to information that indicates the user's tastes in movie genres and specific movies.
[0619] "Viewing history" refers to data that records the history of movies and dramas that a user has watched.
[0620] A "profile" refers to a collection of information that compiles a user's personal information, viewing history, preferences, etc.
[0621] "Basic movie information" refers to basic data such as the movie title, director, cast, release year, and genre.
[0622] A "review" is a text that shows criticism or evaluation of a movie, and is a source of information that allows users to learn the opinions of other viewers and experts.
[0623] An "interview article" refers to a written record of an interview conducted with a film's cast, director, or production staff.
[0624] A "script" refers to a film script or screenplay, a document that contains the actual dialogue and scene details of a film.
[0625] "Subtitles" refers to text that visually displays dialogue or narration in a film.
[0626] "Personalized audio description" refers to audio description that is customized based on a user's individual profile.
[0627] A "speech synthesis engine" is a program for converting text into voice data.
[0628] "Generative AI Model" refers to a model that uses artificial intelligence technology to generate personalized guides based on a user's profile and movie data.
[0629] A "prompt" is an instruction or question input to a generative AI model, and is text used to control the model's output.
[0630] "Smart devices" refers to electronic devices with advanced functions such as smartphones and smart glasses.
[0631] This invention is a system that generates and distributes personalized audio guides that are tailored to the user's circumstances and preferences. The program processing flow of this system is explained below with specific examples.
[0632] 1. Collection of User Information
[0633] When a user first uses the system, the server asks the user to register an account. The user enters basic information such as name, age, language preference, viewing history, movie preferences, etc. The server then creates a user profile and stores it in a database.
[0634] 2. Obtaining movie information
[0635] The server retrieves basic information about the movie selected by the user from online databases and APIs, including the movie title, director, cast, reviews, interviews, script, and subtitle files, and prepares this data for analysis.
[0636] 3. Audio guide generation
[0637] The server analyzes the collected movie data and uses NLP techniques to extract key scenes and emotional nuances. It then generates personalized audio description text based on the user profile. A generative AI model is used to generate the text, and a prompt is input to the model. For example, a prompt might be in the form: "User profile: {Name: User A, Age: 30, Language: Japanese, Favorite movie genre: Mystery} Movie data: {Title: Inception, Director: Christopher Nolan, Cast: {Leonardo DiCaprio: Dom Cobb}, Review: An extremely complex and intricately structured story...} Please generate a personalized movie description."
[0638] After the audio description text is generated, the server passes it to a speech synthesis engine to generate a human-like audio description, adding natural intonation and accents to make it easier to listen to.
[0639] 4. Audio guide distribution
[0640] The generated audio guide is provided to the user's smart device, such as a smartphone or smart glasses, via streaming or download. The server converts the audio guide into a format optimized for the user's device and delivers it appropriately.
[0641] 5. Collecting User Feedback
[0642] Users can provide feedback after using the audio guide. The server collects and analyzes this feedback. The collected feedback is reflected in the user profile and is used to generate audio guides in the future.
[0643] This system makes it possible to provide a personalized movie audio guide tailored to the user's situation and preferences, providing a richer movie-watching experience.
[0644] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0645] Step 1:
[0646] The server requires a user to register an account when using the system for the first time. The user enters basic information such as name, age, language setting, viewing history, and movie preferences. Based on this information, the server creates a user profile and stores it in a database. Input data: name, age, language setting, viewing history, and movie preferences. Output data: user profile.
[0647] Step 2:
[0648] When a user selects a movie to watch, the server retrieves basic information about the movie, including the movie title, director, cast, reviews, interviews, script, and subtitle files. Data is collected from online databases and APIs. Input data: movie title. Output data: movie basic information, reviews, interviews, script, and subtitle files.
[0649] Step 3:
[0650] The server analyzes the collected movie data using NLP technology, extracting key scenes and emotional nuances. Input data: Basic movie information, reviews, interviews, scripts, and subtitle files. Output data: Analyzed movie data.
[0651] Step 4:
[0652] Using the generative AI model, the server generates personalized audio guide text based on the user profile and analyzed movie data. A prompt is input to the generative AI model. For example, a prompt in the following format might be used: "User profile: {Name: User A, Age: 30, Language: Japanese, Favorite movie genre: Mystery} Movie data: {Title: Inception, Director: Generic name, Cast: {Starring actor: Lead role}, Review: Extremely complex and intricately structured story} Please generate a personalized movie audio guide." Input data: User profile, analyzed movie data. Output data: Audio guide text.
[0653] Step 5:
[0654] The server passes the generated audio description text to a speech synthesis engine to generate an audio description that sounds like a human voice, with natural intonation and accent. Input data: Audio description text. Output data: Audio data.
[0655] Step 6:
[0656] The server provides the generated audio guide to the user's smart device, such as a smartphone or smart glasses, in streaming or download format. The audio guide is converted into a format optimized for the user's device. Input data: Audio data. Output data: Audio data in a format suitable for the smart device.
[0657] Step 7:
[0658] Users can provide feedback after using the audio guide. The server collects and analyzes this feedback. The collected feedback is reflected in the user profile and is used to generate audio guides from next time onwards. Input data: User feedback. Output data: Updated user profile.
[0659] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0660] This invention is a system that generates a personalized movie audio guide tailored to the user's situation and preferences, and further combines an emotion engine to recognize the user's emotions and dynamically adjust the content of the audio guide. The program processing flow of this system is explained below with concrete examples.
[0661] Program processing overview
[0662] 1. Collecting User Information
[0663] When a user first uses the system, the server asks the user to register an account. The user enters basic information such as name, age, language preference, viewing history, and movie preferences, and also asks for consent to emotion recognition. Based on this information, the server creates a user profile and stores it in a database.
[0664] Examples:
[0665] User A creates an account and registers in his profile information that he is visually impaired, likes mystery movies, and seeks deep insights. He also agrees to emotion recognition.
[0666] 2. Obtaining movie information
[0667] The user selects a movie they want to watch. The server retrieves basic information about the movie, such as the title, genre, cast, director, and running time, from online databases and APIs. It also collects public reviews, interviews with the cast and director, scripts, and subtitle files.
[0668] 3. Scene Analysis
[0669] The server uses computer vision technology to analyze the film's video data to identify key scenes and character appearances, and NLP technology to analyze the script and subtitle files to extract key dialogue and emotional nuances.
[0670] 4. User Emotion Recognition by Emotion Engine
[0671] The server uses an emotion engine to recognize emotions from the user's facial expressions and voice in real time. While the user is watching a movie, the device uses a camera and microphone to capture the user's facial expressions and voice and transmits the data to the server.
[0672] Examples:
[0673] While User A is watching a movie, the emotion engine recognizes that he is crying at a touching scene. This information is sent to the server.
[0674] 5. Generate and refine personalized guides
[0675] The server generates personalized audio guide text based on the user profile and the recognition results of the emotion engine. Furthermore, it dynamically adjusts the content of the audio guide according to the recognized emotion. For example, if the user is emotional, it adds commentary that reflects that emotion.
[0676] Examples:
[0677] If user A is in tears during a touching scene, the audio guide for that scene will add commentary that delves deeper into the character's emotions.
[0678] 6. Speech Synthesis
[0679] The server then passes the generated audio guide text to a speech synthesis engine, which generates a voice that sounds like a human voice, with natural intonation and accent.
[0680] 7. Audio guide distribution
[0681] The server provides the generated audio guide to the user's terminal in streaming or download format, and the terminal stores the provided audio guide so that the user can use it while watching the movie.
[0682] 8. Gathering Feedback
[0683] After watching a movie, users provide feedback on the quality and content of the audio guide. The server collects the feedback and stores it in a database. The feedback is used to generate future guides.
[0684] 9. Profile Improvements
[0685] The server analyzes the collected feedback and updates the user profile. This updated information is reflected in the next guide generation, providing an audio guide that is more tailored to the user's preferences.
[0686] Overall flow and summary
[0687] This system comprehensively handles everything from user registration to collecting and analyzing movie information, real-time emotion recognition using an emotion engine, generating and adjusting personalized audio guides, delivering them, and collecting feedback. This provides a movie-watching experience tailored to each user's needs and emotions, and addresses the needs of visually impaired people and a diverse audience. The introduction of this system is expected to open up new ways to enjoy movie-going and provide a richer entertainment experience for many people.
[0688] The processing flow will be explained below.
[0689] Step 1: Registering a user
[0690] Users access the system and enter their name, age, language preference, viewing history, favorite movie genre, preferred level of visual and auditory assistance, and consent to emotion recognition on the account creation screen.
[0691] The server receives the information entered by the user and stores it in a database as a new user profile.
[0692] Step 2: Update your profile
[0693] Users can update their profile to reflect changes to their viewing history or new movie preferences by entering and saving changes on the profile page.
[0694] The server updates and saves the changes made by the user to the profile information in the database.
[0695] Step 3: Get movie information
[0696] The user selects the movie they want to watch from within the system.
[0697] The server retrieves basic information about the selected movie, such as title, genre, cast, director, and running time, from online databases or APIs.
[0698] The server also collects reviews of the film, interviews with the cast and director, scripts, and subtitle files.
[0699] Step 4: Analyze the scene
[0700] The server uses computer vision technology to analyze the movie's video data and identify key scenes and scenes in which characters appear.
[0701] The server uses NLP technology to analyze scripts and subtitle files to extract key dialogue and emotional nuances.
[0702] Step 5: Recognizing user emotions with the emotion engine
[0703] The device uses a built-in camera and microphone to capture the user's facial expressions and voice in real time while the movie is playing.
[0704] The terminal sends the captured data to an emotion engine to recognize the user's emotional state.
[0705] The emotion engine (server) analyzes the received data, identifies the user's emotion (such as joy, sadness, surprise, etc.), and generates an emotion status.
[0706] Step 6: Generate and refine your personalized guide
[0707] The server generates personalized audio guide text based on the user profile and the recognition results of the emotion engine.
[0708] The server dynamically adjusts the content of the audio guide depending on the recognized emotion, for example, by adding explanations to calm the user or by emphasizing emotional background information for moving scenes.
[0709] Examples:
[0710] If user A is in tears during a moving scene in a movie, the server can incorporate additional emotional context into the audio description to enhance the viewing experience of that scene.
[0711] Step 7: Text-to-Speech
[0712] The server then passes the generated audio guide text to a speech synthesis engine, which generates a voice that sounds similar to a natural human voice, with natural intonation and accent.
[0713] Step 8: Distributing the audio guide
[0714] The server provides the generated audio guide to the user's terminal in streaming or download format.
[0715] The terminal stores the provided audio description and prepares it for playback as secondary audio during movie playback.
[0716] Step 9: Gather feedback
[0717] After watching a movie, users can enter their feedback on the quality and content of the audio guide on a feedback page within the system.
[0718] The server receives the provided feedback and stores it in a database, which is used to generate subsequent guides.
[0719] Step 10: Refine your profile
[0720] The server analyzes the collected feedback and updates the user profile. This updated information is reflected in the next guide generation, providing an audio guide that is more tailored to the user's preferences.
[0721] Overall flow and summary
[0722] This process comprehensively covers user registration, movie information collection, real-time emotion recognition using an emotion engine, personalized and tailored audio guide generation and delivery, and feedback collection. By dynamically reflecting user emotions, the system can promote a more personalized movie-watching experience and cater to a diverse audience.
[0723] Example 2
[0724] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0725] There is a demand for a variety of viewers, including the visually impaired, to enrich their movie experience and provide personalized guidance that reflects the user's individual emotions and preferences. Conventional systems have difficulty providing guidance that reflects the user's emotions in real time, resulting in only a uniform guidance that makes it difficult to increase user satisfaction.
[0726] The specification process by the specification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for collecting a user's preferences and viewing history and creating a profile; means for acquiring and analyzing basic movie information, reviews, interviews, scripts, and subtitles; means for analyzing visual information using computer vision technology and identifying important scenes and scenes in which characters appear; means for capturing the user's facial expressions and voice through a terminal while watching a movie and recognizing the user's emotions using an emotion engine; means for generating and dynamically adjusting a personalized audio guide based on the user profile and the recognized emotions; means for generating the generated audio guide using a speech synthesis engine; and means for delivering the generated audio guide to the user's terminal. This makes it possible to provide a personalized movie experience tailored to the user's emotions and preferences.
[0727] "User preferences" refer to the tastes and preferences that a user has for a particular genre or content.
[0728] A "viewing history" is a record of movies and programs a user has watched.
[0729] A "profile" is a collection of personal information such as a user's name, age, language preference, viewing history, and movie preferences.
[0730] "Basic movie information" refers to information such as the movie title, genre, cast, director, and screening time.
[0731] A "review" is a piece of writing that expresses an opinion or rating about a movie.
[0732] An "interview article" is a written version of interviews with cast members, directors, and others involved in the film.
[0733] A "script" is a screenplay that contains the dialogue and direction of a movie.
[0734] "Subtitles" are the written versions of the audio in a movie that are displayed on the screen.
[0735] "Computer vision technology" is a technology that allows computers to analyze the content of images and videos.
[0736] An "emotion engine" is a software technology that recognizes emotions by analyzing a user's facial expressions and voice.
[0737] A "speech synthesis engine" is a technology that generates speech by adding natural intonation and accent based on text information.
[0738] A "terminal" is a device that a user uses to watch a movie, including a smartphone, tablet, or PC.
[0739] The present invention provides a system for generating a personalized movie audio guide based on a user's preferences, viewing history, and emotions, thereby improving the user's movie-watching experience. Specific embodiments of the present invention will be described below.
[0740] Hardware and software used
[0741] server:
[0742] Database: Used to store user profiles, movie information, feedback, etc.
[0743] Computer vision technology: Used to analyze film footage and identify key scenes and character appearances.
[0744] Emotion engine: Used to analyze the user's facial expressions and voice to recognize emotions in real time.
[0745] Speech synthesis engine: Used to convert text to speech for personalized audio descriptions.
[0746] Device:
[0747] Camera: Used to capture the user's facial expressions.
[0748] Microphone: Used to capture the user's voice.
[0749] Streaming technology: Used to play audio descriptions provided by the server.
[0750] Explanation of program processing
[0751] Collecting user information
[0752] When a user first uses the system, the server displays an account registration page where the user enters information such as name, age, language preference, viewing history, and preferred genres, and also selects consent to emotion recognition. The server then creates a user profile based on this information and stores it in a database.
[0753] Examples:
[0754] When User A enters information into the registration form, the server receives it, generates a profile and stores it in the database.
[0755] Get movie information
[0756] When a user selects a movie they want to watch, the server retrieves basic information about the movie (title, genre, cast, director, and running time) from an online database, as well as public reviews, interviews with the cast and director, scripts, and subtitle files.
[0757] Examples:
[0758] When User A selects "Mystery Movie X," the server collects basic information about the movie and any associated scripts and subtitle files.
[0759] Scene Analysis
[0760] The server uses computer vision technology to analyze the film's video data to identify key scenes and character appearances, and NLP technology to analyze the script and subtitle files to extract key dialogue and emotional nuances.
[0761] Examples:
[0762] The server breaks down the data for Movie X frame by frame to identify key scenes and characters, while also analyzing the script to extract key lines and moving moments.
[0763] Recognizing user emotions with an emotion engine
[0764] The server uses an emotion engine to recognize emotions in real time from the user's facial expressions and voice. The device uses a camera and microphone to capture the user's facial expressions and voice while watching a movie, and sends the data to the server.
[0765] Examples:
[0766] When user A sheds tears during a touching scene, the device's camera captures this and sends the data to the server.
[0767] Generate and refine personalized guides
[0768] The server generates personalized audio guide text based on the user profile and the recognition results of the emotion engine, and dynamically adjusts the content of the audio guide based on the recognized emotions.
[0769] Examples:
[0770] The server recognizes that user A is moved and adds a specific example to the audio guide for that scene, such as "Character Y's deep emotions are expressed."
[0771] Speech synthesis
[0772] The server passes the generated audio guide text to a speech synthesis engine, which generates natural-sounding speech.
[0773] Examples:
[0774] The server passes the text "Character Y's deep emotions are expressed" to a speech synthesis engine, which generates realistic speech.
[0775] Audio guide distribution
[0776] The server provides the generated audio guide to the user's device in streaming or download format, and the device prepares to play the provided audio guide.
[0777] Examples:
[0778] The server sends the generated audio guide to the terminal, which prepares it for playback.
[0779] Gathering feedback
[0780] After watching a movie, users provide feedback on the quality and content of the audio guide, which the server collects and stores in a database.
[0781] Examples:
[0782] User A provides feedback saying, "The audio guide was easy to understand," and the server receives this and stores it in the database.
[0783] Profile Improvements
[0784] The server analyzes the feedback and updates the user profile, which is then reflected in the next guide generation.
[0785] Examples:
[0786] The server analyzes User A's feedback, adds "I prefer detailed explanations" to his profile, and reflects this in the next audio guide.
[0787] Prompt Sentence Examples
[0788] "Imagine a situation where a user is watching a mystery movie and a touching scene occurs. This makes the user's face begin to shed tears. We would like an explanation of how the emotion engine should recognize this situation and what kind of voice guidance should be generated."
[0789] Examples of prompts:
[0790] "If a user is crying during an emotional scene in a mystery movie, show how an emotion engine can recognize that emotion and add commentary that aligns with the user's emotions."
[0791] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0792] Step 1:
[0793] The server displays an account registration page when a user uses the system for the first time.
[0794] Input: The user enters basic information such as name, age, language preference, viewing history, and movie preferences, and selects consent for emotion recognition.
[0795] Data processing: The server receives this information and creates a user profile.
[0796] Output: The user profile is saved in the database.
[0797] Specific operation: When User A enters information into the registration form and clicks the submit button, the server receives it, generates a profile and saves it in the database.
[0798] Step 2:
[0799] The user selects the movie they want to watch.
[0800] Input: The user selects the movie they want to watch from the movie list.
[0801] Data processing: The server retrieves basic movie information (title, genre, cast, director, running time) from an online database, and collects public reviews, interviews with the cast and director, scripts, and subtitle files.
[0802] Output: The necessary movie information is stored on the server.
[0803] Specific operation: When user A selects the movie "Mystery Movie X," the server collects its basic information and related script and subtitle files.
[0804] Step 3:
[0805] The server analyzes the movie's video data using computer vision technology.
[0806] Input: Collected movie footage data
[0807] Data Processing: Video data is broken down frame by frame to identify key scenes and character appearances, and NLP techniques are used to analyze scripts and subtitle files to extract key dialogue and emotional nuances.
[0808] Output: A dataset containing identified scene and character information, as well as extracted dialogue and emotional nuances.
[0809] What it does: The server breaks down the data of Movie X frame by frame to identify important scenes and characters, while also analyzing the script to extract key lines and moving moments.
[0810] Step 4:
[0811] The device uses a camera and microphone to capture the user's facial expressions and voice while watching a movie, and sends the data to a server.
[0812] Input: facial expressions and voice data of the user while watching a movie
[0813] Data processing: Using an emotion engine, emotions are recognized in real time from the user's facial expressions and voice.
[0814] Output: Data about the user's emotional state.
[0815] Specific operation: When user A sheds tears during a touching scene, the device camera captures this and sends the data to the server.
[0816] Step 5:
[0817] The server generates personalized audio guide text based on the user profile and the recognition results of the emotion engine.
[0818] Input: User profile and real-time emotional data
[0819] Data processing: Generate text that is attuned to the user's preferences and emotions, and dynamically adjust content based on recognized emotions.
[0820] Output: Personalized audio guide text.
[0821] Specific operation: The server recognizes that user A is moved and adds a specific example to the audio description of that scene, such as "Character Y's deep emotions are expressed."
[0822] Step 6:
[0823] The server passes the data to a speech synthesis engine to generate speech from the text.
[0824] Input: The generated audio description text
[0825] Data processing: A speech synthesis engine converts text into speech and adds natural intonation and accent.
[0826] Output: Audio data that sounds similar to a human voice.
[0827] Specific operation: The server passes the text "Character Y's deep emotions are expressed" to the speech synthesis engine, which generates realistic speech.
[0828] Step 7:
[0829] The server provides the generated audio guide to the user's terminal in streaming or download format.
[0830] Input: Generated audio data
[0831] Data processing: The generated audio is converted into an appropriate format and sent to the user's device.
[0832] Output: Audio guide stored on the user's device.
[0833] Specific operation: The server sends the generated audio guide to the terminal, which prepares it for playback.
[0834] Step 8:
[0835] Users provide feedback after watching a movie.
[0836] Input: Feedback information about the quality and content of the audio description
[0837] Data processing: The server stores the collected feedback in a database.
[0838] Output: A database containing the feedback.
[0839] Specific operation: User A provides feedback such as "The audio guide was easy to understand," and the server receives this and stores it in the database.
[0840] Step 9:
[0841] The server analyzes the feedback and updates the user profile.
[0842] Input: Saved Feedback
[0843] Data processing: Analyze feedback and add user preferences and trends to your profile.
[0844] Output: The updated user profile.
[0845] Specific operation: The server analyzes user A's feedback, adds "I prefer detailed explanations" to his profile, and reflects this in the next audio guide.
[0846] (Application example 2)
[0847] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0848] Conventional movie viewing systems have limited means to accommodate visually impaired users and users who want to understand specific scenes or information in depth. In particular, the lack of a function to provide personalized audio guides that are sensitive to the user's emotions results in a poor movie-viewing experience. Furthermore, there is a lack of a mechanism to learn from user feedback and improve the guide content for future viewings. This makes it difficult to improve user satisfaction.
[0849] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0850] In this invention, the server includes means for collecting user preferences and viewing history and creating a profile, means for acquiring and analyzing basic movie information, ratings, related party information, dialogue, and commentary, means for generating a personalized audio guide based on the user profile, means for recognizing the user's emotions and dynamically adjusting the content of the audio guide, means for generating the generated audio guide using a speech synthesis engine, means for delivering the generated audio guide to the user's device, means for analyzing the collected feedback and updating the user profile, means for analyzing visual information using image analysis technology and identifying important scenes and scenes with characters, and means for generating prompt sentences using a generative AI model. This allows visually impaired people and emotionally sensitive users to better understand and enjoy movies.
[0851] "User" refers to a person who uses the system to watch a movie.
[0852] "Preferences" refer to personal preferences that a user has for particular movie genres or content.
[0853] "Viewing history" refers to historical information about movies and video content that a user has viewed in the past.
[0854] A "profile" refers to a data set that compiles information such as a user's preferences and viewing history.
[0855] "Basic movie information" refers to basic data about a movie, such as the title, genre, cast, director, and running time.
[0856] "Rating" refers to the ratings and reviews of movies by users and experts.
[0857] "Related Person Information" refers to detailed information about people such as the film's cast and director.
[0858] "Dialogue" refers to the words spoken by characters in a film.
[0859] "Explanatory text" refers to text that explains the content and background of the film.
[0860] "Emotion" refers to the psychological state a user experiences while watching a movie.
[0861] "Audio guide" refers to audio guidance or explanations provided to the user.
[0862] "Dynamic adjustment" refers to changing content in response to real-time conditions.
[0863] "Speech synthesis engine" refers to software or hardware for converting text data into speech data.
[0864] "Terminal" refers to the electronic device that a user uses to watch movies and use the system.
[0865] "Feedback" refers to opinions and evaluations from users.
[0866] "Image analysis technology" refers to technology for processing visual information and extracting specific patterns and features.
[0867] A "scene" refers to a specific situation or situation within a film.
[0868] A "character scene" refers to a specific moment in a film in which a character appears.
[0869] "Generative AI models" refer to artificial intelligence techniques for performing tasks such as text generation.
[0870] A "prompt sentence" refers to an input sentence that causes a generative AI model to generate text.
[0871] The present invention is a system for improving a user's viewing experience, which includes the following means: A server collects a user's preferences and viewing history and creates a profile, which includes information provided by the user such as name, age, language, viewing history, and movie preferences. The profile is stored in a database.
[0872] The server then retrieves and analyzes information such as the movie title, genre, cast, director, running time, ratings, people involved, dialogue, and description from online databases and APIs, and is then ready to provide detailed movie information to the user.
[0873] Based on the user profile, the server generates a personalized audio guide. To do so, it uses NLP technology to analyze the movie script and subtitles and extract important lines and scenes. It also recognizes the user's emotions in real time and dynamically adjusts the content of the audio guide based on that information. Emotion recognition is performed using, for example, Microsoft Azure Cognitive Services and Google Cloud's Natural Language API. This allows the server to provide commentary that matches the emotions the user feels while watching the movie.
[0874] The generated audio description is converted into speech with natural intonation and accent by a speech synthesis engine (e.g., Amazon Polly or Google Text-to-Speech), and is delivered to the user's device in streaming or download format.
[0875] Users provide feedback after watching a movie, which is collected and analyzed by the server, and their profile is updated to reflect this in the generation of personalized audio guides for future viewings.
[0876] Computer vision techniques (e.g., OpenCV) are used to analyze visual information and identify important scenes and character appearances in the film, which is also incorporated into the personalized audio description.
[0877] The generative AI model is used to generate prompts. An example of a prompt is, "Based on user profile: {'name': 'Mr. A', 'age': 30, 'language': 'jp', 'preferences': ['mystery']}, movie data: {detailed movie information}, emotion: {detailed emotion information}, please generate an in-depth explanation of the key scenes in this movie."
[0878] This allows visually impaired people and emotionally sensitive users to understand and enjoy movies more deeply. This system is ideal for content distribution services and is expected to significantly improve the movie-watching experience.
[0879] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0880] Step 1:
[0881] The server collects the user's preferences and viewing history and creates a profile. The input is information provided by the user about their name, age, language, viewing history, and movie preferences. The information is stored in a database and a user profile is generated. Specifically, the user enters information into an account creation form, which is received by the server.
[0882] Step 2:
[0883] The server retrieves basic movie information (title, genre, cast, director, running time), ratings, related party information, dialogue, and description from an online database or API. The input is the movie title and ID, and the output is the retrieved movie details. The server sends the API request and stores the retrieved data in a database for analysis.
[0884] Step 3:
[0885] The server uses computer vision technology to analyze the visual information of a movie and identify important scenes and scenes featuring characters. The input is the movie's video data, and the output is metadata for the identified scenes. Specifically, OpenCV is used to analyze the video data frame by frame and extract characteristic scenes.
[0886] Step 4:
[0887] The server uses NLP technology to analyze movie scripts and subtitles and extract important lines and scenes. The input is the text data of the script or subtitles, and the output is a list of important lines and scenes. Specifically, it uses NLP libraries such as spaCy to perform text analysis and extract important lines.
[0888] Step 5:
[0889] The device uses a camera and microphone to capture the user's facial expressions and voice and sends the data to a server. An emotion recognition engine analyzes this data. The input is the user's video and audio data, and the output is recognized emotional information. Specifically, the device captures the user's facial expressions and voice at regular intervals and sends them to the server in real time.
[0890] Step 6:
[0891] The server generates a personalized audio guide based on the user profile and real-time emotional data. A prompt is input into the generative AI model to generate the audio guide text. The input is profile information, emotional data, and the prompt, and the output is the audio guide text. An example of a specific prompt is, "Based on user profile: {'name': 'Mr. A', 'age': 30, 'language': 'jp', 'preferences': ['Mystery']}, movie data: {detailed movie information}, emotion: {detailed emotion information}, please generate an in-depth commentary on the important scenes in this movie."
[0892] Step 7:
[0893] The server passes the personalized audio guide text to a speech synthesis engine to generate a voice that sounds similar to a human voice. The input is the audio guide text data, and the output is the audio data. Specifically, the server converts the text to speech using a speech synthesis engine such as Google Text-to-Speech (gTTS).
[0894] Step 8:
[0895] The server distributes the generated audio guide to the user's device. The input is the generated audio data, and the output is the audio guide stored on the user's device. Specifically, the audio guide is provided in streaming format or file download format.
[0896] Step 9:
[0897] After watching a movie, users provide feedback on the quality and content of the audio description. The server collects this feedback and stores it in a database. The input is the user's feedback data, and the output is an updated user profile. For example, opinions can be collected through a web form or an in-app feedback feature.
[0898] Step 10:
[0899] The server improves the user profile based on the feedback and analysis data, and reflects this in future personalized audio guides. The input is the collected feedback and viewing data, and the output is an updated user profile. Specifically, the server analyzes the feedback data and adds new user preference information to the profile.
[0900] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0901] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0902] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0903] [Third embodiment]
[0904] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0905] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0906] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0907] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0908] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0909] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0910] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0911] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0912] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0913] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0914] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0915] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0916] This invention is a system for generating a personalized movie audio guide that is tailored to the user's situation and preferences. The program processing flow of this system will be explained below with specific examples.
[0917] Program processing overview
[0918] 1. Collecting User Information
[0919] When a user first uses the system, the server asks the user to register an account. The user enters basic information such as name, age, language preference, viewing history, movie preferences, etc. The server then creates a user profile and stores it in a database.
[0920] Examples:
[0921] User A creates an account and registers in his profile information that he is visually impaired, likes mystery movies, and seeks deep insights.
[0922] 2. Obtaining movie information
[0923] The server retrieves basic information about the movie selected by the user from online databases and APIs, as well as public reviews and interviews with the cast and director. It also retrieves the movie's script and subtitle files and prepares them as data for analysis.
[0924] 3. Audio guide generation
[0925] The server analyzes the collected movie data, recognizing important scenes and character appearances through scene analysis, and extracts important dialogue and emotional nuances using NLP technology.
[0926] It then generates personalized audio description text based on the user profile, providing detailed image descriptions for the visually impaired and in-depth insights and background information for movie buffs.
[0927] Examples:
[0928] For User A, a detailed audio guide is generated for each scene, and text is created that provides detailed explanations of the main cast's facial expressions and emotional changes, as well as the foreshadowing of the mystery.
[0929] The server then passes this text to a speech synthesis engine to generate a voice description that sounds more human, adding natural intonation and accents to make the description easier to understand.
[0930] 4. Audio guide distribution
[0931] The server provides the generated audio guide to the user's terminal in streaming or download format.
[0932] The terminal can play the provided audio description as secondary audio when the user watches the movie.
[0933] 5. Collecting User Feedback
[0934] The server collects feedback from users after they use the audio guide. Users can provide feedback on the quality and content of the guide, as well as areas for improvement. The server then adds this information to a profile and uses it to generate future guides.
[0935] Examples:
[0936] User A provides feedback stating that the emotional commentary section was very helpful and that he would like more specific scene descriptions. This feedback is reflected in the generation of guides for other movies he watches later.
[0937] Overall flow and summary
[0938] This system comprehensively handles everything from user registration to collecting and analyzing movie information, generating and distributing personalized audio guides, and collecting feedback. This allows for a movie-watching experience tailored to each user's needs, providing a high-value-added viewing experience not only for the visually impaired but also for a diverse audience. The introduction of this system is expected to open up new ways to enjoy movie-watching and provide a richer entertainment experience for many people.
[0939] The processing flow will be explained below.
[0940] Step 1: Registering a user
[0941] Users access the system and enter basic information such as name, age, language preference, viewing history, and movie preferences to create an account.
[0942] The server receives the entered information and creates and stores a user profile in a database.
[0943] Step 2: Update your profile
[0944] Users can update their profile information as their viewing history or preferences change.
[0945] The server stores the updated information in a database to keep the profile up to date.
[0946] Step 3: Get movie information
[0947] The user selects the movie they want to watch.
[0948] The server retrieves basic information such as the movie title, genre, cast, director, and running time from online databases and APIs.
[0949] The server also collects reviews, interviews, scripts, and subtitle files about the film.
[0950] Step 4: Analyze the scene
[0951] The server uses computer vision technology to analyze the movie's video data and identify key scenes and scenes in which characters appear.
[0952] The server uses NLP technology to analyze scripts and subtitle files to extract key dialogue and emotional nuances.
[0953] Step 5: Generate a personalized guide
[0954] The server decides how much detail to include in the audio description based on the user profile: detailed visual descriptions for the visually impaired, for example, or in-depth insights and background information for movie buffs.
[0955] The server creates the text for the audio guide based on the analysis results.
[0956] Step 6: Text-to-Speech
[0957] The server passes the generated audio guide text to a speech synthesis engine, which generates a voice that sounds like a human voice, with natural intonation and accent.
[0958] Step 7: Distributing the audio guide
[0959] The server provides the generated audio guide to the user's terminal in streaming or download format.
[0960] The terminal stores the provided audio guide and keeps it available for the user to use while watching the movie.
[0961] Step 8: Gather feedback
[0962] After watching a movie, users provide feedback on the quality and content of the audio description.
[0963] The server collects the provided feedback and stores it in a database, which is used to generate future guides.
[0964] Step 9: Improve your profile
[0965] The server analyzes the collected feedback and updates the user profile. This updated information is reflected in the next guide generation, providing an audio guide that is more tailored to the user's preferences.
[0966] Through the specific processing steps described above, the system provides a personalized movie-watching experience for each user, and meets the needs of visually impaired people and diverse audiences.
[0967] Example 1
[0968] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0969] Conventional movie audio guide systems have difficulty providing a fully personalized guide for the visually impaired or users who prefer specific movie genres. This limits the movie-watching experience, especially for users who cannot visually confirm the details of each movie scene or the emotional expressions of characters. Furthermore, it is difficult to effectively reflect user feedback and improve the guide, and further technical solutions are needed to increase user satisfaction.
[0970] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0971] In this invention, the server includes means for collecting user preferences and viewing history and creating a profile, means for acquiring and analyzing basic movie information, reviews, interviews, scripts, and subtitles, means for generating a personalized audio guide based on the user profile, means for analyzing the collected movie data and extracting important dialogue and emotional nuances using scene analysis and natural language processing techniques, means for generating the generated audio guide using a speech synthesis engine, means for delivering the generated audio guide to the user's terminal, and means for collecting user feedback and reflecting it in the profile. This allows for the provision of a detailed and personalized movie guide tailored to the user's needs, enabling a variety of users, including visually impaired people, to enjoy a richer movie-watching experience.
[0972] A "user profile" is a collection of data created based on a user's preferences, viewing history, basic information, and so on.
[0973] "Basic movie information" is basic information about a movie, such as the movie title, director, cast, and release date.
[0974] A "review" is text data that describes the ratings, impressions, and opinions posted by people who have watched a movie.
[0975] An "interview article" is text data that describes the contents of interviews with cast members, directors, and others involved in the film.
[0976] A "script" is text data that describes a movie script or screenplay.
[0977] "Subtitles" are text data that contains movie lines and explanations in written form.
[0978] "Scene analysis" is the process of analyzing each scene in a film to identify important scenes and characters.
[0979] "Natural language processing" is a technology that allows computers to understand, analyze, and generate human language.
[0980] An "audio description" is a guide that supplements visual information and provides audio information about the contents of a film.
[0981] A "speech synthesis engine" is a software system that generates natural-sounding speech based on text data.
[0982] "Feedback" refers to information such as evaluations, opinions, and impressions collected from users.
[0983] "Personalization" refers to a state that is customized based on the characteristics and preferences of an individual user.
[0984] The present invention is a system for generating a personalized movie audio guide that is adapted to the user's situation and preferences. Specific embodiments of this system will be described below.
[0985] System configuration
[0986] The system mainly consists of a server and a user device. The server creates user profiles, collects and analyzes movie data, generates audio guides, and collects feedback. The user device selects movies, plays audio guides, and provides feedback. Natural language processing technology and speech synthesis engines are used as software. Specific examples include AWS's Amazon Polly and Google Cloud Text-to-Speech.
[0987] Collecting user information
[0988] When a user first uses the system, the server asks the user to register an account. The user enters basic information such as name, age, language preference, viewing history, movie preferences, etc. The server then creates a user profile and stores it in a database.
[0989] Examples:
[0990] User A creates an account and registers in his profile information that he is visually impaired, likes mystery movies, and seeks deep insights.
[0991] Get movie information
[0992] The server retrieves basic information about the movie selected by the user from online databases and APIs, as well as public reviews and interviews with the cast and director. It also retrieves the movie's script and subtitle files and prepares them as data for analysis.
[0993] Examples:
[0994] When User A selects the movie "Sherlock Holmes," the server collects public information, reviews, scripts, etc. about the movie.
[0995] Audio guide generation
[0996] The server analyzes the collected movie data, recognizing important scenes and scenes in which characters appear through scene analysis, and extracting important lines and emotional nuances using natural language processing technology.
[0997] It then generates personalized audio description text based on the user profile, providing detailed visual descriptions for the visually impaired and in-depth insights and context for movie buffs. The generated text is then passed to a speech synthesis engine (e.g., AWS's Amazon Polly, Google Cloud Text-to-Speech) to generate a human-like audio description.
[0998] Examples:
[0999] A detailed audio guide for each scene is generated for User A, and text is created that provides detailed explanations of the facial expressions and emotional changes of the main cast, as well as hints at the mystery. An example of a prompt would be, "User A is visually impaired and loves mystery movies. Please generate a detailed audio guide for each scene of the movie 'Sherlock Holmes.' Please provide detailed explanations of the facial expressions and emotional changes of the main cast, as well as hints at the mystery."
[1000] Audio guide distribution
[1001] The server provides the generated audio guide to the user's terminal in a streaming or download format, and the terminal can play the provided audio guide as a secondary audio when the user watches the movie.
[1002] Collecting user feedback
[1003] The server collects feedback from users after they use the audio guide. Users can provide feedback on the quality and content of the guide, as well as areas for improvement. The server then adds this information to a profile and uses it to generate future guides.
[1004] Examples:
[1005] User A provides feedback stating that the emotional commentary section was very helpful and that he would like more specific scene descriptions. This feedback is reflected in the generation of guides for other movies he watches later.
[1006] In this way, the system can comprehensively handle everything from user registration to collecting and analyzing movie information, generating and distributing personalized audio guides, and collecting feedback, making it possible to provide a high-value movie-watching experience for a diverse audience, including the visually impaired.
[1007] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1008] Step 1:
[1009] Account Registration
[1010] When a user first accesses the system, the server displays an account registration form. The user enters basic information such as name, age, language preference, viewing history, and movie preferences, and then presses the submit button. The server receives the input data and creates a user profile. The created user profile is stored in a database.
[1011] Input: Basic information entered by the user (name, age, language preference, viewing history, movie preferences)
[1012] Output: User profile stored in the database
[1013] Specific behavior:
[1014] The server displays a registration form through a user interface. The user fills in the form and presses the submit button. The server takes the input data, generates a user profile, and stores it in a database.
[1015] Step 2:
[1016] Movie Selection
[1017] The user selects a movie, enters the movie title in the search box on the user terminal, and presses the search button, and the selected movie title is sent to the server.
[1018] Input: The movie title selected by the user
[1019] Output: Movie title sent to the server
[1020] Specific behavior:
[1021] The user terminal displays a movie title entry form, and when the user enters a title and presses the search button, the selection is sent to the server.
[1022] Step 3:
[1023] Movie data collection
[1024] Based on the received movie title, the server uses online databases and APIs to collect related data such as basic movie information, reviews, interviews, scripts, subtitles, etc. The collected data is then stored for analysis.
[1025] Input: Movie title
[1026] Output: Data stored on the server, including basic movie information, reviews, interviews, scripts, subtitles, etc.
[1027] Specific behavior:
[1028] The server calls a specific API, sends the movie title, and retrieves the relevant data, which is then stored in a database for analysis.
[1029] Step 4:
[1030] Data analysis and audio guide generation
[1031] The server analyzes the stored movie data, performs scene analysis to identify key scenes and character appearances, and uses natural language processing techniques to extract key lines and emotional nuances. Based on the extracted data, it generates personalized audio description text that matches the user profile. The generated text is then passed to a speech synthesis engine.
[1032] Input: saved movie data, user profile
[1033] Output: Personalized audio guide text and audio files
[1034] Specific behavior:
[1035] The server runs an algorithm that analyzes the movie scenes to determine their importance. It then uses natural language processing techniques to analyze the dialogue and extract key elements. It then inputs the prompt into a generative AI model to generate personalized text in a specific format. Finally, it passes the text to a speech synthesis engine to generate an audio file.
[1036] Step 5:
[1037] Audio guide distribution
[1038] The server provides the generated audio guide to the user's device in a streaming or download format, and the device can play the provided audio guide as a secondary audio when the user watches the movie.
[1039] Input: Personalized audio guide audio file
[1040] Output: Audio guide delivered to the user's device
[1041] Specific behavior:
[1042] The server uploads the generated audio guide to cloud storage and generates an access URL, which the user's device uses to play or download the audio guide.
[1043] Step 6:
[1044] Gathering feedback
[1045] After the user has finished watching, the server provides a feedback form. The user can enter their evaluation of the quality and content of the guide and suggestions for improvement in the feedback form and submit it. The server then reflects the collected feedback in the profile and uses it for generating audio guides from the next time onwards.
[1046] Input: User feedback
[1047] Output: Updated user profile
[1048] Specific behavior:
[1049] After the user has finished watching, the server displays a feedback form. The user enters their feedback in the form and presses the submit button. The server receives the feedback data and updates the user profile.
[1050] (Application example 1)
[1051] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1052] Conventional movie audio guide systems have not adequately provided personalized guides based on the viewer's individual preferences and viewing history. As a result, it has been difficult to provide an optimal viewing experience for visually impaired people or users who prefer specific movie genres. Furthermore, continuous improvement is required due to insufficient improvement in the quality of generated audio guides and profile updates based on user feedback.
[1053] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1054] In this invention, the server includes means for collecting user preferences and viewing history and creating a profile, means for acquiring and analyzing basic movie information, reviews, interviews, scripts, and subtitles, means for generating a personalized audio guide based on the user profile, means for generating the generated audio guide using a speech synthesis engine, means for delivering the generated audio guide to the user's terminal, means for generating text for the personalized audio guide based on the user profile and movie data using a generative AI model, means for calling the generative AI model using a prompt sentence in generating the audio guide, and means for converting the generated audio guide into a format suitable for smart devices and delivering it, thereby enabling the provision of an audio guide optimized for each user.
[1055] "User preferences" refers to information that indicates the user's tastes in movie genres and specific movies.
[1056] "Viewing history" refers to data that records the history of movies and dramas that a user has watched.
[1057] A "profile" refers to a collection of information that compiles a user's personal information, viewing history, preferences, etc.
[1058] "Basic movie information" refers to basic data such as the movie title, director, cast, release year, and genre.
[1059] A "review" is a text that shows criticism or evaluation of a movie, and is a source of information that allows users to learn the opinions of other viewers and experts.
[1060] An "interview article" refers to a written record of an interview conducted with a film's cast, director, or production staff.
[1061] A "script" refers to a film script or screenplay, a document that contains the actual dialogue and scene details of a film.
[1062] "Subtitles" refers to text that visually displays dialogue or narration in a film.
[1063] "Personalized audio description" refers to audio description that is customized based on a user's individual profile.
[1064] A "speech synthesis engine" is a program for converting text into voice data.
[1065] "Generative AI Model" refers to a model that uses artificial intelligence technology to generate personalized guides based on a user's profile and movie data.
[1066] A "prompt" is an instruction or question input to a generative AI model, and is text used to control the model's output.
[1067] "Smart devices" refers to electronic devices with advanced functions such as smartphones and smart glasses.
[1068] This invention is a system that generates and distributes personalized audio guides that are tailored to the user's circumstances and preferences. The program processing flow of this system is explained below with specific examples.
[1069] 1. Collection of User Information
[1070] When a user first uses the system, the server asks the user to register an account. The user enters basic information such as name, age, language preference, viewing history, movie preferences, etc. The server then creates a user profile and stores it in a database.
[1071] 2. Obtaining movie information
[1072] The server retrieves basic information about the movie selected by the user from online databases and APIs, including the movie title, director, cast, reviews, interviews, script, and subtitle files, and prepares this data for analysis.
[1073] 3. Audio guide generation
[1074] The server analyzes the collected movie data and uses NLP techniques to extract key scenes and emotional nuances. It then generates personalized audio description text based on the user profile. A generative AI model is used to generate the text, and a prompt is input to the model. For example, a prompt might be in the form: "User profile: {Name: User A, Age: 30, Language: Japanese, Favorite movie genre: Mystery} Movie data: {Title: Inception, Director: Christopher Nolan, Cast: {Leonardo DiCaprio: Dom Cobb}, Review: An extremely complex and intricately structured story...} Please generate a personalized movie description."
[1075] After the audio description text is generated, the server passes it to a speech synthesis engine to generate a human-like audio description, adding natural intonation and accents to make it easier to listen to.
[1076] 4. Audio guide distribution
[1077] The generated audio guide is provided to the user's smart device, such as a smartphone or smart glasses, via streaming or download. The server converts the audio guide into a format optimized for the user's device and delivers it appropriately.
[1078] 5. Collecting User Feedback
[1079] Users can provide feedback after using the audio guide. The server collects and analyzes this feedback. The collected feedback is reflected in the user profile and is used to generate audio guides in the future.
[1080] This system makes it possible to provide a personalized movie audio guide tailored to the user's situation and preferences, providing a richer movie-watching experience.
[1081] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1082] Step 1:
[1083] The server requires a user to register an account when using the system for the first time. The user enters basic information such as name, age, language setting, viewing history, and movie preferences. Based on this information, the server creates a user profile and stores it in a database. Input data: name, age, language setting, viewing history, and movie preferences. Output data: user profile.
[1084] Step 2:
[1085] When a user selects a movie to watch, the server retrieves basic information about the movie, including the movie title, director, cast, reviews, interviews, script, and subtitle files. Data is collected from online databases and APIs. Input data: movie title. Output data: movie basic information, reviews, interviews, script, and subtitle files.
[1086] Step 3:
[1087] The server analyzes the collected movie data using NLP technology, extracting key scenes and emotional nuances. Input data: Basic movie information, reviews, interviews, scripts, and subtitle files. Output data: Analyzed movie data.
[1088] Step 4:
[1089] Using the generative AI model, the server generates personalized audio guide text based on the user profile and analyzed movie data. A prompt is input to the generative AI model. For example, a prompt in the following format might be used: "User profile: {Name: User A, Age: 30, Language: Japanese, Favorite movie genre: Mystery} Movie data: {Title: Inception, Director: Generic name, Cast: {Starring actor: Lead role}, Review: Extremely complex and intricately structured story} Please generate a personalized movie audio guide." Input data: User profile, analyzed movie data. Output data: Audio guide text.
[1090] Step 5:
[1091] The server passes the generated audio description text to a speech synthesis engine to generate an audio description that sounds like a human voice, with natural intonation and accent. Input data: Audio description text. Output data: Audio data.
[1092] Step 6:
[1093] The server provides the generated audio guide to the user's smart device, such as a smartphone or smart glasses, in streaming or download format. The audio guide is converted into a format optimized for the user's device. Input data: Audio data. Output data: Audio data in a format suitable for the smart device.
[1094] Step 7:
[1095] Users can provide feedback after using the audio guide. The server collects and analyzes this feedback. The collected feedback is reflected in the user profile and is used to generate audio guides from next time onwards. Input data: User feedback. Output data: Updated user profile.
[1096] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1097] This invention is a system that generates a personalized movie audio guide tailored to the user's situation and preferences, and further combines an emotion engine to recognize the user's emotions and dynamically adjust the content of the audio guide. The program processing flow of this system is explained below with concrete examples.
[1098] Program processing overview
[1099] 1. Collecting User Information
[1100] When a user first uses the system, the server asks the user to register an account. The user enters basic information such as name, age, language preference, viewing history, and movie preferences, and also asks for consent to emotion recognition. Based on this information, the server creates a user profile and stores it in a database.
[1101] Examples:
[1102] User A creates an account and registers in his profile information that he is visually impaired, likes mystery movies, and seeks deep insights. He also agrees to emotion recognition.
[1103] 2. Obtaining movie information
[1104] The user selects a movie they want to watch. The server retrieves basic information about the movie, such as the title, genre, cast, director, and running time, from online databases and APIs. It also collects public reviews, interviews with the cast and director, scripts, and subtitle files.
[1105] 3. Scene Analysis
[1106] The server uses computer vision technology to analyze the film's video data to identify key scenes and character appearances, and NLP technology to analyze the script and subtitle files to extract key dialogue and emotional nuances.
[1107] 4. User Emotion Recognition by Emotion Engine
[1108] The server uses an emotion engine to recognize emotions from the user's facial expressions and voice in real time. While the user is watching a movie, the device uses a camera and microphone to capture the user's facial expressions and voice and transmits the data to the server.
[1109] Examples:
[1110] While User A is watching a movie, the emotion engine recognizes that he is crying at a touching scene. This information is sent to the server.
[1111] 5. Generate and refine personalized guides
[1112] The server generates personalized audio guide text based on the user profile and the recognition results of the emotion engine. Furthermore, it dynamically adjusts the content of the audio guide according to the recognized emotion. For example, if the user is emotional, it adds commentary that reflects that emotion.
[1113] Examples:
[1114] If user A is in tears during a touching scene, the audio guide for that scene will add commentary that delves deeper into the character's emotions.
[1115] 6. Speech Synthesis
[1116] The server then passes the generated audio guide text to a speech synthesis engine, which generates a voice that sounds like a human voice, with natural intonation and accent.
[1117] 7. Audio guide distribution
[1118] The server provides the generated audio guide to the user's terminal in streaming or download format, and the terminal stores the provided audio guide so that the user can use it while watching the movie.
[1119] 8. Gathering Feedback
[1120] After watching a movie, users provide feedback on the quality and content of the audio guide. The server collects the feedback and stores it in a database. The feedback is used to generate future guides.
[1121] 9. Profile Improvements
[1122] The server analyzes the collected feedback and updates the user profile. This updated information is reflected in the next guide generation, providing an audio guide that is more tailored to the user's preferences.
[1123] Overall flow and summary
[1124] This system comprehensively handles everything from user registration to collecting and analyzing movie information, real-time emotion recognition using an emotion engine, generating and adjusting personalized audio guides, delivering them, and collecting feedback. This provides a movie-watching experience tailored to each user's needs and emotions, and addresses the needs of visually impaired people and a diverse audience. The introduction of this system is expected to open up new ways to enjoy movie-going and provide a richer entertainment experience for many people.
[1125] The processing flow will be explained below.
[1126] Step 1: Registering a user
[1127] Users access the system and enter their name, age, language preference, viewing history, favorite movie genre, preferred level of visual and auditory assistance, and consent to emotion recognition on the account creation screen.
[1128] The server receives the information entered by the user and stores it in a database as a new user profile.
[1129] Step 2: Update your profile
[1130] Users can update their profile to reflect changes to their viewing history or new movie preferences by entering and saving changes on the profile page.
[1131] The server updates and saves the changes made by the user to the profile information in the database.
[1132] Step 3: Get movie information
[1133] The user selects the movie they want to watch from within the system.
[1134] The server retrieves basic information about the selected movie, such as title, genre, cast, director, and running time, from online databases or APIs.
[1135] The server also collects reviews of the film, interviews with the cast and director, scripts, and subtitle files.
[1136] Step 4: Analyze the scene
[1137] The server uses computer vision technology to analyze the movie's video data and identify key scenes and scenes in which characters appear.
[1138] The server uses NLP technology to analyze scripts and subtitle files to extract key dialogue and emotional nuances.
[1139] Step 5: Recognizing user emotions with the emotion engine
[1140] The device uses a built-in camera and microphone to capture the user's facial expressions and voice in real time while the movie is playing.
[1141] The terminal sends the captured data to an emotion engine to recognize the user's emotional state.
[1142] The emotion engine (server) analyzes the received data, identifies the user's emotion (such as joy, sadness, surprise, etc.), and generates an emotion status.
[1143] Step 6: Generate and refine your personalized guide
[1144] The server generates personalized audio guide text based on the user profile and the recognition results of the emotion engine.
[1145] The server dynamically adjusts the content of the audio guide depending on the recognized emotion, for example, by adding explanations to calm the user or by emphasizing emotional background information for moving scenes.
[1146] Examples:
[1147] If user A is in tears during a moving scene in a movie, the server can incorporate additional emotional context into the audio description to enhance the viewing experience of that scene.
[1148] Step 7: Text-to-Speech
[1149] The server then passes the generated audio guide text to a speech synthesis engine, which generates a voice that sounds similar to a natural human voice, with natural intonation and accent.
[1150] Step 8: Distributing the audio guide
[1151] The server provides the generated audio guide to the user's terminal in streaming or download format.
[1152] The terminal stores the provided audio description and prepares it for playback as secondary audio during movie playback.
[1153] Step 9: Gather feedback
[1154] After watching a movie, users can enter their feedback on the quality and content of the audio guide on a feedback page within the system.
[1155] The server receives the provided feedback and stores it in a database, which is used to generate subsequent guides.
[1156] Step 10: Refine your profile
[1157] The server analyzes the collected feedback and updates the user profile. This updated information is reflected in the next guide generation, providing an audio guide that is more tailored to the user's preferences.
[1158] Overall flow and summary
[1159] This process comprehensively covers user registration, movie information collection, real-time emotion recognition using an emotion engine, personalized and tailored audio guide generation and delivery, and feedback collection. By dynamically reflecting user emotions, the system can promote a more personalized movie-watching experience and cater to a diverse audience.
[1160] Example 2
[1161] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1162] There is a demand for a variety of viewers, including the visually impaired, to enrich their movie experience and provide personalized guidance that reflects the user's individual emotions and preferences. Conventional systems have difficulty providing guidance that reflects the user's emotions in real time, resulting in only a uniform guidance that makes it difficult to increase user satisfaction.
[1163] The specification process by the specification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for collecting a user's preferences and viewing history and creating a profile; means for acquiring and analyzing basic movie information, reviews, interviews, scripts, and subtitles; means for analyzing visual information using computer vision technology and identifying important scenes and scenes in which characters appear; means for capturing the user's facial expressions and voice through a terminal while watching a movie and recognizing the user's emotions using an emotion engine; means for generating and dynamically adjusting a personalized audio guide based on the user profile and the recognized emotions; means for generating the generated audio guide using a speech synthesis engine; and means for delivering the generated audio guide to the user's terminal. This makes it possible to provide a personalized movie experience tailored to the user's emotions and preferences.
[1164] "User preferences" refer to the tastes and preferences that a user has for a particular genre or content.
[1165] A "viewing history" is a record of movies and programs a user has watched.
[1166] A "profile" is a collection of personal information such as a user's name, age, language preference, viewing history, and movie preferences.
[1167] "Basic movie information" refers to information such as the movie title, genre, cast, director, and screening time.
[1168] A "review" is a piece of writing that expresses an opinion or rating about a movie.
[1169] An "interview article" is a written version of interviews with cast members, directors, and others involved in the film.
[1170] A "script" is a screenplay that contains the dialogue and direction of a movie.
[1171] "Subtitles" are the written versions of the audio in a movie that are displayed on the screen.
[1172] "Computer vision technology" is a technology that allows computers to analyze the content of images and videos.
[1173] An "emotion engine" is a software technology that recognizes emotions by analyzing a user's facial expressions and voice.
[1174] A "speech synthesis engine" is a technology that generates speech by adding natural intonation and accent based on text information.
[1175] A "terminal" is a device that a user uses to watch a movie, including a smartphone, tablet, or PC.
[1176] The present invention provides a system for generating a personalized movie audio guide based on a user's preferences, viewing history, and emotions, thereby improving the user's movie-watching experience. Specific embodiments of the present invention will be described below.
[1177] Hardware and software used
[1178] server:
[1179] Database: Used to store user profiles, movie information, feedback, etc.
[1180] Computer vision technology: Used to analyze film footage and identify key scenes and character appearances.
[1181] Emotion engine: Used to analyze the user's facial expressions and voice to recognize emotions in real time.
[1182] Speech synthesis engine: Used to convert text to speech for personalized audio descriptions.
[1183] Device:
[1184] Camera: Used to capture the user's facial expressions.
[1185] Microphone: Used to capture the user's voice.
[1186] Streaming technology: Used to play audio descriptions provided by the server.
[1187] Explanation of program processing
[1188] Collecting user information
[1189] When a user first uses the system, the server displays an account registration page where the user enters information such as name, age, language preference, viewing history, and preferred genres, and also selects consent to emotion recognition. The server then creates a user profile based on this information and stores it in a database.
[1190] Examples:
[1191] When User A enters information into the registration form, the server receives it, generates a profile and stores it in the database.
[1192] Get movie information
[1193] When a user selects a movie they want to watch, the server retrieves basic information about the movie (title, genre, cast, director, and running time) from an online database, as well as public reviews, interviews with the cast and director, scripts, and subtitle files.
[1194] Examples:
[1195] When User A selects "Mystery Movie X," the server collects basic information about the movie and any associated scripts and subtitle files.
[1196] Scene Analysis
[1197] The server uses computer vision technology to analyze the film's video data to identify key scenes and character appearances, and NLP technology to analyze the script and subtitle files to extract key dialogue and emotional nuances.
[1198] Examples:
[1199] The server breaks down the data for Movie X frame by frame to identify key scenes and characters, while also analyzing the script to extract key lines and moving moments.
[1200] Recognizing user emotions with an emotion engine
[1201] The server uses an emotion engine to recognize emotions in real time from the user's facial expressions and voice. The device uses a camera and microphone to capture the user's facial expressions and voice while watching a movie, and sends the data to the server.
[1202] Examples:
[1203] When user A sheds tears during a touching scene, the device's camera captures this and sends the data to the server.
[1204] Generate and refine personalized guides
[1205] The server generates personalized audio guide text based on the user profile and the recognition results of the emotion engine, and dynamically adjusts the content of the audio guide based on the recognized emotions.
[1206] Examples:
[1207] The server recognizes that user A is moved and adds a specific example to the audio guide for that scene, such as "Character Y's deep emotions are expressed."
[1208] Speech synthesis
[1209] The server passes the generated audio guide text to a speech synthesis engine, which generates natural-sounding speech.
[1210] Examples:
[1211] The server passes the text "Character Y's deep emotions are expressed" to a speech synthesis engine, which generates realistic speech.
[1212] Audio guide distribution
[1213] The server provides the generated audio guide to the user's device in streaming or download format, and the device prepares to play the provided audio guide.
[1214] Examples:
[1215] The server sends the generated audio guide to the terminal, which prepares it for playback.
[1216] Gathering feedback
[1217] After watching a movie, users provide feedback on the quality and content of the audio guide, which the server collects and stores in a database.
[1218] Examples:
[1219] User A provides feedback saying, "The audio guide was easy to understand," and the server receives this and stores it in the database.
[1220] Profile Improvements
[1221] The server analyzes the feedback and updates the user profile, which is then reflected in the next guide generation.
[1222] Examples:
[1223] The server analyzes User A's feedback, adds "I prefer detailed explanations" to his profile, and reflects this in the next audio guide.
[1224] Prompt Sentence Examples
[1225] "Imagine a situation where a user is watching a mystery movie and a touching scene occurs. This makes the user's face begin to shed tears. We would like an explanation of how the emotion engine should recognize this situation and what kind of voice guidance should be generated."
[1226] Examples of prompts:
[1227] "If a user is crying during an emotional scene in a mystery movie, show how an emotion engine can recognize that emotion and add commentary that aligns with the user's emotions."
[1228] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1229] Step 1:
[1230] The server displays an account registration page when a user uses the system for the first time.
[1231] Input: The user enters basic information such as name, age, language preference, viewing history, and movie preferences, and selects consent for emotion recognition.
[1232] Data processing: The server receives this information and creates a user profile.
[1233] Output: The user profile is saved in the database.
[1234] Specific operation: When User A enters information into the registration form and clicks the submit button, the server receives it, generates a profile and saves it in the database.
[1235] Step 2:
[1236] The user selects the movie they want to watch.
[1237] Input: The user selects the movie they want to watch from the movie list.
[1238] Data processing: The server retrieves basic movie information (title, genre, cast, director, running time) from an online database, and collects public reviews, interviews with the cast and director, scripts, and subtitle files.
[1239] Output: The necessary movie information is stored on the server.
[1240] Specific operation: When user A selects the movie "Mystery Movie X," the server collects its basic information and related script and subtitle files.
[1241] Step 3:
[1242] The server analyzes the movie's video data using computer vision technology.
[1243] Input: Collected movie footage data
[1244] Data Processing: Video data is broken down frame by frame to identify key scenes and character appearances, and NLP techniques are used to analyze scripts and subtitle files to extract key dialogue and emotional nuances.
[1245] Output: A dataset containing identified scene and character information, as well as extracted dialogue and emotional nuances.
[1246] What it does: The server breaks down the data of Movie X frame by frame to identify important scenes and characters, while also analyzing the script to extract key lines and moving moments.
[1247] Step 4:
[1248] The device uses a camera and microphone to capture the user's facial expressions and voice while watching a movie, and sends the data to a server.
[1249] Input: facial expressions and voice data of the user while watching a movie
[1250] Data processing: Using an emotion engine, emotions are recognized in real time from the user's facial expressions and voice.
[1251] Output: Data about the user's emotional state.
[1252] Specific operation: When user A sheds tears during a touching scene, the device camera captures this and sends the data to the server.
[1253] Step 5:
[1254] The server generates personalized audio guide text based on the user profile and the recognition results of the emotion engine.
[1255] Input: User profile and real-time emotional data
[1256] Data processing: Generate text that is attuned to the user's preferences and emotions, and dynamically adjust content based on recognized emotions.
[1257] Output: Personalized audio guide text.
[1258] Specific operation: The server recognizes that user A is moved and adds a specific example to the audio description of that scene, such as "Character Y's deep emotions are expressed."
[1259] Step 6:
[1260] The server passes the data to a speech synthesis engine to generate speech from the text.
[1261] Input: The generated audio description text
[1262] Data processing: A speech synthesis engine converts text into speech and adds natural intonation and accent.
[1263] Output: Audio data that sounds similar to a human voice.
[1264] Specific operation: The server passes the text "Character Y's deep emotions are expressed" to the speech synthesis engine, which generates realistic speech.
[1265] Step 7:
[1266] The server provides the generated audio guide to the user's terminal in streaming or download format.
[1267] Input: Generated audio data
[1268] Data processing: The generated audio is converted into an appropriate format and sent to the user's device.
[1269] Output: Audio guide stored on the user's device.
[1270] Specific operation: The server sends the generated audio guide to the terminal, which prepares it for playback.
[1271] Step 8:
[1272] Users provide feedback after watching a movie.
[1273] Input: Feedback information about the quality and content of the audio description
[1274] Data processing: The server stores the collected feedback in a database.
[1275] Output: A database containing the feedback.
[1276] Specific operation: User A provides feedback such as "The audio guide was easy to understand," and the server receives this and stores it in the database.
[1277] Step 9:
[1278] The server analyzes the feedback and updates the user profile.
[1279] Input: Saved Feedback
[1280] Data processing: Analyze feedback and add user preferences and trends to your profile.
[1281] Output: The updated user profile.
[1282] Specific operation: The server analyzes user A's feedback, adds "I prefer detailed explanations" to his profile, and reflects this in the next audio guide.
[1283] (Application example 2)
[1284] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1285] Conventional movie viewing systems have limited means to accommodate visually impaired users and users who want to understand specific scenes or information in depth. In particular, the lack of a function to provide personalized audio guides that are sensitive to the user's emotions results in a poor movie-viewing experience. Furthermore, there is a lack of a mechanism to learn from user feedback and improve the guide content for future viewings. This makes it difficult to improve user satisfaction.
[1286] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1287] In this invention, the server includes means for collecting user preferences and viewing history and creating a profile, means for acquiring and analyzing basic movie information, ratings, related party information, dialogue, and commentary, means for generating a personalized audio guide based on the user profile, means for recognizing the user's emotions and dynamically adjusting the content of the audio guide, means for generating the generated audio guide using a speech synthesis engine, means for delivering the generated audio guide to the user's device, means for analyzing the collected feedback and updating the user profile, means for analyzing visual information using image analysis technology and identifying important scenes and scenes with characters, and means for generating prompt sentences using a generative AI model. This allows visually impaired people and emotionally sensitive users to better understand and enjoy movies.
[1288] "User" refers to a person who uses the system to watch a movie.
[1289] "Preferences" refer to personal preferences that a user has for particular movie genres or content.
[1290] "Viewing history" refers to historical information about movies and video content that a user has viewed in the past.
[1291] A "profile" refers to a data set that compiles information such as a user's preferences and viewing history.
[1292] "Basic movie information" refers to basic data about a movie, such as the title, genre, cast, director, and running time.
[1293] "Rating" refers to the ratings and reviews of movies by users and experts.
[1294] "Related Person Information" refers to detailed information about people such as the film's cast and director.
[1295] "Dialogue" refers to the words spoken by characters in a film.
[1296] "Explanatory text" refers to text that explains the content and background of the film.
[1297] "Emotion" refers to the psychological state a user experiences while watching a movie.
[1298] "Audio guide" refers to audio guidance or explanations provided to the user.
[1299] "Dynamic adjustment" refers to changing content in response to real-time conditions.
[1300] "Speech synthesis engine" refers to software or hardware for converting text data into speech data.
[1301] "Terminal" refers to the electronic device that a user uses to watch movies and use the system.
[1302] "Feedback" refers to opinions and evaluations from users.
[1303] "Image analysis technology" refers to technology for processing visual information and extracting specific patterns and features.
[1304] A "scene" refers to a specific situation or situation within a film.
[1305] A "character scene" refers to a specific moment in a film in which a character appears.
[1306] "Generative AI models" refer to artificial intelligence techniques for performing tasks such as text generation.
[1307] A "prompt sentence" refers to an input sentence that causes a generative AI model to generate text.
[1308] The present invention is a system for improving a user's viewing experience, which includes the following means: A server collects a user's preferences and viewing history and creates a profile, which includes information provided by the user such as name, age, language, viewing history, and movie preferences. The profile is stored in a database.
[1309] The server then retrieves and analyzes information such as the movie title, genre, cast, director, running time, ratings, people involved, dialogue, and description from online databases and APIs, and is then ready to provide detailed movie information to the user.
[1310] Based on the user profile, the server generates a personalized audio guide. To do so, it uses NLP technology to analyze the movie script and subtitles and extract important lines and scenes. It also recognizes the user's emotions in real time and dynamically adjusts the content of the audio guide based on that information. Emotion recognition is performed using, for example, Microsoft Azure Cognitive Services and Google Cloud's Natural Language API. This allows the server to provide commentary that matches the emotions the user feels while watching the movie.
[1311] The generated audio description is converted into speech with natural intonation and accent by a speech synthesis engine (e.g., Amazon Polly or Google Text-to-Speech), and is delivered to the user's device in streaming or download format.
[1312] Users provide feedback after watching a movie, which is collected and analyzed by the server, and their profile is updated to reflect this in the generation of personalized audio guides for future viewings.
[1313] Computer vision techniques (e.g., OpenCV) are used to analyze visual information and identify important scenes and character appearances in the film, which is also incorporated into the personalized audio description.
[1314] The generative AI model is used to generate prompts. An example of a prompt is, "Based on user profile: {'name': 'Mr. A', 'age': 30, 'language': 'jp', 'preferences': ['mystery']}, movie data: {detailed movie information}, emotion: {detailed emotion information}, please generate an in-depth explanation of the key scenes in this movie."
[1315] This allows visually impaired people and emotionally sensitive users to understand and enjoy movies more deeply. This system is ideal for content distribution services and is expected to significantly improve the movie-watching experience.
[1316] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1317] Step 1:
[1318] The server collects the user's preferences and viewing history and creates a profile. The input is information provided by the user about their name, age, language, viewing history, and movie preferences. The information is stored in a database and a user profile is generated. Specifically, the user enters information into an account creation form, which is received by the server.
[1319] Step 2:
[1320] The server retrieves basic movie information (title, genre, cast, director, running time), ratings, related party information, dialogue, and description from an online database or API. The input is the movie title and ID, and the output is the retrieved movie details. The server sends the API request and stores the retrieved data in a database for analysis.
[1321] Step 3:
[1322] The server uses computer vision technology to analyze the visual information of a movie and identify important scenes and scenes featuring characters. The input is the movie's video data, and the output is metadata for the identified scenes. Specifically, OpenCV is used to analyze the video data frame by frame and extract characteristic scenes.
[1323] Step 4:
[1324] The server uses NLP technology to analyze movie scripts and subtitles and extract important lines and scenes. The input is the text data of the script or subtitles, and the output is a list of important lines and scenes. Specifically, it uses NLP libraries such as spaCy to perform text analysis and extract important lines.
[1325] Step 5:
[1326] The device uses a camera and microphone to capture the user's facial expressions and voice and sends the data to a server. An emotion recognition engine analyzes this data. The input is the user's video and audio data, and the output is recognized emotional information. Specifically, the device captures the user's facial expressions and voice at regular intervals and sends them to the server in real time.
[1327] Step 6:
[1328] The server generates a personalized audio guide based on the user profile and real-time emotional data. A prompt is input into the generative AI model to generate the audio guide text. The input is profile information, emotional data, and the prompt, and the output is the audio guide text. An example of a specific prompt is, "Based on user profile: {'name': 'Mr. A', 'age': 30, 'language': 'jp', 'preferences': ['Mystery']}, movie data: {detailed movie information}, emotion: {detailed emotion information}, please generate an in-depth commentary on the important scenes in this movie."
[1329] Step 7:
[1330] The server passes the personalized audio guide text to a speech synthesis engine to generate a voice that sounds similar to a human voice. The input is the audio guide text data, and the output is the audio data. Specifically, the server converts the text to speech using a speech synthesis engine such as Google Text-to-Speech (gTTS).
[1331] Step 8:
[1332] The server distributes the generated audio guide to the user's device. The input is the generated audio data, and the output is the audio guide stored on the user's device. Specifically, the audio guide is provided in streaming format or file download format.
[1333] Step 9:
[1334] After watching a movie, users provide feedback on the quality and content of the audio description. The server collects this feedback and stores it in a database. The input is the user's feedback data, and the output is an updated user profile. For example, opinions can be collected through a web form or an in-app feedback feature.
[1335] Step 10:
[1336] The server improves the user profile based on the feedback and analysis data, and reflects this in future personalized audio guides. The input is the collected feedback and viewing data, and the output is an updated user profile. Specifically, the server analyzes the feedback data and adds new user preference information to the profile.
[1337] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1338] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1339] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1340] [Fourth embodiment]
[1341] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1342] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1343] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1344] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1345] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1346] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1347] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1348] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1349] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1350] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1351] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1352] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1353] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1354] This invention is a system for generating a personalized movie audio guide that is tailored to the user's situation and preferences. The program processing flow of this system will be explained below with specific examples.
[1355] Program processing overview
[1356] 1. Collecting User Information
[1357] When a user first uses the system, the server asks the user to register an account. The user enters basic information such as name, age, language preference, viewing history, movie preferences, etc. The server then creates a user profile and stores it in a database.
[1358] Examples:
[1359] User A creates an account and registers in his profile information that he is visually impaired, likes mystery movies, and seeks deep insights.
[1360] 2. Obtaining movie information
[1361] The server retrieves basic information about the movie selected by the user from online databases and APIs, as well as public reviews and interviews with the cast and director. It also retrieves the movie's script and subtitle files and prepares them as data for analysis.
[1362] 3. Audio guide generation
[1363] The server analyzes the collected movie data, recognizing important scenes and character appearances through scene analysis, and extracts important dialogue and emotional nuances using NLP technology.
[1364] It then generates personalized audio description text based on the user profile, providing detailed image descriptions for the visually impaired and in-depth insights and background information for movie buffs.
[1365] Examples:
[1366] For User A, a detailed audio guide is generated for each scene, and text is created that provides detailed explanations of the main cast's facial expressions and emotional changes, as well as the foreshadowing of the mystery.
[1367] The server then passes this text to a speech synthesis engine to generate a voice description that sounds more human, adding natural intonation and accents to make the description easier to understand.
[1368] 4. Audio guide distribution
[1369] The server provides the generated audio guide to the user's terminal in streaming or download format.
[1370] The terminal can play the provided audio description as secondary audio when the user watches the movie.
[1371] 5. Collecting User Feedback
[1372] The server collects feedback from users after they use the audio guide. Users can provide feedback on the quality and content of the guide, as well as areas for improvement. The server then adds this information to a profile and uses it to generate future guides.
[1373] Examples:
[1374] User A provides feedback stating that the emotional commentary section was very helpful and that he would like more specific scene descriptions. This feedback is reflected in the generation of guides for other movies he watches later.
[1375] Overall flow and summary
[1376] This system comprehensively handles everything from user registration to collecting and analyzing movie information, generating and distributing personalized audio guides, and collecting feedback. This allows for a movie-watching experience tailored to each user's needs, providing a high-value-added viewing experience not only for the visually impaired but also for a diverse audience. The introduction of this system is expected to open up new ways to enjoy movie-watching and provide a richer entertainment experience for many people.
[1377] The processing flow will be explained below.
[1378] Step 1: Registering a user
[1379] Users access the system and enter basic information such as name, age, language preference, viewing history, and movie preferences to create an account.
[1380] The server receives the entered information and creates and stores a user profile in a database.
[1381] Step 2: Update your profile
[1382] Users can update their profile information as their viewing history or preferences change.
[1383] The server stores the updated information in a database to keep the profile up to date.
[1384] Step 3: Get movie information
[1385] The user selects the movie they want to watch.
[1386] The server retrieves basic information such as the movie title, genre, cast, director, and running time from online databases and APIs.
[1387] The server also collects reviews, interviews, scripts, and subtitle files about the film.
[1388] Step 4: Analyze the scene
[1389] The server uses computer vision technology to analyze the movie's video data and identify key scenes and scenes in which characters appear.
[1390] The server uses NLP technology to analyze scripts and subtitle files to extract key dialogue and emotional nuances.
[1391] Step 5: Generate a personalized guide
[1392] The server decides how much detail to include in the audio description based on the user profile: detailed visual descriptions for the visually impaired, for example, or in-depth insights and background information for movie buffs.
[1393] The server creates the text for the audio guide based on the analysis results.
[1394] Step 6: Text-to-Speech
[1395] The server passes the generated audio guide text to a speech synthesis engine, which generates a voice that sounds like a human voice, with natural intonation and accent.
[1396] Step 7: Distributing the audio guide
[1397] The server provides the generated audio guide to the user's terminal in streaming or download format.
[1398] The terminal stores the provided audio guide and keeps it available for the user to use while watching the movie.
[1399] Step 8: Gather feedback
[1400] After watching a movie, users provide feedback on the quality and content of the audio description.
[1401] The server collects the provided feedback and stores it in a database, which is used to generate future guides.
[1402] Step 9: Improve your profile
[1403] The server analyzes the collected feedback and updates the user profile. This updated information is reflected in the next guide generation, providing an audio guide that is more tailored to the user's preferences.
[1404] Through the specific processing steps described above, the system provides a personalized movie-watching experience for each user, and meets the needs of visually impaired people and diverse audiences.
[1405] Example 1
[1406] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1407] Conventional movie audio guide systems have difficulty providing a fully personalized guide for the visually impaired or users who prefer specific movie genres. This limits the movie-watching experience, especially for users who cannot visually confirm the details of each movie scene or the emotional expressions of characters. Furthermore, it is difficult to effectively reflect user feedback and improve the guide, and further technical solutions are needed to increase user satisfaction.
[1408] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1409] In this invention, the server includes means for collecting user preferences and viewing history and creating a profile, means for acquiring and analyzing basic movie information, reviews, interviews, scripts, and subtitles, means for generating a personalized audio guide based on the user profile, means for analyzing the collected movie data and extracting important dialogue and emotional nuances using scene analysis and natural language processing techniques, means for generating the generated audio guide using a speech synthesis engine, means for delivering the generated audio guide to the user's terminal, and means for collecting user feedback and reflecting it in the profile. This allows for the provision of a detailed and personalized movie guide tailored to the user's needs, enabling a variety of users, including visually impaired people, to enjoy a richer movie-watching experience.
[1410] A "user profile" is a collection of data created based on a user's preferences, viewing history, basic information, and so on.
[1411] "Basic movie information" is basic information about a movie, such as the movie title, director, cast, and release date.
[1412] A "review" is text data that describes the ratings, impressions, and opinions posted by people who have watched a movie.
[1413] An "interview article" is text data that describes the contents of interviews with cast members, directors, and others involved in the film.
[1414] A "script" is text data that describes a movie script or screenplay.
[1415] "Subtitles" are text data that contains movie lines and explanations in written form.
[1416] "Scene analysis" is the process of analyzing each scene in a film to identify important scenes and characters.
[1417] "Natural language processing" is a technology that allows computers to understand, analyze, and generate human language.
[1418] An "audio description" is a guide that supplements visual information and provides audio information about the contents of a film.
[1419] A "speech synthesis engine" is a software system that generates natural-sounding speech based on text data.
[1420] "Feedback" refers to information such as evaluations, opinions, and impressions collected from users.
[1421] "Personalization" refers to a state that is customized based on the characteristics and preferences of an individual user.
[1422] The present invention is a system for generating a personalized movie audio guide that is adapted to the user's situation and preferences. Specific embodiments of this system will be described below.
[1423] System configuration
[1424] The system mainly consists of a server and a user device. The server creates user profiles, collects and analyzes movie data, generates audio guides, and collects feedback. The user device selects movies, plays audio guides, and provides feedback. Natural language processing technology and speech synthesis engines are used as software. Specific examples include AWS's Amazon Polly and Google Cloud Text-to-Speech.
[1425] Collecting user information
[1426] When a user first uses the system, the server asks the user to register an account. The user enters basic information such as name, age, language preference, viewing history, movie preferences, etc. The server then creates a user profile and stores it in a database.
[1427] Examples:
[1428] User A creates an account and registers in his profile information that he is visually impaired, likes mystery movies, and seeks deep insights.
[1429] Get movie information
[1430] The server retrieves basic information about the movie selected by the user from online databases and APIs, as well as public reviews and interviews with the cast and director. It also retrieves the movie's script and subtitle files and prepares them as data for analysis.
[1431] Examples:
[1432] When User A selects the movie "Sherlock Holmes," the server collects public information, reviews, scripts, etc. about the movie.
[1433] Audio guide generation
[1434] The server analyzes the collected movie data, recognizing important scenes and scenes in which characters appear through scene analysis, and extracting important lines and emotional nuances using natural language processing technology.
[1435] It then generates personalized audio description text based on the user profile, providing detailed visual descriptions for the visually impaired and in-depth insights and context for movie buffs. The generated text is then passed to a speech synthesis engine (e.g., AWS's Amazon Polly, Google Cloud Text-to-Speech) to generate a human-like audio description.
[1436] Examples:
[1437] A detailed audio guide for each scene is generated for User A, and text is created that provides detailed explanations of the facial expressions and emotional changes of the main cast, as well as hints at the mystery. An example of a prompt would be, "User A is visually impaired and loves mystery movies. Please generate a detailed audio guide for each scene of the movie 'Sherlock Holmes.' Please provide detailed explanations of the facial expressions and emotional changes of the main cast, as well as hints at the mystery."
[1438] Audio guide distribution
[1439] The server provides the generated audio guide to the user's terminal in a streaming or download format, and the terminal can play the provided audio guide as a secondary audio when the user watches the movie.
[1440] Collecting user feedback
[1441] The server collects feedback from users after they use the audio guide. Users can provide feedback on the quality and content of the guide, as well as areas for improvement. The server then adds this information to a profile and uses it to generate future guides.
[1442] Examples:
[1443] User A provides feedback stating that the emotional commentary section was very helpful and that he would like more specific scene descriptions. This feedback is reflected in the generation of guides for other movies he watches later.
[1444] In this way, the system can comprehensively handle everything from user registration to collecting and analyzing movie information, generating and distributing personalized audio guides, and collecting feedback, making it possible to provide a high-value movie-watching experience for a diverse audience, including the visually impaired.
[1445] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1446] Step 1:
[1447] Account Registration
[1448] When a user first accesses the system, the server displays an account registration form. The user enters basic information such as name, age, language preference, viewing history, and movie preferences, and then presses the submit button. The server receives the input data and creates a user profile. The created user profile is stored in a database.
[1449] Input: Basic information entered by the user (name, age, language preference, viewing history, movie preferences)
[1450] Output: User profile stored in the database
[1451] Specific behavior:
[1452] The server displays a registration form through a user interface. The user fills in the form and presses the submit button. The server takes the input data, generates a user profile, and stores it in a database.
[1453] Step 2:
[1454] Movie Selection
[1455] The user selects a movie, enters the movie title in the search box on the user terminal, and presses the search button, and the selected movie title is sent to the server.
[1456] Input: The movie title selected by the user
[1457] Output: Movie title sent to the server
[1458] Specific behavior:
[1459] The user terminal displays a movie title entry form, and when the user enters a title and presses the search button, the selection is sent to the server.
[1460] Step 3:
[1461] Movie data collection
[1462] Based on the received movie title, the server uses online databases and APIs to collect related data such as basic movie information, reviews, interviews, scripts, subtitles, etc. The collected data is then stored for analysis.
[1463] Input: Movie title
[1464] Output: Data stored on the server, including basic movie information, reviews, interviews, scripts, subtitles, etc.
[1465] Specific behavior:
[1466] The server calls a specific API, sends the movie title, and retrieves the relevant data, which is then stored in a database for analysis.
[1467] Step 4:
[1468] Data analysis and audio guide generation
[1469] The server analyzes the stored movie data, performs scene analysis to identify key scenes and character appearances, and uses natural language processing techniques to extract key lines and emotional nuances. Based on the extracted data, it generates personalized audio description text that matches the user profile. The generated text is then passed to a speech synthesis engine.
[1470] Input: saved movie data, user profile
[1471] Output: Personalized audio guide text and audio files
[1472] Specific behavior:
[1473] The server runs an algorithm that analyzes the movie scenes to determine their importance. It then uses natural language processing techniques to analyze the dialogue and extract key elements. It then inputs the prompt into a generative AI model to generate personalized text in a specific format. Finally, it passes the text to a speech synthesis engine to generate an audio file.
[1474] Step 5:
[1475] Audio guide distribution
[1476] The server provides the generated audio guide to the user's device in a streaming or download format, and the device can play the provided audio guide as a secondary audio when the user watches the movie.
[1477] Input: Personalized audio guide audio file
[1478] Output: Audio guide delivered to the user's device
[1479] Specific behavior:
[1480] The server uploads the generated audio guide to cloud storage and generates an access URL, which the user's device uses to play or download the audio guide.
[1481] Step 6:
[1482] Gathering feedback
[1483] After the user has finished watching, the server provides a feedback form. The user can enter their evaluation of the quality and content of the guide and suggestions for improvement in the feedback form and submit it. The server then reflects the collected feedback in the profile and uses it for generating audio guides from the next time onwards.
[1484] Input: User feedback
[1485] Output: Updated user profile
[1486] Specific behavior:
[1487] After the user has finished watching, the server displays a feedback form. The user enters their feedback in the form and presses the submit button. The server receives the feedback data and updates the user profile.
[1488] (Application example 1)
[1489] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1490] Conventional movie audio guide systems have not adequately provided personalized guides based on the viewer's individual preferences and viewing history. As a result, it has been difficult to provide an optimal viewing experience for visually impaired people or users who prefer specific movie genres. Furthermore, continuous improvement is required due to insufficient improvement in the quality of generated audio guides and profile updates based on user feedback.
[1491] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1492] In this invention, the server includes means for collecting user preferences and viewing history and creating a profile, means for acquiring and analyzing basic movie information, reviews, interviews, scripts, and subtitles, means for generating a personalized audio guide based on the user profile, means for generating the generated audio guide using a speech synthesis engine, means for delivering the generated audio guide to the user's terminal, means for generating text for the personalized audio guide based on the user profile and movie data using a generative AI model, means for calling the generative AI model using a prompt sentence in generating the audio guide, and means for converting the generated audio guide into a format suitable for smart devices and delivering it, thereby enabling the provision of an audio guide optimized for each user.
[1493] "User preferences" refers to information that indicates the user's tastes in movie genres and specific movies.
[1494] "Viewing history" refers to data that records the history of movies and dramas that a user has watched.
[1495] A "profile" refers to a collection of information that compiles a user's personal information, viewing history, preferences, etc.
[1496] "Basic movie information" refers to basic data such as the movie title, director, cast, release year, and genre.
[1497] A "review" is a text that shows criticism or evaluation of a movie, and is a source of information that allows users to learn the opinions of other viewers and experts.
[1498] An "interview article" refers to a written record of an interview conducted with a film's cast, director, or production staff.
[1499] A "script" refers to a film script or screenplay, a document that contains the actual dialogue and scene details of a film.
[1500] "Subtitles" refers to text that visually displays dialogue or narration in a film.
[1501] "Personalized audio description" refers to audio description that is customized based on a user's individual profile.
[1502] A "speech synthesis engine" is a program for converting text into voice data.
[1503] "Generative AI Model" refers to a model that uses artificial intelligence technology to generate personalized guides based on a user's profile and movie data.
[1504] A "prompt" is an instruction or question input to a generative AI model, and is text used to control the model's output.
[1505] "Smart devices" refers to electronic devices with advanced functions such as smartphones and smart glasses.
[1506] This invention is a system that generates and distributes personalized audio guides that are tailored to the user's circumstances and preferences. The program processing flow of this system is explained below with specific examples.
[1507] 1. Collection of User Information
[1508] When a user first uses the system, the server asks the user to register an account. The user enters basic information such as name, age, language preference, viewing history, movie preferences, etc. The server then creates a user profile and stores it in a database.
[1509] 2. Obtaining movie information
[1510] The server retrieves basic information about the movie selected by the user from online databases and APIs, including the movie title, director, cast, reviews, interviews, script, and subtitle files, and prepares this data for analysis.
[1511] 3. Audio guide generation
[1512] The server analyzes the collected movie data and uses NLP techniques to extract key scenes and emotional nuances. It then generates personalized audio description text based on the user profile. A generative AI model is used to generate the text, and a prompt is input to the model. For example, a prompt might be in the form: "User profile: {Name: User A, Age: 30, Language: Japanese, Favorite movie genre: Mystery} Movie data: {Title: Inception, Director: Christopher Nolan, Cast: {Leonardo DiCaprio: Dom Cobb}, Review: An extremely complex and intricately structured story...} Please generate a personalized movie description."
[1513] After the audio description text is generated, the server passes it to a speech synthesis engine to generate a human-like audio description, adding natural intonation and accents to make it easier to listen to.
[1514] 4. Audio guide distribution
[1515] The generated audio guide is provided to the user's smart device, such as a smartphone or smart glasses, via streaming or download. The server converts the audio guide into a format optimized for the user's device and delivers it appropriately.
[1516] 5. Collecting User Feedback
[1517] Users can provide feedback after using the audio guide. The server collects and analyzes this feedback. The collected feedback is reflected in the user profile and is used to generate audio guides in the future.
[1518] This system makes it possible to provide a personalized movie audio guide tailored to the user's situation and preferences, providing a richer movie-watching experience.
[1519] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1520] Step 1:
[1521] The server requires a user to register an account when using the system for the first time. The user enters basic information such as name, age, language setting, viewing history, and movie preferences. Based on this information, the server creates a user profile and stores it in a database. Input data: name, age, language setting, viewing history, and movie preferences. Output data: user profile.
[1522] Step 2:
[1523] When a user selects a movie to watch, the server retrieves basic information about the movie, including the movie title, director, cast, reviews, interviews, script, and subtitle files. Data is collected from online databases and APIs. Input data: movie title. Output data: movie basic information, reviews, interviews, script, and subtitle files.
[1524] Step 3:
[1525] The server analyzes the collected movie data using NLP technology, extracting key scenes and emotional nuances. Input data: Basic movie information, reviews, interviews, scripts, and subtitle files. Output data: Analyzed movie data.
[1526] Step 4:
[1527] Using the generative AI model, the server generates personalized audio guide text based on the user profile and analyzed movie data. A prompt is input to the generative AI model. For example, a prompt in the following format might be used: "User profile: {Name: User A, Age: 30, Language: Japanese, Favorite movie genre: Mystery} Movie data: {Title: Inception, Director: Generic name, Cast: {Starring actor: Lead role}, Review: Extremely complex and intricately structured story} Please generate a personalized movie audio guide." Input data: User profile, analyzed movie data. Output data: Audio guide text.
[1528] Step 5:
[1529] The server passes the generated audio description text to a speech synthesis engine to generate an audio description that sounds like a human voice, with natural intonation and accent. Input data: Audio description text. Output data: Audio data.
[1530] Step 6:
[1531] The server provides the generated audio guide to the user's smart device, such as a smartphone or smart glasses, in streaming or download format. The audio guide is converted into a format optimized for the user's device. Input data: Audio data. Output data: Audio data in a format suitable for the smart device.
[1532] Step 7:
[1533] Users can provide feedback after using the audio guide. The server collects and analyzes this feedback. The collected feedback is reflected in the user profile and is used to generate audio guides from next time onwards. Input data: User feedback. Output data: Updated user profile.
[1534] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1535] This invention is a system that generates a personalized movie audio guide tailored to the user's situation and preferences, and further combines an emotion engine to recognize the user's emotions and dynamically adjust the content of the audio guide. The program processing flow of this system is explained below with concrete examples.
[1536] Program processing overview
[1537] 1. Collecting User Information
[1538] When a user first uses the system, the server asks the user to register an account. The user enters basic information such as name, age, language preference, viewing history, and movie preferences, and also asks for consent to emotion recognition. Based on this information, the server creates a user profile and stores it in a database.
[1539] Examples:
[1540] User A creates an account and registers in his profile information that he is visually impaired, likes mystery movies, and seeks deep insights. He also agrees to emotion recognition.
[1541] 2. Obtaining movie information
[1542] The user selects a movie they want to watch. The server retrieves basic information about the movie, such as the title, genre, cast, director, and running time, from online databases and APIs. It also collects public reviews, interviews with the cast and director, scripts, and subtitle files.
[1543] 3. Scene Analysis
[1544] The server uses computer vision technology to analyze the film's video data to identify key scenes and character appearances, and NLP technology to analyze the script and subtitle files to extract key dialogue and emotional nuances.
[1545] 4. User Emotion Recognition by Emotion Engine
[1546] The server uses an emotion engine to recognize emotions from the user's facial expressions and voice in real time. While the user is watching a movie, the device uses a camera and microphone to capture the user's facial expressions and voice and transmits the data to the server.
[1547] Examples:
[1548] While User A is watching a movie, the emotion engine recognizes that he is crying at a touching scene. This information is sent to the server.
[1549] 5. Generate and refine personalized guides
[1550] The server generates personalized audio guide text based on the user profile and the recognition results of the emotion engine. Furthermore, it dynamically adjusts the content of the audio guide according to the recognized emotion. For example, if the user is emotional, it adds commentary that reflects that emotion.
[1551] Examples:
[1552] If user A is in tears during a touching scene, the audio guide for that scene will add commentary that delves deeper into the character's emotions.
[1553] 6. Speech Synthesis
[1554] The server then passes the generated audio guide text to a speech synthesis engine, which generates a voice that sounds like a human voice, with natural intonation and accent.
[1555] 7. Audio guide distribution
[1556] The server provides the generated audio guide to the user's terminal in streaming or download format, and the terminal stores the provided audio guide so that the user can use it while watching the movie.
[1557] 8. Gathering Feedback
[1558] After watching a movie, users provide feedback on the quality and content of the audio guide. The server collects the feedback and stores it in a database. The feedback is used to generate future guides.
[1559] 9. Profile Improvements
[1560] The server analyzes the collected feedback and updates the user profile. This updated information is reflected in the next guide generation, providing an audio guide that is more tailored to the user's preferences.
[1561] Overall flow and summary
[1562] This system comprehensively handles everything from user registration to collecting and analyzing movie information, real-time emotion recognition using an emotion engine, generating and adjusting personalized audio guides, delivering them, and collecting feedback. This provides a movie-watching experience tailored to each user's needs and emotions, and addresses the needs of visually impaired people and a diverse audience. The introduction of this system is expected to open up new ways to enjoy movie-going and provide a richer entertainment experience for many people.
[1563] The processing flow will be explained below.
[1564] Step 1: Registering a user
[1565] Users access the system and enter their name, age, language preference, viewing history, favorite movie genre, preferred level of visual and auditory assistance, and consent to emotion recognition on the account creation screen.
[1566] The server receives the information entered by the user and stores it in a database as a new user profile.
[1567] Step 2: Update your profile
[1568] Users can update their profile to reflect changes to their viewing history or new movie preferences by entering and saving changes on the profile page.
[1569] The server updates and saves the changes made by the user to the profile information in the database.
[1570] Step 3: Get movie information
[1571] The user selects the movie they want to watch from within the system.
[1572] The server retrieves basic information about the selected movie, such as title, genre, cast, director, and running time, from online databases or APIs.
[1573] The server also collects reviews of the film, interviews with the cast and director, scripts, and subtitle files.
[1574] Step 4: Analyze the scene
[1575] The server uses computer vision technology to analyze the movie's video data and identify key scenes and scenes in which characters appear.
[1576] The server uses NLP technology to analyze scripts and subtitle files to extract key dialogue and emotional nuances.
[1577] Step 5: Recognizing user emotions with the emotion engine
[1578] The device uses a built-in camera and microphone to capture the user's facial expressions and voice in real time while the movie is playing.
[1579] The terminal sends the captured data to an emotion engine to recognize the user's emotional state.
[1580] The emotion engine (server) analyzes the received data, identifies the user's emotion (such as joy, sadness, surprise, etc.), and generates an emotion status.
[1581] Step 6: Generate and refine your personalized guide
[1582] The server generates personalized audio guide text based on the user profile and the recognition results of the emotion engine.
[1583] The server dynamically adjusts the content of the audio guide depending on the recognized emotion, for example, by adding explanations to calm the user or by emphasizing emotional background information for moving scenes.
[1584] Examples:
[1585] If user A is in tears during a moving scene in a movie, the server can incorporate additional emotional context into the audio description to enhance the viewing experience of that scene.
[1586] Step 7: Text-to-Speech
[1587] The server then passes the generated audio guide text to a speech synthesis engine, which generates a voice that sounds similar to a natural human voice, with natural intonation and accent.
[1588] Step 8: Distributing the audio guide
[1589] The server provides the generated audio guide to the user's terminal in streaming or download format.
[1590] The terminal stores the provided audio description and prepares it for playback as secondary audio during movie playback.
[1591] Step 9: Gather feedback
[1592] After watching a movie, users can enter their feedback on the quality and content of the audio guide on a feedback page within the system.
[1593] The server receives the provided feedback and stores it in a database, which is used to generate subsequent guides.
[1594] Step 10: Refine your profile
[1595] The server analyzes the collected feedback and updates the user profile. This updated information is reflected in the next guide generation, providing an audio guide that is more tailored to the user's preferences.
[1596] Overall flow and summary
[1597] This process comprehensively covers user registration, movie information collection, real-time emotion recognition using an emotion engine, personalized and tailored audio guide generation and delivery, and feedback collection. By dynamically reflecting user emotions, the system can promote a more personalized movie-watching experience and cater to a diverse audience.
[1598] Example 2
[1599] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1600] There is a demand for a variety of viewers, including the visually impaired, to enrich their movie experience and provide personalized guidance that reflects the user's individual emotions and preferences. Conventional systems have difficulty providing guidance that reflects the user's emotions in real time, resulting in only a uniform guidance that makes it difficult to increase user satisfaction.
[1601] The specification process by the specification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for collecting a user's preferences and viewing history and creating a profile; means for acquiring and analyzing basic movie information, reviews, interviews, scripts, and subtitles; means for analyzing visual information using computer vision technology and identifying important scenes and scenes in which characters appear; means for capturing the user's facial expressions and voice through a terminal while watching a movie and recognizing the user's emotions using an emotion engine; means for generating and dynamically adjusting a personalized audio guide based on the user profile and the recognized emotions; means for generating the generated audio guide using a speech synthesis engine; and means for delivering the generated audio guide to the user's terminal. This makes it possible to provide a personalized movie experience tailored to the user's emotions and preferences.
[1602] "User preferences" refer to the tastes and preferences that a user has for a particular genre or content.
[1603] A "viewing history" is a record of movies and programs a user has watched.
[1604] A "profile" is a collection of personal information such as a user's name, age, language preference, viewing history, and movie preferences.
[1605] "Basic movie information" refers to information such as the movie title, genre, cast, director, and screening time.
[1606] A "review" is a piece of writing that expresses an opinion or rating about a movie.
[1607] An "interview article" is a written version of interviews with cast members, directors, and others involved in the film.
[1608] A "script" is a screenplay that contains the dialogue and direction of a movie.
[1609] "Subtitles" are the written versions of the audio in a movie that are displayed on the screen.
[1610] "Computer vision technology" is a technology that allows computers to analyze the content of images and videos.
[1611] An "emotion engine" is a software technology that recognizes emotions by analyzing a user's facial expressions and voice.
[1612] A "speech synthesis engine" is a technology that generates speech by adding natural intonation and accent based on text information.
[1613] A "terminal" is a device that a user uses to watch a movie, including a smartphone, tablet, or PC.
[1614] The present invention provides a system for generating a personalized movie audio guide based on a user's preferences, viewing history, and emotions, thereby improving the user's movie-watching experience. Specific embodiments of the present invention will be described below.
[1615] Hardware and software used
[1616] server:
[1617] Database: Used to store user profiles, movie information, feedback, etc.
[1618] Computer vision technology: Used to analyze film footage and identify key scenes and character appearances.
[1619] Emotion engine: Used to analyze the user's facial expressions and voice to recognize emotions in real time.
[1620] Speech synthesis engine: Used to convert text to speech for personalized audio descriptions.
[1621] Device:
[1622] Camera: Used to capture the user's facial expressions.
[1623] Microphone: Used to capture the user's voice.
[1624] Streaming technology: Used to play audio descriptions provided by the server.
[1625] Explanation of program processing
[1626] Collecting user information
[1627] When a user first uses the system, the server displays an account registration page where the user enters information such as name, age, language preference, viewing history, and preferred genres, and also selects consent to emotion recognition. The server then creates a user profile based on this information and stores it in a database.
[1628] Examples:
[1629] When User A enters information into the registration form, the server receives it, generates a profile and stores it in the database.
[1630] Get movie information
[1631] When a user selects a movie they want to watch, the server retrieves basic information about the movie (title, genre, cast, director, and running time) from an online database, as well as public reviews, interviews with the cast and director, scripts, and subtitle files.
[1632] Examples:
[1633] When User A selects "Mystery Movie X," the server collects basic information about the movie and any associated scripts and subtitle files.
[1634] Scene Analysis
[1635] The server uses computer vision technology to analyze the film's video data to identify key scenes and character appearances, and NLP technology to analyze the script and subtitle files to extract key dialogue and emotional nuances.
[1636] Examples:
[1637] The server breaks down the data for Movie X frame by frame to identify key scenes and characters, while also analyzing the script to extract key lines and moving moments.
[1638] Recognizing user emotions with an emotion engine
[1639] The server uses an emotion engine to recognize emotions in real time from the user's facial expressions and voice. The device uses a camera and microphone to capture the user's facial expressions and voice while watching a movie, and sends the data to the server.
[1640] Examples:
[1641] When user A sheds tears during a touching scene, the device's camera captures this and sends the data to the server.
[1642] Generate and refine personalized guides
[1643] The server generates personalized audio guide text based on the user profile and the recognition results of the emotion engine, and dynamically adjusts the content of the audio guide based on the recognized emotions.
[1644] Examples:
[1645] The server recognizes that user A is moved and adds a specific example to the audio guide for that scene, such as "Character Y's deep emotions are expressed."
[1646] Speech synthesis
[1647] The server passes the generated audio guide text to a speech synthesis engine, which generates natural-sounding speech.
[1648] Examples:
[1649] The server passes the text "Character Y's deep emotions are expressed" to a speech synthesis engine, which generates realistic speech.
[1650] Audio guide distribution
[1651] The server provides the generated audio guide to the user's device in streaming or download format, and the device prepares to play the provided audio guide.
[1652] Examples:
[1653] The server sends the generated audio guide to the terminal, which prepares it for playback.
[1654] Gathering feedback
[1655] After watching a movie, users provide feedback on the quality and content of the audio guide, which the server collects and stores in a database.
[1656] Examples:
[1657] User A provides feedback saying, "The audio guide was easy to understand," and the server receives this and stores it in the database.
[1658] Profile Improvements
[1659] The server analyzes the feedback and updates the user profile, which is then reflected in the next guide generation.
[1660] Examples:
[1661] The server analyzes User A's feedback, adds "I prefer detailed explanations" to his profile, and reflects this in the next audio guide.
[1662] Prompt Sentence Examples
[1663] "Imagine a situation where a user is watching a mystery movie and a touching scene occurs. This makes the user's face begin to shed tears. We would like an explanation of how the emotion engine should recognize this situation and what kind of voice guidance should be generated."
[1664] Examples of prompts:
[1665] "If a user is crying during an emotional scene in a mystery movie, show how an emotion engine can recognize that emotion and add commentary that aligns with the user's emotions."
[1666] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1667] Step 1:
[1668] The server displays an account registration page when a user uses the system for the first time.
[1669] Input: The user enters basic information such as name, age, language preference, viewing history, and movie preferences, and selects consent for emotion recognition.
[1670] Data processing: The server receives this information and creates a user profile.
[1671] Output: The user profile is saved in the database.
[1672] Specific operation: When User A enters information into the registration form and clicks the submit button, the server receives it, generates a profile and saves it in the database.
[1673] Step 2:
[1674] The user selects the movie they want to watch.
[1675] Input: The user selects the movie they want to watch from the movie list.
[1676] Data processing: The server retrieves basic movie information (title, genre, cast, director, running time) from an online database, and collects public reviews, interviews with the cast and director, scripts, and subtitle files.
[1677] Output: The necessary movie information is stored on the server.
[1678] Specific operation: When user A selects the movie "Mystery Movie X," the server collects its basic information and related script and subtitle files.
[1679] Step 3:
[1680] The server analyzes the movie's video data using computer vision technology.
[1681] Input: Collected movie footage data
[1682] Data Processing: Video data is broken down frame by frame to identify key scenes and character appearances, and NLP techniques are used to analyze scripts and subtitle files to extract key dialogue and emotional nuances.
[1683] Output: A dataset containing identified scene and character information, as well as extracted dialogue and emotional nuances.
[1684] What it does: The server breaks down the data of Movie X frame by frame to identify important scenes and characters, while also analyzing the script to extract key lines and moving moments.
[1685] Step 4:
[1686] The device uses a camera and microphone to capture the user's facial expressions and voice while watching a movie, and sends the data to a server.
[1687] Input: facial expressions and voice data of the user while watching a movie
[1688] Data processing: Using an emotion engine, emotions are recognized in real time from the user's facial expressions and voice.
[1689] Output: Data about the user's emotional state.
[1690] Specific operation: When user A sheds tears during a touching scene, the device camera captures this and sends the data to the server.
[1691] Step 5:
[1692] The server generates personalized audio guide text based on the user profile and the recognition results of the emotion engine.
[1693] Input: User profile and real-time emotional data
[1694] Data processing: Generate text that is attuned to the user's preferences and emotions, and dynamically adjust content based on recognized emotions.
[1695] Output: Personalized audio guide text.
[1696] Specific operation: The server recognizes that user A is moved and adds a specific example to the audio description of that scene, such as "Character Y's deep emotions are expressed."
[1697] Step 6:
[1698] The server passes the data to a speech synthesis engine to generate speech from the text.
[1699] Input: The generated audio description text
[1700] Data processing: A speech synthesis engine converts text into speech and adds natural intonation and accent.
[1701] Output: Audio data that sounds similar to a human voice.
[1702] Specific operation: The server passes the text "Character Y's deep emotions are expressed" to the speech synthesis engine, which generates realistic speech.
[1703] Step 7:
[1704] The server provides the generated audio guide to the user's terminal in streaming or download format.
[1705] Input: Generated audio data
[1706] Data processing: The generated audio is converted into an appropriate format and sent to the user's device.
[1707] Output: Audio guide stored on the user's device.
[1708] Specific operation: The server sends the generated audio guide to the terminal, which prepares it for playback.
[1709] Step 8:
[1710] Users provide feedback after watching a movie.
[1711] Input: Feedback information about the quality and content of the audio description
[1712] Data processing: The server stores the collected feedback in a database.
[1713] Output: A database containing the feedback.
[1714] Specific operation: User A provides feedback such as "The audio guide was easy to understand," and the server receives this and stores it in the database.
[1715] Step 9:
[1716] The server analyzes the feedback and updates the user profile.
[1717] Input: Saved Feedback
[1718] Data processing: Analyze feedback and add user preferences and trends to your profile.
[1719] Output: The updated user profile.
[1720] Specific operation: The server analyzes user A's feedback, adds "I prefer detailed explanations" to his profile, and reflects this in the next audio guide.
[1721] (Application example 2)
[1722] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1723] Conventional movie viewing systems have limited means to accommodate visually impaired users and users who want to understand specific scenes or information in depth. In particular, the lack of a function to provide personalized audio guides that are sensitive to the user's emotions results in a poor movie-viewing experience. Furthermore, there is a lack of a mechanism to learn from user feedback and improve the guide content for future viewings. This makes it difficult to improve user satisfaction.
[1724] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1725] In this invention, the server includes means for collecting user preferences and viewing history and creating a profile, means for acquiring and analyzing basic movie information, ratings, related party information, dialogue, and commentary, means for generating a personalized audio guide based on the user profile, means for recognizing the user's emotions and dynamically adjusting the content of the audio guide, means for generating the generated audio guide using a speech synthesis engine, means for delivering the generated audio guide to the user's device, means for analyzing the collected feedback and updating the user profile, means for analyzing visual information using image analysis technology and identifying important scenes and scenes with characters, and means for generating prompt sentences using a generative AI model. This allows visually impaired people and emotionally sensitive users to better understand and enjoy movies.
[1726] "User" refers to a person who uses the system to watch a movie.
[1727] "Preferences" refer to personal preferences that a user has for particular movie genres or content.
[1728] "Viewing history" refers to historical information about movies and video content that a user has viewed in the past.
[1729] A "profile" refers to a data set that compiles information such as a user's preferences and viewing history.
[1730] "Basic movie information" refers to basic data about a movie, such as the title, genre, cast, director, and running time.
[1731] "Rating" refers to the ratings and reviews of movies by users and experts.
[1732] "Related Person Information" refers to detailed information about people such as the film's cast and director.
[1733] "Dialogue" refers to the words spoken by characters in a film.
[1734] "Explanatory text" refers to text that explains the content and background of the film.
[1735] "Emotion" refers to the psychological state a user experiences while watching a movie.
[1736] "Audio guide" refers to audio guidance or explanations provided to the user.
[1737] "Dynamic adjustment" refers to changing content in response to real-time conditions.
[1738] "Speech synthesis engine" refers to software or hardware for converting text data into speech data.
[1739] "Terminal" refers to the electronic device that a user uses to watch movies and use the system.
[1740] "Feedback" refers to opinions and evaluations from users.
[1741] "Image analysis technology" refers to technology for processing visual information and extracting specific patterns and features.
[1742] A "scene" refers to a specific situation or situation within a film.
[1743] A "character scene" refers to a specific moment in a film in which a character appears.
[1744] "Generative AI models" refer to artificial intelligence techniques for performing tasks such as text generation.
[1745] A "prompt sentence" refers to an input sentence that causes a generative AI model to generate text.
[1746] The present invention is a system for improving a user's viewing experience, which includes the following means: A server collects a user's preferences and viewing history and creates a profile, which includes information provided by the user such as name, age, language, viewing history, and movie preferences. The profile is stored in a database.
[1747] The server then retrieves and analyzes information such as the movie title, genre, cast, director, running time, ratings, people involved, dialogue, and description from online databases and APIs, and is then ready to provide detailed movie information to the user.
[1748] Based on the user profile, the server generates a personalized audio guide. To do so, it uses NLP technology to analyze the movie script and subtitles and extract important lines and scenes. It also recognizes the user's emotions in real time and dynamically adjusts the content of the audio guide based on that information. Emotion recognition is performed using, for example, Microsoft Azure Cognitive Services and Google Cloud's Natural Language API. This allows the server to provide commentary that matches the emotions the user feels while watching the movie.
[1749] The generated audio description is converted into speech with natural intonation and accent by a speech synthesis engine (e.g., Amazon Polly or Google Text-to-Speech), and is delivered to the user's device in streaming or download format.
[1750] Users provide feedback after watching a movie, which is collected and analyzed by the server, and their profile is updated to reflect this in the generation of personalized audio guides for future viewings.
[1751] Computer vision techniques (e.g., OpenCV) are used to analyze visual information and identify important scenes and character appearances in the film, which is also incorporated into the personalized audio description.
[1752] The generative AI model is used to generate prompts. An example of a prompt is, "Based on user profile: {'name': 'Mr. A', 'age': 30, 'language': 'jp', 'preferences': ['mystery']}, movie data: {detailed movie information}, emotion: {detailed emotion information}, please generate an in-depth explanation of the key scenes in this movie."
[1753] This allows visually impaired people and emotionally sensitive users to understand and enjoy movies more deeply. This system is ideal for content distribution services and is expected to significantly improve the movie-watching experience.
[1754] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1755] Step 1:
[1756] The server collects the user's preferences and viewing history and creates a profile. The input is information provided by the user about their name, age, language, viewing history, and movie preferences. The information is stored in a database and a user profile is generated. Specifically, the user enters information into an account creation form, which is received by the server.
[1757] Step 2:
[1758] The server retrieves basic movie information (title, genre, cast, director, running time), ratings, related party information, dialogue, and description from an online database or API. The input is the movie title and ID, and the output is the retrieved movie details. The server sends the API request and stores the retrieved data in a database for analysis.
[1759] Step 3:
[1760] The server uses computer vision technology to analyze the visual information of a movie and identify important scenes and scenes featuring characters. The input is the movie's video data, and the output is metadata for the identified scenes. Specifically, OpenCV is used to analyze the video data frame by frame and extract characteristic scenes.
[1761] Step 4:
[1762] The server uses NLP technology to analyze movie scripts and subtitles and extract important lines and scenes. The input is the text data of the script or subtitles, and the output is a list of important lines and scenes. Specifically, it uses NLP libraries such as spaCy to perform text analysis and extract important lines.
[1763] Step 5:
[1764] The device uses a camera and microphone to capture the user's facial expressions and voice and sends the data to a server. An emotion recognition engine analyzes this data. The input is the user's video and audio data, and the output is recognized emotional information. Specifically, the device captures the user's facial expressions and voice at regular intervals and sends them to the server in real time.
[1765] Step 6:
[1766] The server generates a personalized audio guide based on the user profile and real-time emotional data. A prompt is input into the generative AI model to generate the audio guide text. The input is profile information, emotional data, and the prompt, and the output is the audio guide text. An example of a specific prompt is, "Based on user profile: {'name': 'Mr. A', 'age': 30, 'language': 'jp', 'preferences': ['Mystery']}, movie data: {detailed movie information}, emotion: {detailed emotion information}, please generate an in-depth commentary on the important scenes in this movie."
[1767] Step 7:
[1768] The server passes the personalized audio guide text to a speech synthesis engine to generate a voice that sounds similar to a human voice. The input is the audio guide text data, and the output is the audio data. Specifically, the server converts the text to speech using a speech synthesis engine such as Google Text-to-Speech (gTTS).
[1769] Step 8:
[1770] The server distributes the generated audio guide to the user's device. The input is the generated audio data, and the output is the audio guide stored on the user's device. Specifically, the audio guide is provided in streaming format or file download format.
[1771] Step 9:
[1772] After watching a movie, users provide feedback on the quality and content of the audio description. The server collects this feedback and stores it in a database. The input is the user's feedback data, and the output is an updated user profile. For example, opinions can be collected through a web form or an in-app feedback feature.
[1773] Step 10:
[1774] The server improves the user profile based on the feedback and analysis data, and reflects this in future personalized audio guides. The input is the collected feedback and viewing data, and the output is an updated user profile. Specifically, the server analyzes the feedback data and adds new user preference information to the profile.
[1775] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1776] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1777] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1778] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1779] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1780] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1781] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1782] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1783] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1784] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1785] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1786] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1787] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1788] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1789] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1790] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1791] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1792] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1793] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1794] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1795] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1796] The following is further disclosed regarding the above embodiment.
[1797] (Claim 1)
[1798] means for collecting user preferences and viewing history and creating a profile;
[1799] A means to obtain and analyze basic movie information, reviews, interviews, scripts and subtitles,
[1800] means for generating a personalized audio guide based on a user profile;
[1801] means for generating the generated audio guide by a speech synthesis engine;
[1802] A system including means for distributing the generated audio guide to a user's terminal.
[1803] (Claim 2)
[1804] 10. The system of claim 1, further comprising means for analyzing the collected feedback and updating the user profile.
[1805] (Claim 3)
[1806] 10. The system of claim 1, further comprising means for analyzing the visual information using computer vision techniques to identify important scenes and scenes in which characters appear.
[1807] "Example 1"
[1808] (Claim 1)
[1809] means for collecting user preferences and viewing history and creating a profile;
[1810] A means to obtain and analyze basic movie information, reviews, interviews, scripts and subtitles,
[1811] means for generating a personalized audio guide based on a user profile;
[1812] A means of analyzing the collected movie data and extracting key lines and emotional nuances using scene analysis and natural language processing techniques;
[1813] means for generating the generated audio guide by a speech synthesis engine;
[1814] means for delivering the generated audio guide to a user's terminal;
[1815] A system that includes a means of collecting user feedback and incorporating it into profiles.
[1816] (Claim 2)
[1817] 10. The system of claim 1, further comprising means for analyzing the collected feedback and updating the user profile.
[1818] (Claim 3)
[1819] 10. The system of claim 1, further comprising means for analyzing the visual information using computer vision techniques to identify important scenes and scenes in which characters appear.
[1820] "Application Example 1"
[1821] (Claim 1)
[1822] means for collecting user preferences and viewing history and creating a profile;
[1823] A means to obtain and analyze basic movie information, reviews, interviews, scripts and subtitles,
[1824] means for generating a personalized audio guide based on a user profile;
[1825] means for generating the generated audio guide by a speech synthesis engine;
[1826] means for delivering the generated audio guide to a user's terminal;
[1827] a means for generating personalized audio description text based on a user profile and movie data using a generative AI model;
[1828] A means for invoking a generative AI model using a prompt sentence in generating an audio guide;
[1829] A means for converting the generated audio guide into a format suitable for smart devices and distributing it;
[1830] A system including:
[1831] (Claim 2)
[1832] 10. The system of claim 1, further comprising means for analyzing the collected feedback and updating the user profile.
[1833] (Claim 3)
[1834] 10. The system of claim 1, further comprising means for analyzing the visual information using computer vision techniques to identify important scenes and scenes in which characters appear.
[1835] "Example 2: Combining Emotion Engines"
[1836] (Claim 1)
[1837] means for collecting user preferences and viewing history and creating a profile;
[1838] A means to obtain and analyze basic movie information, reviews, interviews, scripts and subtitles,
[1839] A means of analyzing visual information using computer vision technology to identify important scenes and scenes in which characters appear;
[1840] A means for capturing a user's facial expressions and voice through a terminal while watching a movie and recognizing the user's emotions using an emotion engine;
[1841] means for generating and dynamically adjusting personalized audio guidance based on a user profile and recognized emotions;
[1842] means for generating the generated audio guide by a speech synthesis engine;
[1843] A means for delivering the generated audio guide to the user's device
[1844] A system including:
[1845] (Claim 2)
[1846] 10. The system of claim 1, further comprising means for analyzing the collected feedback and updating the user profile.
[1847] (Claim 3)
[1848] A method of using computer vision and natural language processing technologies to analyze movie video data, scripts, and subtitles to extract important dialogue and emotional nuances.
[1849] 10. The system of claim 1, further comprising:
[1850] "Application example 2 when combining emotion engines"
[1851] (Claim 1)
[1852] means for collecting user preferences and viewing history and creating a profile;
[1853] A means to acquire and analyze basic information about movies, ratings, information about people involved, lines, and commentary,
[1854] means for generating a personalized audio guide based on a user profile;
[1855] means for recognizing a user's emotion and dynamically adjusting the content of the audio guide;
[1856] means for generating the generated audio guide by a speech synthesis engine;
[1857] means for delivering the generated audio guide to a user's terminal;
[1858] means for analyzing the collected feedback and updating the user profile;
[1859] A means of analyzing visual information using image analysis technology to identify important scenes and scenes with characters,
[1860] A system including means for generating a prompt sentence using a generative AI model.
[1861] (Claim 2)
[1862] 10. The system of claim 1, wherein the collected feedback is analyzed and the user profile is updated.
[1863] (Claim 3)
[1864] The system according to claim 1, wherein the visual information is analyzed using analytical techniques to identify important scenes and scenes with characters. [Explanation of symbols]
[1865] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. means for collecting user preferences and viewing history and creating a profile; A means to obtain and analyze basic movie information, reviews, interviews, scripts and subtitles, means for generating a personalized audio guide based on a user profile; means for generating the generated audio guide by a speech synthesis engine; A system including means for distributing the generated audio guide to a user's terminal.
2. The system of claim 1 , further comprising means for analyzing the collected feedback and updating the user profile.
3. 10. The system according to claim 1, further comprising means for analyzing the visual information using computer vision techniques to identify important scenes and scenes in which characters appear.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A