system

By combining user-selectable application programming interfaces (APIs) and generative artificial intelligence with music distribution services, landscapes and virtual characters synchronized with music are generated in real time. This solves the problems of insufficient visual experience and high cost in window-type display services, and provides a personalized and interactive immersive experience.

JP2026085736APending Publication Date: 2026-05-25SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-11-13
Publication Date
2026-05-25

AI Technical Summary

Technical Problem

Existing window-type display services suffer from reduced visual experience due to building obstruction and repetitive video, leading to user boredom, while also incurring high production costs.

Method used

By using a user-selectable application programming interface (API), combined with music distribution services and generative artificial intelligence, the system generates real-time landscapes and virtual characters synchronized with the music, providing an interactive experience and using user feedback to optimize the system.

Benefits of technology

It achieves a low-cost immersive experience, enhances visual and auditory interactivity, allows users to customize content according to their personal preferences, and the system continuously improves based on feedback.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026085736000001_ABST
    Figure 2026085736000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] A means of providing an interface for users to select an experience category, A means of receiving and analyzing music data, A method that utilizes a generative AI to generate landscapes based on analysis results, A means of generating and interacting with characters synchronized with music, A system that includes means for presenting generated video and music to the user in sync.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0004] , , , ,

[0005] , , , , ,

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, and includes steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance as a response to the user utterance.

Prior Art Documents

Patent Documents

[0003] <统一码转义序列: / / 这里推测原文有误,按照原样保留

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] There is a problem that the visual enjoyment is reduced because the view from the window in the urban area is restricted by adjacent buildings. In addition, in the existing window-type display service, the user is likely to get bored because the same video is repeatedly displayed, and there is also a problem that the production cost soars. In response to these problems, it is required to provide an immersive experience that utilizes realistic or fantasy landscapes and synergies with music at a low cost.

Means for Solving the Problems

[0005] This invention provides an interface that allows users to select a category they wish to experience through an application and receive music data from a music distribution service. The server analyzes the music data and generates a landscape in real time using a generative AI based on the analysis results. Furthermore, it generates a character based on information obtained from the music and enables interaction with this character. As a result, the generated video and music are synchronized and presented to the user as a new experience. In addition, user feedback is saved as learning data, and further customization options are provided to improve user immersion and satisfaction.

[0006] A "user" is a person who uses this system and is the entity that selects experience categories and customizes settings.

[0007] An "experience category" is an option that users can choose from, a classification that provides experiences based on specific scenery or themes.

[0008] An "interface" refers to the screens or tools that users use to access and operate a system.

[0009] "Music data" refers to digital data, including audio information, obtained from music streaming services.

[0010] "Analysis" is a series of processes in which a system analyzes music data and identifies its characteristics.

[0011] "Generative AI" is an artificial intelligence technology that uses analysis results to create new landscapes and images.

[0012] "Landscape" refers to the visual scenes and backgrounds created by generative AI and presented to the user.

[0013] A "character" is a virtual person or entity created for the purpose of interacting with the user.

[0014] "Dialogue" refers to the communication by the character with the user through language and actions.

[0015] "Feedback" refers to the reactions and opinions provided by the user regarding the experience, and is used for system improvement and adaptation.

[0016] "Customization options" are the choices that allow the user to adjust the system's operation and display content according to their preferences.

Brief Description of the Drawings

[0017] [Figure 1] It is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] It is a conceptual diagram showing an example of the main functions of a data processing device and a smart device according to the first embodiment. [Figure 3] It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which multiple emotions are mapped. [Figure 10] It shows an emotion map to which multiple emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Embodiment 2 when the emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when the emotion engine is combined.

Mode for Carrying Out the Invention

[0018] Hereinafter, an example of an embodiment of the system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0019] First, the terms used in the following description will be explained.

[0020] In the following embodiments, the numbered processor (hereinafter simply referred to as "processor") may be one arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be one type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), APU (Accelerated Processing Unit), etc.

[0021] In the following embodiments, the numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0022] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.

[0023] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0024] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0025] [First Embodiment]

[0026] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0027] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0028] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0029] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0030] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0031] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0032] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0033] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0034] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0035] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0036] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0037] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0038] This invention provides a system that allows users to enjoy the synergy of real-time generated scenery and music. This system mainly consists of a user's terminal, a server, and a music distribution service.

[0039] First, the user launches the application provided on their device and has the function to select the category of scenery they want to experience. The information selected by the user is then transferred to the server via the internet.

[0040] The server prepares to generate the appropriate landscape based on the selected category. It also begins streaming music data accessed by the device from a music streaming service. The music data is sent to the server for analysis of the music's rhythm, tempo, and other characteristics.

[0041] The server uses the results of this music analysis to generate appropriate landscapes in real time via a generating AI. The generated landscapes are linked to the characteristics of the music; for example, calm music will automatically create a tranquil natural landscape, while fast-paced music will create a breathtaking city nightscape.

[0042] Furthermore, the server generates characters that synchronize with the music and interact with the user. The characters speak different lines depending on the content and tone of the music, and can even engage in humorous conversations. This provides a more immersive experience.

[0043] The generated video and music are streamed to the device, which then displays them together on the screen. Users can enjoy the real-time changing scenery and music, creating a multi-faceted experience that engages both sight and sound.

[0044] The system's ability to collect user feedback, which is then analyzed and recorded by the server, is also crucial. This feedback information is used to improve the generating AI. Furthermore, users have the ability to customize their experience, adjusting settings such as scenery and dialogue according to their preferences. This customization information is also sent to the server and used to improve future experiences.

[0045] As a concrete example, consider a scenario where the user selects two different categories: "natural scenery" and "fantasy world." In the natural scenery category, gentle folk music plays, and a tranquil sunrise scene over a lake is displayed. The character is in the shape of a bird and provides simple trivia about nature. In contrast, in the fantasy world category, a colorful magical world is displayed along with dance music, and the character dances while greeting the user in a friendly manner. Thus, the present invention is a system that dynamically provides diverse experiences based on the user's choices.

[0046] The following describes the processing flow.

[0047] Step 1:

[0048] The user launches the application on their device and selects the category they want to experience from the displayed interface. This selection information is sent from the device to the server.

[0049] Step 2:

[0050] The device connects to a music streaming service via the internet and receives music data in real time based on the user's preferences. The received music data is then relayed to a server.

[0051] Step 3:

[0052] The server analyzes the received music data. This analysis includes extracting musical elements such as rhythm, tempo, and melody characteristics. The analysis results are used as input data for the generating AI.

[0053] Step 4:

[0054] Based on the analysis of the music, the server uses a generative AI to generate a landscape in real time that fits the selected category. The landscape changes in accordance with the pace and tone of the music.

[0055] Step 5:

[0056] The server generates characters within the same context as the generated landscape. The generating AI determines the characters' actions and dialogue based on the music's theme and mood. The characters' dialogue scenarios are also constructed at this stage.

[0057] Step 6:

[0058] The server streams generated scenery and character data to the device, which then displays it on the screen. The synchronization of music and video timing provides a visually and aurally consistent experience.

[0059] Step 7:

[0060] Users provide feedback through the interface presented during their experience. This feedback is sent from the device to the server and used to improve the system and update user profiles.

[0061] Step 8:

[0062] After the experience ends, users can use customization options to adjust the scenery, character movements, dialogue, and other elements. This customization information is sent from the device to the server and reflected in the next experience.

[0063] (Example 1)

[0064] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0065] In modern society, there is a demand for more personalized and immersive music listening experiences. However, existing systems struggle to provide dynamic visual experiences synchronized with music, and they lack options for users to finely customize that experience. Furthermore, there is a challenge in the insufficient mechanism for incorporating user feedback and continuously improving the system.

[0066] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0067] In this invention, the server includes means for providing a user screen for the user to select the type of experience, means for receiving and analyzing audio data, means for utilizing an artificial intelligence model that generates a landscape based on the analysis results, and means for generating a virtual character that is synchronized with the audio and interacting with it. This provides the user with an immersive experience that links music and visuals, further allows the user to customize the experience, and enables continuous improvement of the system through feedback.

[0068] "User interface for selecting the type of experience" refers to the interface that allows users to select scenery and experience content that are linked to their desired music.

[0069] "Receiving and analyzing audio data" refers to the process of analyzing the characteristics of music data, such as rhythm and tempo.

[0070] An "artificial intelligence model for generating landscapes" refers to an algorithm that generates appropriate visuals in real time based on the results of analyzing music data.

[0071] A "virtual character linked to sound" refers to a digital character that is generated according to the characteristics and content of the music and is capable of interacting with the user.

[0072] "Engaging in interaction" refers to virtual characters providing users with music-related information or offering interactive experiences through dialogue.

[0073] "Synchronizing and providing generated video and audio to users" refers to seamlessly presenting real-time generated visuals and music to the user.

[0074] "Providing choices" means offering within the system options that allow users to customize their experience and adjust the scenery and dialogue content to their own preferences.

[0075] This invention is a system that allows users to enjoy a real-time fusion of music and scenery. The system consists of a user's terminal, a server, and a music distribution service.

[0076] First, the user launches a dedicated application on their device and selects the category of scenery they wish to experience. A user interface is provided for selection, allowing the user to easily determine their preferred type. This selection information is then transmitted to a server via the internet.

[0077] The server receives the selected category information and prepares to access the music streaming service. Music data is streamed through the terminal, and the server receives and analyzes the music data. The analysis utilizes audio analysis software to extract the rhythm, tempo, and other features of the music.

[0078] The server also uses a generative AI model to generate appropriate landscapes in real time based on the analysis results. The generated landscapes are linked to the characteristics of the music; for example, calm music will create a tranquil natural landscape, while dynamic music will create an impressive urban landscape in real time.

[0079] Furthermore, the server generates a virtual character synchronized with the audio, enabling two-way interaction with the user. This character provides a deeper sense of immersion by using dialogue based on the content and tone of the music. Through interaction with the character, users can enjoy an experience where music and scenery are integrated.

[0080] The video and music are synchronized on the device, and the generated content is delivered seamlessly to the user. Users can also use customization options to adjust the scenery and character interactions to their liking.

[0081] For example, when a user selects the "Natural Landscape" category, the device receives calming music from a music streaming service. The server generates a tranquil lake landscape corresponding to this music, displays a bird character, and provides fun facts about nature. In this way, the user can enjoy a unique audiovisual experience based on the prompt message, "Please select a tranquil landscape and set it so that a character speaks simple nature facts accompanied by music that evokes a morning atmosphere."

[0082] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0083] Step 1:

[0084] The user launches the application through their device and selects the category they wish to experience from the displayed user screen. This input information is sent to the server via the internet. The selection information is stored in a database and used in the server's next processing step. Specifically, the user selects the "Natural Landscapes" category, and this selection information is sent to the server.

[0085] Step 2:

[0086] The server prepares to communicate with the music streaming service based on the selected category information. The server sends a request to the terminal and starts streaming music data. The streamed music data is received by the server, and analysis of the music's rhythm, tempo, tone, etc., is performed. The input data is a music file, and the output is the analysis results showing the characteristics of the music. For example, data on a gentle rhythm and tone may be extracted.

[0087] Step 3:

[0088] The server uses the analysis results and leverages a generative AI model to generate landscapes that match the music in real time. In this process, the algorithm selects appropriate visual elements based on the characteristics of the analyzed music, dynamically constructing the landscape. The input is the characteristics of the music, and the output is a visually generated content in real time. For example, for calm music, landscapes of lakes and forests are generated.

[0089] Step 4:

[0090] The server then generates virtual characters based on musical characteristics and assigns them scripts. These characters initiate interaction with the user using dialogue synchronized with the music. Input is the music and user selections, while output is dialogue and actions for interaction. Specifically, the server generates a bird character and provides topics related to nature.

[0091] Step 5:

[0092] The terminal receives generated scenery and character information transmitted from the server and integrates and displays it on the screen. At this time, the video and music are synchronized, providing the user with a seamless experience. The input is the generated scenery and character data, and the output is integrated visual and auditory content for the user. For example, scenery and characters synchronized with gentle music are displayed on the screen.

[0093] Step 6:

[0094] User feedback is collected and sent to the server. The server stores this feedback in a database and uses it to improve future AI models. The input is user feedback data, and the output is training data aimed at improving the system. For example, based on the feedback, the AI ​​model is adjusted so that a more preferred landscape is generated in the next session.

[0095] (Application Example 1)

[0096] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0097] When users enjoy the synergy of music and visuals in real time, it is essential that the generation of scenery and characters is smoothly adapted to the characteristics of the music, thereby enhancing the quality of the viewing experience. Furthermore, incorporating user customization and feedback is a challenge in providing an experience that better suits individual needs.

[0098] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0099] In this invention, the server includes means for providing an operating unit for the user to select the type of experience, means for receiving and analyzing acoustic signals, and means for utilizing a generation AI to generate scenes based on the analysis results. This enables the user to experience immersive viewing experiences tailored to their individual needs by generating scenery and characters in harmony with music in real time.

[0100] A "user" refers to someone who uses and operates a system or application.

[0101] "Experience types" refer to the various themes and situations that users can choose from.

[0102] The term "operation unit" refers to the interface through which the user makes selections and gives instructions to the system.

[0103] "Acoustic signal" refers to sound signals that the system analyzes, such as music data and voice data.

[0104] "Analysis" refers to the process of analyzing the characteristics of acoustic signals and extracting guidelines for generating landscapes and characters.

[0105] "Scene" refers to the collective term for the images and visuals presented to the user.

[0106] "Generative AI" refers to artificial intelligence that uses machine learning techniques to automatically generate landscapes, characters, and other images.

[0107] A "moving character" refers to a character that appears on the screen in conjunction with an audio signal and interacts with the user through dialogue and actions.

[0108] The system implementing this invention uses a user terminal, a server, and a generative AI with music analysis capabilities. Processing begins when the user selects the type of experience they wish to experience via the control panel on the terminal. The terminal transmits this selection information to the server. The server receives audio signals via streaming from a music distribution service over the internet and analyzes those signals.

[0109] Based on the analysis results, the server uses a generative AI to generate corresponding scenes in real time. This generative AI has the ability to analyze characteristic data such as rhythm, tempo, and melody of the acoustic signal and design a scene that is appropriate for it. Furthermore, an action character is also generated according to the attributes of the acoustic signal and interacts with the user.

[0110] The generated scenes and sounds are streamed to the device in real time. The device integrates and displays this, allowing the user to enjoy a harmonious visual and auditory experience.

[0111] As a concrete example, if the user selects relaxing music, the AI ​​will present a tranquil forest landscape in sync with that music. Within that forest, a moving object in the shape of a small bird will appear, creating a scene that provides the user with interesting facts about nature.

[0112] As an example, the prompt given to the generation AI is, "Generate character movements that harmonize with a tranquil nighttime cityscape, set to jazz music." In this way, interaction between the AI ​​and the user experience is realized.

[0113] This system can continuously improve the user experience through user customization and feedback. Specifically, user customizations and feedback information are recorded on the server and used as training data for the generating AI. This enables the provision of a more personalized experience.

[0114] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0115] Step 1:

[0116] The user selects the type of experience they wish to have using the controls on their device. This input determines the theme of the experience and is sent to the server as data. At this stage, the user's selection information is collected on the server.

[0117] Step 2:

[0118] The server receives audio signals related to a specified experience via streaming from a music streaming service. This audio data is used as input data to extract the characteristics of the music. The server analyzes the audio signals and performs data processing to identify rhythm, tempo, pitch, and other characteristics.

[0119] Step 3:

[0120] The server utilizes a generative AI model based on the analysis results and generates a suitable scene using prompt text. This AI takes the analyzed musical characteristics as input data and generates a visual scene that matches the prompt as output. For example, a prompt text for calm music might be "Generate a calm forest landscape."

[0121] Step 4:

[0122] The server simultaneously generates characters based on the tone and content of the sound. A generation AI model is used to depict the characters' actions and dialogue. In this process, the volume and rhythm of the sound are taken as input, and the character's movements and speech are output.

[0123] Step 5:

[0124] The generated scenes and audio are streamed in real time from the server to the terminal, which then integrates and displays them to the user. The terminal receives the data stream from the server and outputs synchronized audio and video to provide the user with an experience.

[0125] Step 6:

[0126] After the experience ends, the user provides feedback, which the device sends to the server. This feedback information is recorded on the server as input and used as training data for the generated AI model. The server uses this data to perform calculations to improve the quality of future experiences.

[0127] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0128] This invention provides a system that allows users to enjoy a synergy based on real-time generated scenery and music, and in particular, by combining it with an emotion engine that recognizes the user's emotions, it offers a more personalized experience. This system consists of a user's terminal, a server, a music distribution service, and an emotion engine.

[0129] First, the user launches an application on their device and selects a category of scenery they want to experience. This information is sent from the device to the server, and the device's emotion engine analyzes the user's facial expressions and voice captured through the camera and microphone to recognize the user's current emotional state.

[0130] The server receives this selection and emotional information and retrieves music data from a music streaming service. The music data is transmitted in real time, and the server analyzes the rhythm, tempo, and other musical characteristics. The analyzed information is used by a generative AI to create a landscape in real time that is appropriate to the user's emotional state.

[0131] The generating AI adjusts the color tone, composition, and dynamic elements of the landscape based on the output of the emotion engine. For example, it can generate a calm landscape when the user is relaxed and a lively landscape when they are excited. Characters also evolve to display appropriate expressions and actions according to emotions, supporting natural conversations.

[0132] The generated video and music are streamed to the device. The device seamlessly integrates the video and music, presenting it as a new, optimized experience for the user. The user can enjoy this immersive experience and simultaneously provide feedback through the interface. This feedback is sent from the device to the server and used as data to improve future experiences.

[0133] As a concrete example, consider a scenario where a user selects "natural scenery" with the goal of "relaxing," and the emotion engine recognizes the user's calm state. In this case, the server generates scenery of a quiet forest or a calm beach and selects music with a soothing melody. The character speaks to the user in a friendly and gentle tone. In this way, the present invention flexibly provides experiences based on the user's emotions and choices, making it possible to enrich daily life in urban areas.

[0134] The following describes the processing flow.

[0135] Step 1:

[0136] The user launches the application on their device and selects the category they want to experience from the interface. This selection information is sent from the device to the server.

[0137] Step 2:

[0138] The device uses its built-in camera and microphone to capture the user's facial expressions and voice. An emotion engine analyzes this data to identify the user's emotional state. The identified emotion information is then sent to the server.

[0139] Step 3:

[0140] The device accesses a music streaming service and receives music data tailored to the user's preferences in real time. The received music data is then transferred to a server.

[0141] Step 4:

[0142] The server analyzes the music data and extracts features such as tempo, rhythm, and genre. These analysis results are then used as a set of data necessary for the landscape generation process.

[0143] Step 5:

[0144] The server combines the user's emotional state, recognized by the emotion engine, with the music analysis results, and inputs this data into the generative AI. The generative AI then generates an appropriate scene in real time.

[0145] Step 6:

[0146] The generating AI adjusts the color scheme and composition of the landscape to match the user's emotions, creating a landscape with dynamic elements. Simultaneously, the server generates a character synchronized with the music, ready to interact with the user.

[0147] Step 7:

[0148] The server streams landscape data and character dialogue scripts generated by the server to the terminal. The terminal displays these on its screen, ensuring that music and video are seamlessly integrated.

[0149] Step 8:

[0150] Users utilize a feature to provide feedback during the experience. The device sends the entered feedback to a server for recording. This feedback is used to improve future experiences.

[0151] Step 9:

[0152] After the user finishes the experience, they can adjust the scenery and dialogue using customization options. The device sends the customization information to the server, and it is reflected in the next experience.

[0153] (Example 2)

[0154] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0155] In modern society, customized entertainment experiences tailored to the emotional state of users are limited. In particular, there is a lack of means to create immersive experiences where music and visual content are closely integrated. Furthermore, there is a need for mechanisms that incorporate user feedback into the system to improve individual experiences.

[0156] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0157] In this invention, the server includes means for providing a display device for the user to select the type of experience, means for acquiring and analyzing sound information, and means for utilizing a generative model that generates a landscape based on the analysis results and the user's emotional information. This makes it possible for music and visual content to be synchronized according to the user's emotions, providing a personalized and immersive experience.

[0158] "User" refers to an individual or group that operates the system and selects the type of experience.

[0159] "Experience type" refers to the category or theme of the content that the user wants to experience.

[0160] A "display device" refers to a device or software that provides an interface used by users to select the type of experience they are experiencing.

[0161] "Audio information" refers to data that a system processes, including music and other sound data.

[0162] "Analysis" refers to the process of processing and calculating acquired sound and emotional information.

[0163] "Emotional information" refers to the emotional state estimated from facial expressions and voices acquired through the user's camera and microphone.

[0164] A "generative model" refers to a system that uses artificial intelligence to generate new landscapes and content based on analysis results and emotional information.

[0165] A "virtual character" is a generated character that interacts with sound information and enables dialogue with the user.

[0166] "Visual information" refers to generated video data such as landscapes, which is presented to the user in conjunction with audio information.

[0167] This invention is a system that enables users to enjoy a customized entertainment experience tailored to their emotions in real time. This system mainly consists of three main elements: a terminal, a server, and an audio service.

[0168] First, users can select their preferred content category via a display on their device that allows them to choose the type of experience. This selection information is then sent from the device to the server. The device uses its built-in emotion engine to analyze emotional information from the user's facial expressions and voice captured through the camera and microphone. This information indicates the user's current emotions and is used to customize the content provided by the system.

[0169] The server uses the selection and emotion information received from the user to retrieve sound information from the audio service. The retrieved sound information is then analyzed on the server, and features such as rhythm and tempo are identified. Based on these features and emotion information, the server generates prompts and inputs them into a generative AI model. This model generates visual information such as landscapes and characters based on the specified information. For example, if the user selects "I want to relax," the generative AI model will depict calm natural scenery and friendly characters.

[0170] The generated visual information and acquired audio information are transmitted from the server to the terminal in real time. The terminal integrates these and presents them seamlessly to the user, providing an immersive experience. For example, if the user selects "I want to relax" and the system recognizes that the user is in a calm state based on emotional information, calming music and visual natural scenery appropriate to that situation will be presented.

[0171] An example of a prompt message would be: "The user has selected the category 'Natural Landscapes' and their emotional state is 'Calm.' Based on this, please generate a relaxing landscape and corresponding music." In this way, the system can provide an entertainment experience that matches the user's individual emotions.

[0172] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0173] Step 1:

[0174] The user launches the application on their device and selects the type of content they want to experience. This input information is saved on the device and then transferred to a server via the network. Users can choose categories such as "natural landscapes" or "urban scenery."

[0175] Step 2:

[0176] The device operates a built-in emotion engine and uses the user's camera and microphone to collect facial expressions and voice. This collected data is then analyzed by an emotion analysis algorithm to identify the user's emotional state (e.g., "relaxed," "excited," etc.). The analysis output is sent to the server as the user's emotion information.

[0177] Step 3:

[0178] The server receives selection and emotion information from the user. Based on this information, the server connects to an audio service and obtains appropriate sound information. Here, the server requests the optimal music data according to the user's selection and emotion, and the server analyzes the characteristics of the obtained music, such as rhythm, tempo, and melody. The output of this analysis becomes part of the generation prompt.

[0179] Step 4:

[0180] The server generates prompts for visual information based on the analyzed sound and emotion information. These prompts include information about the sound rhythm, the user's emotional state, and the landscape according to the selected category. For example, a specific prompt statement such as "The user's selected category is 'Natural Landscape', and their emotional state is 'Calm'" is passed to the generating AI model.

[0181] Step 5:

[0182] The generation AI model begins generating landscapes and characters based on prompts received from the server. The model adjusts color tones and behavioral elements based on emotional information, creating, for example, a forest landscape with gentle colors or a character with a friendly expression. The generated video data is then returned to the server.

[0183] Step 6:

[0184] The server streams the generated visual information and analyzed audio information to the device. The server transfers data in real time, and the device adjusts to seamlessly present the experience to the user. Through this integrated view, the user can experience an immersive experience.

[0185] (Application Example 2)

[0186] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0187] Analyzing users' emotions in real time and providing a personalized experience that integrates music and visuals has been difficult with conventional technology. Rapidly generating content suited to the user's emotions and creating an immersive experience through interaction is a challenge that needs to be addressed technologically.

[0188] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0189] In this invention, the server includes means for analyzing the user's emotional information to personalize the experience, means for analyzing and generating music and video information, and means for presenting the generated content to the user in real time. This enables the user to quickly enjoy a personalized experience tailored to their emotions.

[0190] "User" refers to a person or entity that uses the system.

[0191] "User interface" refers to the screens and input devices used by users to interact with a system.

[0192] "Music information" refers to audio data and related metadata, which constitute the audio content that the system analyzes.

[0193] "Data analysis" refers to the process of processing received information based on algorithms to obtain the intended result.

[0194] "Generative AI" refers to artificial intelligence technology that automatically generates new content based on received data.

[0195] "Visual information" refers to visual elements such as images and screen displays presented to the user.

[0196] "Emotional information" is data that represents the user's current emotional state and is acquired through facial expressions and voice.

[0197] "Personalization" refers to optimizing content and services to suit the individual user's preferences and circumstances.

[0198] The system for realizing this invention analyzes the user's emotional information and provides personalized visual information and music to the user using generative AI. The system consists of a user interface, an emotional analysis engine, generative AI, and servers and terminals responsible for data synchronization.

[0199] The server performs data analysis to integrate music and visual information based on the user's selected experience topic and emotional information. For this analysis, a camera and microphone for acquiring emotional data must be installed on the terminal. Software such as "Affectiva" or "Microsoft® Azure® Emotion API" is used for analyzing emotional information. Generative AI such as "OpenAI® GPT-4®" or "Midjourney" is used to generate visual information, while music information is acquired and analyzed in real time from music streaming services.

[0200] In this invention, the terminal synchronizes generated visual information and music transmitted from the server, presenting it to the user as an immersive experience. The terminal also receives feedback from the user, transmits it to the server, and stores it as data for future experience improvements.

[0201] As a concrete example, if a user who has finished a stressful day selects "relaxation" as the theme in the application, the generating AI will provide a tranquil forest scene and soothing piano music. In this case, an example of a prompt message would be: "The user's emotions are calm, and we want to generate a relaxing scene. Specifically, we will provide a combination of a quiet forest scene and gentle piano music."

[0202] This system allows users to instantly enjoy the experience that best suits their emotions at that moment.

[0203] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0204] Step 1:

[0205] The user launches the application and selects a theme to determine the subject they wish to experience. The user's theme selection information is provided as input and sent to the system via the user interface. The selected theme is then sent to the server as output.

[0206] Step 2:

[0207] The device uses a camera and microphone to record facial expressions and voice in order to acquire user emotional information. The input includes real-time video and audio of the user. An emotion engine analyzes this data and outputs the user's current emotional state. This emotional information is sent to a server.

[0208] Step 3:

[0209] Based on the subject selection information and sentiment information received by the server, music information is retrieved from a music streaming service. Subject and sentiment information are used as input, and the corresponding music tracks are output. This music is analyzed by the server, and its characteristics are extracted.

[0210] Step 4:

[0211] The server generates visual information using a generative AI model based on the analyzed musical characteristics and emotional information. The input for this step is musical characteristics and emotional information, and the output is a video tailored to the user. The generative AI constructs the specific video using prompts.

[0212] Step 5:

[0213] Visual information and music generated from the server are streamed to the terminal. The generated data is used as input, and the terminal receives and synchronizes it to present it to the user. The output is a harmonious presentation of visuals and music experienced by the user.

[0214] Step 6:

[0215] Users provide feedback on the experience they are given. As input, the user's feedback information is recorded on the device. As output, this data is sent to a server and stored as learning data for improving future experiences.

[0216] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0217] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include those described above. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions shown by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0218] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0219] [Second Embodiment]

[0220] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0221] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0222] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0223] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0224] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0225] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0226] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0227] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0228] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0229] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0230] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0231] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0232] This invention provides a system that allows users to enjoy the synergy of real-time generated scenery and music. This system mainly consists of a user's terminal, a server, and a music distribution service.

[0233] First, the user launches the application provided on their device and has the function to select the category of scenery they want to experience. The information selected by the user is then transferred to the server via the internet.

[0234] The server prepares to generate the appropriate landscape based on the selected category. It also begins streaming music data accessed by the device from a music streaming service. The music data is sent to the server for analysis of the music's rhythm, tempo, and other characteristics.

[0235] The server uses the results of this music analysis to generate appropriate landscapes in real time via a generating AI. The generated landscapes are linked to the characteristics of the music; for example, calm music will automatically create a tranquil natural landscape, while fast-paced music will create a breathtaking city nightscape.

[0236] Furthermore, the server generates characters that synchronize with the music and interact with the user. The characters speak different lines depending on the content and tone of the music, and can even engage in humorous conversations. This provides a more immersive experience.

[0237] The generated video and music are streamed to the device, which then displays them together on the screen. Users can enjoy the real-time changing scenery and music, creating a multi-faceted experience that engages both sight and sound.

[0238] The system's ability to collect user feedback, which is then analyzed and recorded by the server, is also crucial. This feedback information is used to improve the generating AI. Furthermore, users have the ability to customize their experience, adjusting settings such as scenery and dialogue according to their preferences. This customization information is also sent to the server and used to improve future experiences.

[0239] As a concrete example, consider a scenario where the user selects two different categories: "natural scenery" and "fantasy world." In the natural scenery category, gentle folk music plays, and a tranquil sunrise scene over a lake is displayed. The character is in the shape of a bird and provides simple trivia about nature. In contrast, in the fantasy world category, a colorful magical world is displayed along with dance music, and the character dances while greeting the user in a friendly manner. Thus, the present invention is a system that dynamically provides diverse experiences based on the user's choices.

[0240] The following describes the processing flow.

[0241] Step 1:

[0242] The user launches the application on their device and selects the category they want to experience from the displayed interface. This selection information is sent from the device to the server.

[0243] Step 2:

[0244] The device connects to a music streaming service via the internet and receives music data in real time based on the user's preferences. The received music data is then relayed to a server.

[0245] Step 3:

[0246] The server analyzes the received music data. This analysis includes extracting musical elements such as rhythm, tempo, and melody characteristics. The analysis results are used as input data for the generating AI.

[0247] Step 4:

[0248] Based on the analysis of the music, the server uses a generative AI to generate a landscape in real time that fits the selected category. The landscape changes in accordance with the pace and tone of the music.

[0249] Step 5:

[0250] The server generates characters within the same context as the generated landscape. The generating AI determines the characters' actions and dialogue based on the music's theme and mood. The characters' dialogue scenarios are also constructed at this stage.

[0251] Step 6:

[0252] The server streams generated scenery and character data to the device, which then displays it on the screen. The synchronization of music and video timing provides a visually and aurally consistent experience.

[0253] Step 7:

[0254] Users provide feedback through the interface presented during their experience. This feedback is sent from the device to the server and used to improve the system and update user profiles.

[0255] Step 8:

[0256] After the experience ends, users can use customization options to adjust the scenery, character movements, dialogue, and other elements. This customization information is sent from the device to the server and reflected in the next experience.

[0257] (Example 1)

[0258] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0259] In modern society, there is a demand for more personalized and immersive music listening experiences. However, existing systems struggle to provide dynamic visual experiences synchronized with music, and they lack options for users to finely customize that experience. Furthermore, there is a challenge in the insufficient mechanism for incorporating user feedback and continuously improving the system.

[0260] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0261] In this invention, the server includes means for providing a user screen for the user to select the type of experience, means for receiving and analyzing audio data, means for utilizing an artificial intelligence model that generates a landscape based on the analysis results, and means for generating a virtual character that is synchronized with the audio and interacting with it. This provides the user with an immersive experience that links music and visuals, further allows the user to customize the experience, and enables continuous improvement of the system through feedback.

[0262] "User interface for selecting the type of experience" refers to the interface that allows users to select scenery and experience content that are linked to their desired music.

[0263] "Receiving and analyzing audio data" refers to the process of analyzing the characteristics of music data, such as rhythm and tempo.

[0264] An "artificial intelligence model for generating landscapes" refers to an algorithm that generates appropriate visuals in real time based on the results of analyzing music data.

[0265] A "virtual character linked to sound" refers to a digital character that is generated according to the characteristics and content of the music and is capable of interacting with the user.

[0266] "Engaging in interaction" refers to virtual characters providing users with music-related information or offering interactive experiences through dialogue.

[0267] "Synchronizing and providing generated video and audio to users" refers to seamlessly presenting real-time generated visuals and music to the user.

[0268] "Providing choices" means offering within the system options that allow users to customize their experience and adjust the scenery and dialogue content to their own preferences.

[0269] This invention is a system that allows users to enjoy a real-time fusion of music and scenery. The system consists of a user's terminal, a server, and a music distribution service.

[0270] First, the user launches a dedicated application on their device and selects the category of scenery they wish to experience. A user interface is provided for selection, allowing the user to easily determine their preferred type. This selection information is then transmitted to a server via the internet.

[0271] The server receives the selected category information and prepares to access the music streaming service. Music data is streamed through the terminal, and the server receives and analyzes the music data. The analysis utilizes audio analysis software to extract the rhythm, tempo, and other features of the music.

[0272] The server also uses a generative AI model to generate appropriate landscapes in real time based on the analysis results. The generated landscapes are linked to the characteristics of the music; for example, calm music will create a tranquil natural landscape, while dynamic music will create an impressive urban landscape in real time.

[0273] Furthermore, the server generates a virtual character synchronized with the audio, enabling two-way interaction with the user. This character provides a deeper sense of immersion by using dialogue based on the content and tone of the music. Through interaction with the character, users can enjoy an experience where music and scenery are integrated.

[0274] The video and music are synchronized on the device, and the generated content is delivered seamlessly to the user. Users can also use customization options to adjust the scenery and character interactions to their liking.

[0275] For example, when a user selects the "Natural Landscape" category, the device receives calming music from a music streaming service. The server generates a tranquil lake landscape corresponding to this music, displays a bird character, and provides fun facts about nature. In this way, the user can enjoy a unique audiovisual experience based on the prompt message, "Please select a tranquil landscape and set it so that a character speaks simple nature facts accompanied by music that evokes a morning atmosphere."

[0276] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0277] Step 1:

[0278] The user launches the application through their device and selects the category they wish to experience from the displayed user screen. This input information is sent to the server via the internet. The selection information is stored in a database and used in the server's next processing step. Specifically, the user selects the "Natural Landscapes" category, and this selection information is sent to the server.

[0279] Step 2:

[0280] The server prepares for communication with the music distribution service based on the selected category information. The server sends a request to the terminal and starts streaming music data. The streamed music data is received by the server, and analysis such as the rhythm, tempo, and tone of the music is performed. The input data is a music file, and the output is an analysis result indicating the characteristics of the music. For example, data on a gentle rhythm and tone is extracted.

[0281] Step 3:

[0282] The server utilizes the analysis result, employs the generative AI model, and generates a scenery matching the music in real time. In this process, based on the analyzed characteristics of the music, the algorithm selects appropriate visual elements and dynamically constructs the scenery. The input is the characteristics of the music, and the output is the visually generated content in real time. For example, for gentle music, a scenery of a lake or forest is generated.

[0283] Step

[0284] The server further generates a virtual character based on the music characteristics and assigns a script. The character starts interacting with the user using lines associated with the music. The input is the music and the user's selection information, and the output is the lines and actions for interaction. Specifically, the server generates a bird character and provides topics related to nature.

[0285] Step 5:

[0286] The terminal receives the information on the generated scenery and character sent from the server and integrates and displays them on the display. At this time, the video and music are synchronized, and a seamless experience is provided to the user. The input is the data on the generated scenery and character, and the output is the integrated visual and auditory content for the user. For example, the scenery and character matching the gentle music are displayed on the screen.

[0287] Step 6:

[0288] User feedback is collected and sent to the server. The server stores this feedback in a database and uses it to improve future AI models. The input is user feedback data, and the output is training data aimed at improving the system. For example, based on the feedback, the AI ​​model is adjusted so that a more preferred landscape is generated in the next session.

[0289] (Application Example 1)

[0290] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0291] When users enjoy the synergy of music and visuals in real time, it is essential that the generation of scenery and characters is smoothly adapted to the characteristics of the music, thereby enhancing the quality of the viewing experience. Furthermore, incorporating user customization and feedback is a challenge in providing an experience that better suits individual needs.

[0292] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0293] In this invention, the server includes means for providing an operating unit for the user to select the type of experience, means for receiving and analyzing acoustic signals, and means for utilizing a generation AI to generate scenes based on the analysis results. This enables the user to experience immersive viewing experiences tailored to their individual needs by generating scenery and characters in harmony with music in real time.

[0294] A "user" refers to someone who uses and operates a system or application.

[0295] "Experience types" refer to the various themes and situations that users can choose from.

[0296] The term "operation unit" refers to the interface through which the user makes selections and gives instructions to the system.

[0297] "Acoustic signal" refers to sound signals that the system analyzes, such as music data and voice data.

[0298] "Analysis" refers to the process of analyzing the characteristics of acoustic signals and extracting guidelines for generating landscapes and characters.

[0299] "Scene" refers to the collective term for the images and visuals presented to the user.

[0300] "Generative AI" refers to artificial intelligence that uses machine learning techniques to automatically generate landscapes, characters, and other images.

[0301] A "moving character" refers to a character that appears on the screen in conjunction with an audio signal and interacts with the user through dialogue and actions.

[0302] The system implementing this invention uses a user terminal, a server, and a generative AI with music analysis capabilities. Processing begins when the user selects the type of experience they wish to experience via the control panel on the terminal. The terminal transmits this selection information to the server. The server receives audio signals via streaming from a music distribution service over the internet and analyzes those signals.

[0303] Based on the analysis results, the server uses a generative AI to generate corresponding scenes in real time. This generative AI has the ability to analyze characteristic data such as rhythm, tempo, and melody of the acoustic signal and design a scene that is appropriate for it. Furthermore, an action character is also generated according to the attributes of the acoustic signal and interacts with the user.

[0304] The generated scenes and sounds are streamed to the device in real time. The device integrates and displays this, allowing the user to enjoy a harmonious visual and auditory experience.

[0305] As a specific example, when a user selects music that allows them to relax, the generative AI presents a serene forest landscape in accordance with the music. Within that forest, an animated object in the shape of a bird appears, creating a scenario where natural trivia is provided to the user.

[0306] As an example sentence, a prompt sentence such as "Please generate the actions of a character that harmonizes with the serene night cityscape in accordance with jazz music." is used for the generative AI. In this way, the interaction between the AI and the user experience is realized.

[0307] This system can continuously improve the experience through customization and feedback by the user. Specifically, the content customized by the user and the feedback information are recorded on the server and utilized as learning data for the generative AI. This enables the provision of a more personalized experience.

[0308] The flow of the specific process in Application Example 1 will be described using FIG. 12.

[0309] Step 1:

[0310] The user selects the type of experience they want to have using the operation unit on the terminal. This input determines the theme of the experience and is transmitted to the server as data. At this stage, the user's selection information is accumulated on the server.

[0311] [[ID=​​​​​​​​​​​The server utilizes a generative AI model based on the analysis results and generates a suitable scene using prompt text. This AI takes the analyzed musical characteristics as input data and generates a visual scene that matches the prompt as output. For example, a prompt text for calm music might be "Generate a calm forest landscape."

[0315] Step 4:

[0316] The server simultaneously generates characters based on the tone and content of the sound. A generation AI model is used to depict the characters' actions and dialogue. In this process, the volume and rhythm of the sound are taken as input, and the character's movements and speech are output.

[0317] Step 5:

[0318] The generated scenes and audio are streamed in real time from the server to the terminal, which then integrates and displays them to the user. The terminal receives the data stream from the server and outputs synchronized audio and video to provide the user with an experience.

[0319] Step 6:

[0320] After the experience ends, the user provides feedback, which the device sends to the server. This feedback information is recorded on the server as input and used as training data for the generated AI model. The server uses this data to perform calculations to improve the quality of future experiences.

[0321] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0322] This invention provides a system that allows users to enjoy a synergy based on real-time generated scenery and music, and in particular, by combining it with an emotion engine that recognizes the user's emotions, it offers a more personalized experience. This system consists of a user's terminal, a server, a music distribution service, and an emotion engine.

[0323] First, the user launches an application on their device and selects a category of scenery they want to experience. This information is sent from the device to the server, and the device's emotion engine analyzes the user's facial expressions and voice captured through the camera and microphone to recognize the user's current emotional state.

[0324] The server receives this selection and emotional information and retrieves music data from a music streaming service. The music data is transmitted in real time, and the server analyzes the rhythm, tempo, and other musical characteristics. The analyzed information is used by a generative AI to create a landscape in real time that is appropriate to the user's emotional state.

[0325] The generating AI adjusts the color tone, composition, and dynamic elements of the landscape based on the output of the emotion engine. For example, it can generate a calm landscape when the user is relaxed and a lively landscape when they are excited. Characters also evolve to display appropriate expressions and actions according to emotions, supporting natural conversations.

[0326] The generated video and music are streamed to the device. The device seamlessly integrates the video and music, presenting it as a new, optimized experience for the user. The user can enjoy this immersive experience and simultaneously provide feedback through the interface. This feedback is sent from the device to the server and used as data to improve future experiences.

[0327] As a concrete example, consider a scenario where a user selects "natural scenery" with the goal of "relaxing," and the emotion engine recognizes the user's calm state. In this case, the server generates scenery of a quiet forest or a calm beach and selects music with a soothing melody. The character speaks to the user in a friendly and gentle tone. In this way, the present invention flexibly provides experiences based on the user's emotions and choices, making it possible to enrich daily life in urban areas.

[0328] The following describes the processing flow.

[0329] Step 1:

[0330] The user launches the application on their device and selects the category they want to experience from the interface. This selection information is sent from the device to the server.

[0331] Step 2:

[0332] The device uses its built-in camera and microphone to capture the user's facial expressions and voice. An emotion engine analyzes this data to identify the user's emotional state. The identified emotion information is then sent to the server.

[0333] Step 3:

[0334] The device accesses a music streaming service and receives music data tailored to the user's preferences in real time. The received music data is then transferred to a server.

[0335] Step 4:

[0336] The server analyzes the music data and extracts features such as tempo, rhythm, and genre. These analysis results are then used as a set of data necessary for the landscape generation process.

[0337] Step 5:

[0338] The server combines the user's emotional state, recognized by the emotion engine, with the music analysis results, and inputs this data into the generative AI. The generative AI then generates an appropriate scene in real time.

[0339] Step 6:

[0340] The generating AI adjusts the color scheme and composition of the landscape to match the user's emotions, creating a landscape with dynamic elements. Simultaneously, the server generates a character synchronized with the music, ready to interact with the user.

[0341] Step 7:

[0342] The server streams landscape data and character dialogue scripts generated by the server to the terminal. The terminal displays these on its screen, ensuring that music and video are seamlessly integrated.

[0343] Step 8:

[0344] Users utilize a feature to provide feedback during the experience. The device sends the entered feedback to a server for recording. This feedback is used to improve future experiences.

[0345] Step 9:

[0346] After the user finishes the experience, they can adjust the scenery and dialogue using customization options. The device sends the customization information to the server, and it is reflected in the next experience.

[0347] (Example 2)

[0348] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0349] In modern society, customized entertainment experiences tailored to the emotional state of users are limited. In particular, there is a lack of means to create immersive experiences where music and visual content are closely integrated. Furthermore, there is a need for mechanisms that incorporate user feedback into the system to improve individual experiences.

[0350] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0351] In this invention, the server includes means for providing a display device for the user to select the type of experience, means for acquiring and analyzing sound information, and means for utilizing a generative model that generates a landscape based on the analysis results and the user's emotional information. This makes it possible for music and visual content to be synchronized according to the user's emotions, providing a personalized and immersive experience.

[0352] "User" refers to an individual or group that operates the system and selects the type of experience.

[0353] "Experience type" refers to the category or theme of the content that the user wants to experience.

[0354] A "display device" refers to a device or software that provides an interface used by users to select the type of experience they are experiencing.

[0355] "Audio information" refers to data that a system processes, including music and other sound data.

[0356] "Analysis" refers to the process of processing and calculating acquired sound and emotional information.

[0357] "Emotional information" refers to the emotional state estimated from facial expressions and voices acquired through the user's camera and microphone.

[0358] A "generative model" refers to a system that uses artificial intelligence to generate new landscapes and content based on analysis results and emotional information.

[0359] A "virtual character" is a generated character that interacts with sound information and enables dialogue with the user.

[0360] "Visual information" refers to generated video data such as landscapes, which is presented to the user in conjunction with audio information.

[0361] This invention is a system that enables users to enjoy a customized entertainment experience tailored to their emotions in real time. This system mainly consists of three main elements: a terminal, a server, and an audio service.

[0362] First, users can select their preferred content category via a display on their device that allows them to choose the type of experience. This selection information is then sent from the device to the server. The device uses its built-in emotion engine to analyze emotional information from the user's facial expressions and voice captured through the camera and microphone. This information indicates the user's current emotions and is used to customize the content provided by the system.

[0363] The server uses the selection and emotion information received from the user to retrieve sound information from the audio service. The retrieved sound information is then analyzed on the server, and features such as rhythm and tempo are identified. Based on these features and emotion information, the server generates prompts and inputs them into a generative AI model. This model generates visual information such as landscapes and characters based on the specified information. For example, if the user selects "I want to relax," the generative AI model will depict calm natural scenery and friendly characters.

[0364] The generated visual information and acquired audio information are transmitted from the server to the terminal in real time. The terminal integrates these and presents them seamlessly to the user, providing an immersive experience. For example, if the user selects "I want to relax" and the system recognizes that the user is in a calm state based on emotional information, calming music and visual natural scenery appropriate to that situation will be presented.

[0365] An example of a prompt message would be: "The user has selected the category 'Natural Landscapes' and their emotional state is 'Calm.' Based on this, please generate a relaxing landscape and corresponding music." In this way, the system can provide an entertainment experience that matches the user's individual emotions.

[0366] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0367] Step 1:

[0368] The user launches the application on their device and selects the type of content they want to experience. This input information is saved on the device and then transferred to a server via the network. Users can choose categories such as "natural landscapes" or "urban scenery."

[0369] Step 2:

[0370] The device operates a built-in emotion engine and uses the user's camera and microphone to collect facial expressions and voice. This collected data is then analyzed by an emotion analysis algorithm to identify the user's emotional state (e.g., "relaxed," "excited," etc.). The analysis output is sent to the server as the user's emotion information.

[0371] Step 3:

[0372] The server receives selection and emotion information from the user. Based on this information, the server connects to an audio service and obtains appropriate sound information. Here, the server requests the optimal music data according to the user's selection and emotion, and the server analyzes the characteristics of the obtained music, such as rhythm, tempo, and melody. The output of this analysis becomes part of the generation prompt.

[0373] Step 4:

[0374] The server generates prompts for visual information based on the analyzed sound and emotion information. These prompts include information about the sound rhythm, the user's emotional state, and the landscape according to the selected category. For example, a specific prompt statement such as "The user's selected category is 'Natural Landscape', and their emotional state is 'Calm'" is passed to the generating AI model.

[0375] Step 5:

[0376] The generation AI model begins generating landscapes and characters based on prompts received from the server. The model adjusts color tones and behavioral elements based on emotional information, creating, for example, a forest landscape with gentle colors or a character with a friendly expression. The generated video data is then returned to the server.

[0377] Step 6:

[0378] The server streams the generated visual information and analyzed audio information to the device. The server transfers data in real time, and the device adjusts to seamlessly present the experience to the user. Through this integrated view, the user can experience an immersive experience.

[0379] (Application Example 2)

[0380] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0381] Analyzing users' emotions in real time and providing a personalized experience that integrates music and visuals has been difficult with conventional technology. Rapidly generating content suited to the user's emotions and creating an immersive experience through interaction is a challenge that needs to be addressed technologically.

[0382] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0383] In this invention, the server includes means for analyzing the user's emotional information to personalize the experience, means for analyzing and generating music and video information, and means for presenting the generated content to the user in real time. This enables the user to quickly enjoy a personalized experience tailored to their emotions.

[0384] "User" refers to a person or entity that uses the system.

[0385] "User interface" refers to the screens and input devices used by users to interact with a system.

[0386] "Music information" refers to audio data and related metadata, which constitute the audio content that the system analyzes.

[0387] "Data analysis" refers to the process of processing received information based on algorithms to obtain the intended result.

[0388] "Generative AI" refers to artificial intelligence technology that automatically generates new content based on received data.

[0389] "Visual information" refers to visual elements such as images and screen displays presented to the user.

[0390] "Emotional information" is data that represents the user's current emotional state and is acquired through facial expressions and voice.

[0391] "Personalization" refers to optimizing content and services to suit the individual user's preferences and circumstances.

[0392] The system for realizing this invention analyzes the user's emotional information and provides personalized visual information and music to the user using generative AI. The system consists of a user interface, an emotional analysis engine, generative AI, and servers and terminals responsible for data synchronization.

[0393] The server performs data analysis to integrate music and visual information based on the user's selected experience topic and emotional information. For this analysis, a camera and microphone for acquiring emotional data must be installed on the terminal. Software such as "Affectiva" or "Microsoft Azure Emotion API" are used for analyzing emotional information. Generative AI such as "OpenAI's GPT-4" or "Midjourney" are used to generate visual information, and music information is acquired and analyzed in real time from music streaming services.

[0394] In this invention, the terminal synchronizes generated visual information and music transmitted from the server, presenting it to the user as an immersive experience. The terminal also receives feedback from the user, transmits it to the server, and stores it as data for future experience improvements.

[0395] As a concrete example, if a user who has finished a stressful day selects "relaxation" as the theme in the application, the generating AI will provide a tranquil forest scene and soothing piano music. In this case, an example of a prompt message would be: "The user's emotions are calm, and we want to generate a relaxing scene. Specifically, we will provide a combination of a quiet forest scene and gentle piano music."

[0396] This system allows users to instantly enjoy the experience that best suits their emotions at that moment.

[0397] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0398] Step 1:

[0399] The user launches the application and selects a theme to determine the subject they wish to experience. The user's theme selection information is provided as input and sent to the system via the user interface. The selected theme is then sent to the server as output.

[0400] Step 2:

[0401] The device uses a camera and microphone to record facial expressions and voice in order to acquire user emotional information. The input includes real-time video and audio of the user. An emotion engine analyzes this data and outputs the user's current emotional state. This emotional information is sent to a server.

[0402] Step 3:

[0403] Based on the subject selection information and sentiment information received by the server, music information is retrieved from a music streaming service. Subject and sentiment information are used as input, and the corresponding music tracks are output. This music is analyzed by the server, and its characteristics are extracted.

[0404] Step 4:

[0405] The server generates visual information using a generative AI model based on the analyzed musical characteristics and emotional information. The input for this step is musical characteristics and emotional information, and the output is a video tailored to the user. The generative AI constructs the specific video using prompts.

[0406] Step 5:

[0407] Visual information and music generated from the server are streamed to the terminal. The generated data is used as input, and the terminal receives and synchronizes it to present it to the user. The output is a harmonious presentation of visuals and music experienced by the user.

[0408] Step 6:

[0409] Users provide feedback on the experience they are given. As input, the user's feedback information is recorded on the device. As output, this data is sent to a server and stored as learning data for improving future experiences.

[0410] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0411] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include those described above. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions shown by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0412] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0413] [Third Embodiment]

[0414] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0415] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0416] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0417] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0418] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0419] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0420] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0421] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0422] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0423] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0424] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0425] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0426] This invention provides a system that allows users to enjoy the synergy of real-time generated scenery and music. This system mainly consists of a user's terminal, a server, and a music distribution service.

[0427] First, the user launches the application provided on their device and has the function to select the category of scenery they want to experience. The information selected by the user is then transferred to the server via the internet.

[0428] The server prepares to generate the appropriate landscape based on the selected category. It also begins streaming music data accessed by the device from a music streaming service. The music data is sent to the server for analysis of the music's rhythm, tempo, and other characteristics.

[0429] The server uses the results of this music analysis to generate appropriate landscapes in real time via a generating AI. The generated landscapes are linked to the characteristics of the music; for example, calm music will automatically create a tranquil natural landscape, while fast-paced music will create a breathtaking city nightscape.

[0430] Furthermore, the server generates characters that synchronize with the music and interact with the user. The characters speak different lines depending on the content and tone of the music, and can even engage in humorous conversations. This provides a more immersive experience.

[0431] The generated video and music are streamed to the device, which then displays them together on the screen. Users can enjoy the real-time changing scenery and music, creating a multi-faceted experience that engages both sight and sound.

[0432] The system's ability to collect user feedback, which is then analyzed and recorded by the server, is also crucial. This feedback information is used to improve the generating AI. Furthermore, users have the ability to customize their experience, adjusting settings such as scenery and dialogue according to their preferences. This customization information is also sent to the server and used to improve future experiences.

[0433] As a concrete example, consider a scenario where the user selects two different categories: "natural scenery" and "fantasy world." In the natural scenery category, gentle folk music plays, and a tranquil sunrise scene over a lake is displayed. The character is in the shape of a bird and provides simple trivia about nature. In contrast, in the fantasy world category, a colorful magical world is displayed along with dance music, and the character dances while greeting the user in a friendly manner. Thus, the present invention is a system that dynamically provides diverse experiences based on the user's choices.

[0434] The following describes the processing flow.

[0435] Step 1:

[0436] The user launches the application on their device and selects the category they want to experience from the displayed interface. This selection information is sent from the device to the server.

[0437] Step 2:

[0438] The device connects to a music streaming service via the internet and receives music data in real time based on the user's preferences. The received music data is then relayed to a server.

[0439] Step 3:

[0440] The server analyzes the received music data. This analysis includes extracting musical elements such as rhythm, tempo, and melody characteristics. The analysis results are used as input data for the generating AI.

[0441] Step 4:

[0442] Based on the analysis of the music, the server uses a generative AI to generate a landscape in real time that fits the selected category. The landscape changes in accordance with the pace and tone of the music.

[0443] Step 5:

[0444] The server generates characters within the same context as the generated landscape. The generating AI determines the characters' actions and dialogue based on the music's theme and mood. The characters' dialogue scenarios are also constructed at this stage.

[0445] Step 6:

[0446] The server streams generated scenery and character data to the device, which then displays it on the screen. The synchronization of music and video timing provides a visually and aurally consistent experience.

[0447] Step 7:

[0448] Users provide feedback through the interface presented during their experience. This feedback is sent from the device to the server and used to improve the system and update user profiles.

[0449] Step 8:

[0450] After the experience ends, users can use customization options to adjust the scenery, character movements, dialogue, and other elements. This customization information is sent from the device to the server and reflected in the next experience.

[0451] (Example 1)

[0452] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0453] In modern society, there is a demand for more personalized and immersive music listening experiences. However, existing systems struggle to provide dynamic visual experiences synchronized with music, and they lack options for users to finely customize that experience. Furthermore, there is a challenge in the insufficient mechanism for incorporating user feedback and continuously improving the system.

[0454] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0455] In this invention, the server includes means for providing a user screen for the user to select the type of experience, means for receiving and analyzing audio data, means for utilizing an artificial intelligence model that generates a landscape based on the analysis results, and means for generating a virtual character that is synchronized with the audio and interacting with it. This provides the user with an immersive experience that links music and visuals, further allows the user to customize the experience, and enables continuous improvement of the system through feedback.

[0456] "User interface for selecting the type of experience" refers to the interface that allows users to select scenery and experience content that are linked to their desired music.

[0457] "Receiving and analyzing audio data" refers to the process of analyzing the characteristics of music data, such as rhythm and tempo.

[0458] An "artificial intelligence model for generating landscapes" refers to an algorithm that generates appropriate visuals in real time based on the results of analyzing music data.

[0459] A "virtual character linked to sound" refers to a digital character that is generated according to the characteristics and content of the music and is capable of interacting with the user.

[0460] "Engaging in interaction" refers to virtual characters providing users with music-related information or offering interactive experiences through dialogue.

[0461] "Synchronizing and providing generated video and audio to users" refers to seamlessly presenting real-time generated visuals and music to the user.

[0462] "Providing choices" means offering within the system options that allow users to customize their experience and adjust the scenery and dialogue content to their own preferences.

[0463] This invention is a system that allows users to enjoy a real-time fusion of music and scenery. The system consists of a user's terminal, a server, and a music distribution service.

[0464] First, the user launches a dedicated application on their device and selects the category of scenery they wish to experience. A user interface is provided for selection, allowing the user to easily determine their preferred type. This selection information is then transmitted to a server via the internet.

[0465] The server receives the selected category information and prepares to access the music streaming service. Music data is streamed through the terminal, and the server receives and analyzes the music data. The analysis utilizes audio analysis software to extract the rhythm, tempo, and other features of the music.

[0466] The server also uses a generative AI model to generate appropriate landscapes in real time based on the analysis results. The generated landscapes are linked to the characteristics of the music; for example, calm music will create a tranquil natural landscape, while dynamic music will create an impressive urban landscape in real time.

[0467] Furthermore, the server generates a virtual character synchronized with the audio, enabling two-way interaction with the user. This character provides a deeper sense of immersion by using dialogue based on the content and tone of the music. Through interaction with the character, users can enjoy an experience where music and scenery are integrated.

[0468] The video and music are synchronized on the device, and the generated content is delivered seamlessly to the user. Users can also use customization options to adjust the scenery and character interactions to their liking.

[0469] For example, when a user selects the "Natural Landscape" category, the device receives calming music from a music streaming service. The server generates a tranquil lake landscape corresponding to this music, displays a bird character, and provides fun facts about nature. In this way, the user can enjoy a unique audiovisual experience based on the prompt message, "Please select a tranquil landscape and set it so that a character speaks simple nature facts accompanied by music that evokes a morning atmosphere."

[0470] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0471] Step 1:

[0472] The user launches the application through their device and selects the category they wish to experience from the displayed user screen. This input information is sent to the server via the internet. The selection information is stored in a database and used in the server's next processing step. Specifically, the user selects the "Natural Landscapes" category, and this selection information is sent to the server.

[0473] Step 2:

[0474] The server prepares to communicate with the music streaming service based on the selected category information. The server sends a request to the terminal and starts streaming music data. The streamed music data is received by the server, and analysis of the music's rhythm, tempo, tone, etc., is performed. The input data is a music file, and the output is the analysis results showing the characteristics of the music. For example, data on a gentle rhythm and tone may be extracted.

[0475] Step 3:

[0476] The server uses the analysis results and leverages a generative AI model to generate landscapes that match the music in real time. In this process, the algorithm selects appropriate visual elements based on the characteristics of the analyzed music, dynamically constructing the landscape. The input is the characteristics of the music, and the output is a visually generated content in real time. For example, for calm music, landscapes of lakes and forests are generated.

[0477] Step 4:

[0478] The server then generates virtual characters based on musical characteristics and assigns them scripts. These characters initiate interaction with the user using dialogue synchronized with the music. Input is the music and user selections, while output is dialogue and actions for interaction. Specifically, the server generates a bird character and provides topics related to nature.

[0479] Step 5:

[0480] The terminal receives generated scenery and character information transmitted from the server and integrates and displays it on the screen. At this time, the video and music are synchronized, providing the user with a seamless experience. The input is the generated scenery and character data, and the output is integrated visual and auditory content for the user. For example, scenery and characters synchronized with gentle music are displayed on the screen.

[0481] Step 6:

[0482] User feedback is collected and sent to the server. The server stores this feedback in a database and uses it to improve future AI models. The input is user feedback data, and the output is training data aimed at improving the system. For example, based on the feedback, the AI ​​model is adjusted so that a more preferred landscape is generated in the next session.

[0483] (Application Example 1)

[0484] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0485] When users enjoy the synergy of music and visuals in real time, it is essential that the generation of scenery and characters is smoothly adapted to the characteristics of the music, thereby enhancing the quality of the viewing experience. Furthermore, incorporating user customization and feedback is a challenge in providing an experience that better suits individual needs.

[0486] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0487] In this invention, the server includes means for providing an operating unit for the user to select the type of experience, means for receiving and analyzing acoustic signals, and means for utilizing a generation AI to generate scenes based on the analysis results. This enables the user to experience immersive viewing experiences tailored to their individual needs by generating scenery and characters in harmony with music in real time.

[0488] A "user" refers to someone who uses and operates a system or application.

[0489] "Experience types" refer to the various themes and situations that users can choose from.

[0490] The term "operation unit" refers to the interface through which the user makes selections and gives instructions to the system.

[0491] "Acoustic signal" refers to sound signals that the system analyzes, such as music data and voice data.

[0492] "Analysis" refers to the process of analyzing the characteristics of acoustic signals and extracting guidelines for generating landscapes and characters.

[0493] "Scene" refers to the collective term for the images and visuals presented to the user.

[0494] "Generative AI" refers to artificial intelligence that uses machine learning techniques to automatically generate landscapes, characters, and other images.

[0495] A "moving character" refers to a character that appears on the screen in conjunction with an audio signal and interacts with the user through dialogue and actions.

[0496] The system implementing this invention uses a user terminal, a server, and a generative AI with music analysis capabilities. Processing begins when the user selects the type of experience they wish to experience via the control panel on the terminal. The terminal transmits this selection information to the server. The server receives audio signals via streaming from a music distribution service over the internet and analyzes those signals.

[0497] Based on the analysis results, the server uses a generative AI to generate corresponding scenes in real time. This generative AI has the ability to analyze characteristic data such as rhythm, tempo, and melody of the acoustic signal and design a scene that is appropriate for it. Furthermore, an action character is also generated according to the attributes of the acoustic signal and interacts with the user.

[0498] The generated scenes and sounds are streamed to the device in real time. The device integrates and displays this, allowing the user to enjoy a harmonious visual and auditory experience.

[0499] As a concrete example, if the user selects relaxing music, the AI ​​will present a tranquil forest landscape in sync with that music. Within that forest, a moving object in the shape of a small bird will appear, creating a scene that provides the user with interesting facts about nature.

[0500] As an example, the prompt given to the generation AI is, "Generate character movements that harmonize with a tranquil nighttime cityscape, set to jazz music." In this way, interaction between the AI ​​and the user experience is realized.

[0501] This system can continuously improve the user experience through user customization and feedback. Specifically, user customizations and feedback information are recorded on the server and used as training data for the generating AI. This enables the provision of a more personalized experience.

[0502] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0503] Step 1:

[0504] The user selects the type of experience they wish to have using the controls on their device. This input determines the theme of the experience and is sent to the server as data. At this stage, the user's selection information is collected on the server.

[0505] Step 2:

[0506] The server receives audio signals related to a specified experience via streaming from a music streaming service. This audio data is used as input data to extract the characteristics of the music. The server analyzes the audio signals and performs data processing to identify rhythm, tempo, pitch, and other characteristics.

[0507] Step 3:

[0508] The server utilizes a generative AI model based on the analysis results and generates a suitable scene using prompt text. This AI takes the analyzed musical characteristics as input data and generates a visual scene that matches the prompt as output. For example, a prompt text for calm music might be "Generate a calm forest landscape."

[0509] Step 4:

[0510] The server simultaneously generates characters based on the tone and content of the sound. A generation AI model is used to depict the characters' actions and dialogue. In this process, the volume and rhythm of the sound are taken as input, and the character's movements and speech are output.

[0511] Step 5:

[0512] The generated scenes and audio are streamed in real time from the server to the terminal, which then integrates and displays them to the user. The terminal receives the data stream from the server and outputs synchronized audio and video to provide the user with an experience.

[0513] Step 6:

[0514] After the experience ends, the user provides feedback, which the device sends to the server. This feedback information is recorded on the server as input and used as training data for the generated AI model. The server uses this data to perform calculations to improve the quality of future experiences.

[0515] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0516] This invention provides a system that allows users to enjoy a synergy based on real-time generated scenery and music, and in particular, by combining it with an emotion engine that recognizes the user's emotions, it offers a more personalized experience. This system consists of a user's terminal, a server, a music distribution service, and an emotion engine.

[0517] First, the user launches an application on their device and selects a category of scenery they want to experience. This information is sent from the device to the server, and the device's emotion engine analyzes the user's facial expressions and voice captured through the camera and microphone to recognize the user's current emotional state.

[0518] The server receives this selection and emotional information and retrieves music data from a music streaming service. The music data is transmitted in real time, and the server analyzes the rhythm, tempo, and other musical characteristics. The analyzed information is used by a generative AI to create a landscape in real time that is appropriate to the user's emotional state.

[0519] The generating AI adjusts the color tone, composition, and dynamic elements of the landscape based on the output of the emotion engine. For example, it can generate a calm landscape when the user is relaxed and a lively landscape when they are excited. Characters also evolve to display appropriate expressions and actions according to emotions, supporting natural conversations.

[0520] The generated video and music are streamed to the device. The device seamlessly integrates the video and music, presenting it as a new, optimized experience for the user. The user can enjoy this immersive experience and simultaneously provide feedback through the interface. This feedback is sent from the device to the server and used as data to improve future experiences.

[0521] As a concrete example, consider a scenario where a user selects "natural scenery" with the goal of "relaxing," and the emotion engine recognizes the user's calm state. In this case, the server generates scenery of a quiet forest or a calm beach and selects music with a soothing melody. The character speaks to the user in a friendly and gentle tone. In this way, the present invention flexibly provides experiences based on the user's emotions and choices, making it possible to enrich daily life in urban areas.

[0522] The following describes the processing flow.

[0523] Step 1:

[0524] The user launches the application on their device and selects the category they want to experience from the interface. This selection information is sent from the device to the server.

[0525] Step 2:

[0526] The device uses its built-in camera and microphone to capture the user's facial expressions and voice. An emotion engine analyzes this data to identify the user's emotional state. The identified emotion information is then sent to the server.

[0527] Step 3:

[0528] The device accesses a music streaming service and receives music data tailored to the user's preferences in real time. The received music data is then transferred to a server.

[0529] Step 4:

[0530] The server analyzes the music data and extracts features such as tempo, rhythm, and genre. These analysis results are then used as a set of data necessary for the landscape generation process.

[0531] Step 5:

[0532] The server combines the user's emotional state, recognized by the emotion engine, with the music analysis results, and inputs this data into the generative AI. The generative AI then generates an appropriate scene in real time.

[0533] Step 6:

[0534] The generating AI adjusts the color scheme and composition of the landscape to match the user's emotions, creating a landscape with dynamic elements. Simultaneously, the server generates a character synchronized with the music, ready to interact with the user.

[0535] Step 7:

[0536] The server streams landscape data and character dialogue scripts generated by the server to the terminal. The terminal displays these on its screen, ensuring that music and video are seamlessly integrated.

[0537] Step 8:

[0538] Users utilize a feature to provide feedback during the experience. The device sends the entered feedback to a server for recording. This feedback is used to improve future experiences.

[0539] Step 9:

[0540] After the user finishes the experience, they can adjust the scenery and dialogue using customization options. The device sends the customization information to the server, and it is reflected in the next experience.

[0541] (Example 2)

[0542] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0543] In modern society, customized entertainment experiences tailored to the emotional state of users are limited. In particular, there is a lack of means to create immersive experiences where music and visual content are closely integrated. Furthermore, there is a need for mechanisms that incorporate user feedback into the system to improve individual experiences.

[0544] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0545] In this invention, the server includes means for providing a display device for the user to select the type of experience, means for acquiring and analyzing sound information, and means for utilizing a generative model that generates a landscape based on the analysis results and the user's emotional information. This makes it possible for music and visual content to be synchronized according to the user's emotions, providing a personalized and immersive experience.

[0546] "User" refers to an individual or group that operates the system and selects the type of experience.

[0547] "Experience type" refers to the category or theme of the content that the user wants to experience.

[0548] A "display device" refers to a device or software that provides an interface used by users to select the type of experience they are experiencing.

[0549] "Audio information" refers to data that a system processes, including music and other sound data.

[0550] "Analysis" refers to the process of processing and calculating acquired sound and emotional information.

[0551] "Emotional information" refers to the emotional state estimated from facial expressions and voices acquired through the user's camera and microphone.

[0552] A "generative model" refers to a system that uses artificial intelligence to generate new landscapes and content based on analysis results and emotional information.

[0553] A "virtual character" is a generated character that interacts with sound information and enables dialogue with the user.

[0554] "Visual information" refers to generated video data such as landscapes, which is presented to the user in conjunction with audio information.

[0555] This invention is a system that enables users to enjoy a customized entertainment experience tailored to their emotions in real time. This system mainly consists of three main elements: a terminal, a server, and an audio service.

[0556] First, users can select their preferred content category via a display on their device that allows them to choose the type of experience. This selection information is then sent from the device to the server. The device uses its built-in emotion engine to analyze emotional information from the user's facial expressions and voice captured through the camera and microphone. This information indicates the user's current emotions and is used to customize the content provided by the system.

[0557] The server uses the selection and emotion information received from the user to retrieve sound information from the audio service. The retrieved sound information is then analyzed on the server, and features such as rhythm and tempo are identified. Based on these features and emotion information, the server generates prompts and inputs them into a generative AI model. This model generates visual information such as landscapes and characters based on the specified information. For example, if the user selects "I want to relax," the generative AI model will depict calm natural scenery and friendly characters.

[0558] The generated visual information and acquired audio information are transmitted from the server to the terminal in real time. The terminal integrates these and presents them seamlessly to the user, providing an immersive experience. For example, if the user selects "I want to relax" and the system recognizes that the user is in a calm state based on emotional information, calming music and visual natural scenery appropriate to that situation will be presented.

[0559] An example of a prompt message would be: "The user has selected the category 'Natural Landscapes' and their emotional state is 'Calm.' Based on this, please generate a relaxing landscape and corresponding music." In this way, the system can provide an entertainment experience that matches the user's individual emotions.

[0560] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0561] Step 1:

[0562] The user launches the application on their device and selects the type of content they want to experience. This input information is saved on the device and then transferred to a server via the network. Users can choose categories such as "natural landscapes" or "urban scenery."

[0563] Step 2:

[0564] The device operates a built-in emotion engine and uses the user's camera and microphone to collect facial expressions and voice. This collected data is then analyzed by an emotion analysis algorithm to identify the user's emotional state (e.g., "relaxed," "excited," etc.). The analysis output is sent to the server as the user's emotion information.

[0565] Step 3:

[0566] The server receives selection and emotion information from the user. Based on this information, the server connects to an audio service and obtains appropriate sound information. Here, the server requests the optimal music data according to the user's selection and emotion, and the server analyzes the characteristics of the obtained music, such as rhythm, tempo, and melody. The output of this analysis becomes part of the generation prompt.

[0567] Step 4:

[0568] The server generates prompts for visual information based on the analyzed sound and emotion information. These prompts include information about the sound rhythm, the user's emotional state, and the landscape according to the selected category. For example, a specific prompt statement such as "The user's selected category is 'Natural Landscape', and their emotional state is 'Calm'" is passed to the generating AI model.

[0569] Step 5:

[0570] The generation AI model begins generating landscapes and characters based on prompts received from the server. The model adjusts color tones and behavioral elements based on emotional information, creating, for example, a forest landscape with gentle colors or a character with a friendly expression. The generated video data is then returned to the server.

[0571] Step 6:

[0572] The server streams the generated visual information and analyzed audio information to the device. The server transfers data in real time, and the device adjusts to seamlessly present the experience to the user. Through this integrated view, the user can experience an immersive experience.

[0573] (Application Example 2)

[0574] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0575] Analyzing users' emotions in real time and providing a personalized experience that integrates music and visuals has been difficult with conventional technology. Rapidly generating content suited to the user's emotions and creating an immersive experience through interaction is a challenge that needs to be addressed technologically.

[0576] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0577] In this invention, the server includes means for analyzing the user's emotional information to personalize the experience, means for analyzing and generating music and video information, and means for presenting the generated content to the user in real time. This enables the user to quickly enjoy a personalized experience tailored to their emotions.

[0578] "User" refers to a person or entity that uses the system.

[0579] "User interface" refers to the screens and input devices used by users to interact with a system.

[0580] "Music information" refers to audio data and related metadata, which constitute the audio content that the system analyzes.

[0581] "Data analysis" refers to the process of processing received information based on algorithms to obtain the intended result.

[0582] "Generative AI" refers to artificial intelligence technology that automatically generates new content based on received data.

[0583] "Visual information" refers to visual elements such as images and screen displays presented to the user.

[0584] "Emotional information" is data that represents the user's current emotional state and is acquired through facial expressions and voice.

[0585] "Personalization" refers to optimizing content and services to suit the individual user's preferences and circumstances.

[0586] The system for realizing this invention analyzes the user's emotional information and provides personalized visual information and music to the user using generative AI. The system consists of a user interface, an emotional analysis engine, generative AI, and servers and terminals responsible for data synchronization.

[0587] The server performs data analysis to integrate music and visual information based on the user's selected experience topic and emotional information. For this analysis, a camera and microphone for acquiring emotional data must be installed on the terminal. Software such as "Affectiva" or "Microsoft Azure Emotion API" are used for analyzing emotional information. Generative AI such as "OpenAI's GPT-4" or "Midjourney" are used to generate visual information, and music information is acquired and analyzed in real time from music streaming services.

[0588] In this invention, the terminal synchronizes generated visual information and music transmitted from the server, presenting it to the user as an immersive experience. The terminal also receives feedback from the user, transmits it to the server, and stores it as data for future experience improvements.

[0589] As a concrete example, if a user who has finished a stressful day selects "relaxation" as the theme in the application, the generating AI will provide a tranquil forest scene and soothing piano music. In this case, an example of a prompt message would be: "The user's emotions are calm, and we want to generate a relaxing scene. Specifically, we will provide a combination of a quiet forest scene and gentle piano music."

[0590] This system allows users to instantly enjoy the experience that best suits their emotions at that moment.

[0591] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0592] Step 1:

[0593] The user launches the application and selects a theme to determine the subject they wish to experience. The user's theme selection information is provided as input and sent to the system via the user interface. The selected theme is then sent to the server as output.

[0594] Step 2:

[0595] The device uses a camera and microphone to record facial expressions and voice in order to acquire user emotional information. The input includes real-time video and audio of the user. An emotion engine analyzes this data and outputs the user's current emotional state. This emotional information is sent to a server.

[0596] Step 3:

[0597] Based on the subject selection information and sentiment information received by the server, music information is retrieved from a music streaming service. Subject and sentiment information are used as input, and the corresponding music tracks are output. This music is analyzed by the server, and its characteristics are extracted.

[0598] Step 4:

[0599] The server generates visual information using a generative AI model based on the analyzed musical characteristics and emotional information. The input for this step is musical characteristics and emotional information, and the output is a video tailored to the user. The generative AI constructs the specific video using prompts.

[0600] Step 5:

[0601] Visual information and music generated from the server are streamed to the terminal. The generated data is used as input, and the terminal receives and synchronizes it to present it to the user. The output is a harmonious presentation of visuals and music experienced by the user.

[0602] Step 6:

[0603] Users provide feedback on the experience they are given. As input, the user's feedback information is recorded on the device. As output, this data is sent to a server and stored as learning data for improving future experiences.

[0604] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0605] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include those described above. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions shown by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0606] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0607] [Fourth Embodiment]

[0608] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0609] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0610] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0611] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0612] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0613] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0614] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0615] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0616] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0617] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0618] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0619] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0620] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0621] This invention provides a system that allows users to enjoy the synergy of real-time generated scenery and music. This system mainly consists of a user's terminal, a server, and a music distribution service.

[0622] First, the user launches the application provided on their device and has the function to select the category of scenery they want to experience. The information selected by the user is then transferred to the server via the internet.

[0623] The server prepares to generate the appropriate landscape based on the selected category. It also begins streaming music data accessed by the device from a music streaming service. The music data is sent to the server for analysis of the music's rhythm, tempo, and other characteristics.

[0624] The server uses the results of this music analysis to generate appropriate landscapes in real time via a generating AI. The generated landscapes are linked to the characteristics of the music; for example, calm music will automatically create a tranquil natural landscape, while fast-paced music will create a breathtaking city nightscape.

[0625] Furthermore, the server generates characters that synchronize with the music and interact with the user. The characters speak different lines depending on the content and tone of the music, and can even engage in humorous conversations. This provides a more immersive experience.

[0626] The generated video and music are streamed to the device, which then displays them together on the screen. Users can enjoy the real-time changing scenery and music, creating a multi-faceted experience that engages both sight and sound.

[0627] The system's ability to collect user feedback, which is then analyzed and recorded by the server, is also crucial. This feedback information is used to improve the generating AI. Furthermore, users have the ability to customize their experience, adjusting settings such as scenery and dialogue according to their preferences. This customization information is also sent to the server and used to improve future experiences.

[0628] As a concrete example, consider a scenario where the user selects two different categories: "natural scenery" and "fantasy world." In the natural scenery category, gentle folk music plays, and a tranquil sunrise scene over a lake is displayed. The character is in the shape of a bird and provides simple trivia about nature. In contrast, in the fantasy world category, a colorful magical world is displayed along with dance music, and the character dances while greeting the user in a friendly manner. Thus, the present invention is a system that dynamically provides diverse experiences based on the user's choices.

[0629] The following describes the processing flow.

[0630] Step 1:

[0631] The user launches the application on their device and selects the category they want to experience from the displayed interface. This selection information is sent from the device to the server.

[0632] Step 2:

[0633] The device connects to a music streaming service via the internet and receives music data in real time based on the user's preferences. The received music data is then relayed to a server.

[0634] Step 3:

[0635] The server analyzes the received music data. This analysis includes extracting musical elements such as rhythm, tempo, and melody characteristics. The analysis results are used as input data for the generating AI.

[0636] Step 4:

[0637] Based on the analysis of the music, the server uses a generative AI to generate a landscape in real time that fits the selected category. The landscape changes in accordance with the pace and tone of the music.

[0638] Step 5:

[0639] The server generates characters within the same context as the generated landscape. The generating AI determines the characters' actions and dialogue based on the music's theme and mood. The characters' dialogue scenarios are also constructed at this stage.

[0640] Step 6:

[0641] The server streams generated scenery and character data to the device, which then displays it on the screen. The synchronization of music and video timing provides a visually and aurally consistent experience.

[0642] Step 7:

[0643] Users provide feedback through the interface presented during their experience. This feedback is sent from the device to the server and used to improve the system and update user profiles.

[0644] Step 8:

[0645] After the experience ends, users can use customization options to adjust the scenery, character movements, dialogue, and other elements. This customization information is sent from the device to the server and reflected in the next experience.

[0646] (Example 1)

[0647] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0648] In modern society, there is a demand for more personalized and immersive music listening experiences. However, existing systems struggle to provide dynamic visual experiences synchronized with music, and they lack options for users to finely customize that experience. Furthermore, there is a challenge in the insufficient mechanism for incorporating user feedback and continuously improving the system.

[0649] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0650] In this invention, the server includes means for providing a user screen for the user to select the type of experience, means for receiving and analyzing audio data, means for utilizing an artificial intelligence model that generates a landscape based on the analysis results, and means for generating a virtual character that is synchronized with the audio and interacting with it. This provides the user with an immersive experience that links music and visuals, further allows the user to customize the experience, and enables continuous improvement of the system through feedback.

[0651] "User interface for selecting the type of experience" refers to the interface that allows users to select scenery and experience content that are linked to their desired music.

[0652] "Receiving and analyzing audio data" refers to the process of analyzing the characteristics of music data, such as rhythm and tempo.

[0653] An "artificial intelligence model for generating landscapes" refers to an algorithm that generates appropriate visuals in real time based on the results of analyzing music data.

[0654] A "virtual character linked to sound" refers to a digital character that is generated according to the characteristics and content of the music and is capable of interacting with the user.

[0655] "Engaging in interaction" refers to virtual characters providing users with music-related information or offering interactive experiences through dialogue.

[0656] "Synchronizing and providing generated video and audio to users" refers to seamlessly presenting real-time generated visuals and music to the user.

[0657] "Providing choices" means offering within the system options that allow users to customize their experience and adjust the scenery and dialogue content to their own preferences.

[0658] This invention is a system that allows users to enjoy a real-time fusion of music and scenery. The system consists of a user's terminal, a server, and a music distribution service.

[0659] First, the user launches a dedicated application on their device and selects the category of scenery they wish to experience. A user interface is provided for selection, allowing the user to easily determine their preferred type. This selection information is then transmitted to a server via the internet.

[0660] The server receives the selected category information and prepares to access the music streaming service. Music data is streamed through the terminal, and the server receives and analyzes the music data. The analysis utilizes audio analysis software to extract the rhythm, tempo, and other features of the music.

[0661] The server also uses a generative AI model to generate appropriate landscapes in real time based on the analysis results. The generated landscapes are linked to the characteristics of the music; for example, calm music will create a tranquil natural landscape, while dynamic music will create an impressive urban landscape in real time.

[0662] Furthermore, the server generates a virtual character synchronized with the audio, enabling two-way interaction with the user. This character provides a deeper sense of immersion by using dialogue based on the content and tone of the music. Through interaction with the character, users can enjoy an experience where music and scenery are integrated.

[0663] The video and music are synchronized on the device, and the generated content is delivered seamlessly to the user. Users can also use customization options to adjust the scenery and character interactions to their liking.

[0664] For example, when a user selects the "Natural Landscape" category, the device receives calming music from a music streaming service. The server generates a tranquil lake landscape corresponding to this music, displays a bird character, and provides fun facts about nature. In this way, the user can enjoy a unique audiovisual experience based on the prompt message, "Please select a tranquil landscape and set it so that a character speaks simple nature facts accompanied by music that evokes a morning atmosphere."

[0665] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0666] Step 1:

[0667] The user launches the application through their device and selects the category they wish to experience from the displayed user screen. This input information is sent to the server via the internet. The selection information is stored in a database and used in the server's next processing step. Specifically, the user selects the "Natural Landscapes" category, and this selection information is sent to the server.

[0668] Step 2:

[0669] The server prepares to communicate with the music streaming service based on the selected category information. The server sends a request to the terminal and starts streaming music data. The streamed music data is received by the server, and analysis of the music's rhythm, tempo, tone, etc., is performed. The input data is a music file, and the output is the analysis results showing the characteristics of the music. For example, data on a gentle rhythm and tone may be extracted.

[0670] Step 3:

[0671] The server uses the analysis results and leverages a generative AI model to generate landscapes that match the music in real time. In this process, the algorithm selects appropriate visual elements based on the characteristics of the analyzed music, dynamically constructing the landscape. The input is the characteristics of the music, and the output is a visually generated content in real time. For example, for calm music, landscapes of lakes and forests are generated.

[0672] Step 4:

[0673] The server then generates virtual characters based on musical characteristics and assigns them scripts. These characters initiate interaction with the user using dialogue synchronized with the music. Input is the music and user selections, while output is dialogue and actions for interaction. Specifically, the server generates a bird character and provides topics related to nature.

[0674] Step 5:

[0675] The terminal receives generated scenery and character information transmitted from the server and integrates and displays it on the screen. At this time, the video and music are synchronized, providing the user with a seamless experience. The input is the generated scenery and character data, and the output is integrated visual and auditory content for the user. For example, scenery and characters synchronized with gentle music are displayed on the screen.

[0676] Step 6:

[0677] User feedback is collected and sent to the server. The server stores this feedback in a database and uses it to improve future AI models. The input is user feedback data, and the output is training data aimed at improving the system. For example, based on the feedback, the AI ​​model is adjusted so that a more preferred landscape is generated in the next session.

[0678] (Application Example 1)

[0679] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0680] When users enjoy the synergy of music and visuals in real time, it is essential that the generation of scenery and characters is smoothly adapted to the characteristics of the music, thereby enhancing the quality of the viewing experience. Furthermore, incorporating user customization and feedback is a challenge in providing an experience that better suits individual needs.

[0681] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0682] In this invention, the server includes means for providing an operating unit for the user to select the type of experience, means for receiving and analyzing acoustic signals, and means for utilizing a generation AI to generate scenes based on the analysis results. This enables the user to experience immersive viewing experiences tailored to their individual needs by generating scenery and characters in harmony with music in real time.

[0683] A "user" refers to someone who uses and operates a system or application.

[0684] "Experience types" refer to the various themes and situations that users can choose from.

[0685] The term "operation unit" refers to the interface through which the user makes selections and gives instructions to the system.

[0686] "Acoustic signal" refers to sound signals that the system analyzes, such as music data and voice data.

[0687] "Analysis" refers to the process of analyzing the characteristics of acoustic signals and extracting guidelines for generating landscapes and characters.

[0688] "Scene" refers to the collective term for the images and visuals presented to the user.

[0689] "Generative AI" refers to artificial intelligence that uses machine learning techniques to automatically generate landscapes, characters, and other images.

[0690] A "moving character" refers to a character that appears on the screen in conjunction with an audio signal and interacts with the user through dialogue and actions.

[0691] The system implementing this invention uses a user terminal, a server, and a generative AI with music analysis capabilities. Processing begins when the user selects the type of experience they wish to experience via the control panel on the terminal. The terminal transmits this selection information to the server. The server receives audio signals via streaming from a music distribution service over the internet and analyzes those signals.

[0692] Based on the analysis results, the server uses a generative AI to generate corresponding scenes in real time. This generative AI has the ability to analyze characteristic data such as rhythm, tempo, and melody of the acoustic signal and design a scene that is appropriate for it. Furthermore, an action character is also generated according to the attributes of the acoustic signal and interacts with the user.

[0693] The generated scenes and sounds are streamed to the device in real time. The device integrates and displays this, allowing the user to enjoy a harmonious visual and auditory experience.

[0694] As a concrete example, if the user selects relaxing music, the AI ​​will present a tranquil forest landscape in sync with that music. Within that forest, a moving object in the shape of a small bird will appear, creating a scene that provides the user with interesting facts about nature.

[0695] As an example, the prompt given to the generation AI is, "Generate character movements that harmonize with a tranquil nighttime cityscape, set to jazz music." In this way, interaction between the AI ​​and the user experience is realized.

[0696] This system can continuously improve the user experience through user customization and feedback. Specifically, user customizations and feedback information are recorded on the server and used as training data for the generating AI. This enables the provision of a more personalized experience.

[0697] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0698] Step 1:

[0699] The user selects the type of experience they wish to have using the controls on their device. This input determines the theme of the experience and is sent to the server as data. At this stage, the user's selection information is collected on the server.

[0700] Step 2:

[0701] The server receives audio signals related to a specified experience via streaming from a music streaming service. This audio data is used as input data to extract the characteristics of the music. The server analyzes the audio signals and performs data processing to identify rhythm, tempo, pitch, and other characteristics.

[0702] Step 3:

[0703] The server utilizes a generative AI model based on the analysis results and generates a suitable scene using prompt text. This AI takes the analyzed musical characteristics as input data and generates a visual scene that matches the prompt as output. For example, a prompt text for calm music might be "Generate a calm forest landscape."

[0704] Step 4:

[0705] The server simultaneously generates characters based on the tone and content of the sound. A generation AI model is used to depict the characters' actions and dialogue. In this process, the volume and rhythm of the sound are taken as input, and the character's movements and speech are output.

[0706] Step 5:

[0707] The generated scenes and audio are streamed in real time from the server to the terminal, which then integrates and displays them to the user. The terminal receives the data stream from the server and outputs synchronized audio and video to provide the user with an experience.

[0708] Step 6:

[0709] After the experience ends, the user provides feedback, which the device sends to the server. This feedback information is recorded on the server as input and used as training data for the generated AI model. The server uses this data to perform calculations to improve the quality of future experiences.

[0710] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0711] This invention provides a system that allows users to enjoy a synergy based on real-time generated scenery and music, and in particular, by combining it with an emotion engine that recognizes the user's emotions, it offers a more personalized experience. This system consists of a user's terminal, a server, a music distribution service, and an emotion engine.

[0712] First, the user launches an application on their device and selects a category of scenery they want to experience. This information is sent from the device to the server, and the device's emotion engine analyzes the user's facial expressions and voice captured through the camera and microphone to recognize the user's current emotional state.

[0713] The server receives this selection and emotional information and retrieves music data from a music streaming service. The music data is transmitted in real time, and the server analyzes the rhythm, tempo, and other musical characteristics. The analyzed information is used by a generative AI to create a landscape in real time that is appropriate to the user's emotional state.

[0714] The generating AI adjusts the color tone, composition, and dynamic elements of the landscape based on the output of the emotion engine. For example, it can generate a calm landscape when the user is relaxed and a lively landscape when they are excited. Characters also evolve to display appropriate expressions and actions according to emotions, supporting natural conversations.

[0715] The generated video and music are streamed to the device. The device seamlessly integrates the video and music, presenting it as a new, optimized experience for the user. The user can enjoy this immersive experience and simultaneously provide feedback through the interface. This feedback is sent from the device to the server and used as data to improve future experiences.

[0716] As a concrete example, consider a scenario where a user selects "natural scenery" with the goal of "relaxing," and the emotion engine recognizes the user's calm state. In this case, the server generates scenery of a quiet forest or a calm beach and selects music with a soothing melody. The character speaks to the user in a friendly and gentle tone. In this way, the present invention flexibly provides experiences based on the user's emotions and choices, making it possible to enrich daily life in urban areas.

[0717] The following describes the processing flow.

[0718] Step 1:

[0719] The user launches the application on their device and selects the category they want to experience from the interface. This selection information is sent from the device to the server.

[0720] Step 2:

[0721] The device uses its built-in camera and microphone to capture the user's facial expressions and voice. An emotion engine analyzes this data to identify the user's emotional state. The identified emotion information is then sent to the server.

[0722] Step 3:

[0723] The device accesses a music streaming service and receives music data tailored to the user's preferences in real time. The received music data is then transferred to a server.

[0724] Step 4:

[0725] The server analyzes the music data and extracts features such as tempo, rhythm, and genre. These analysis results are then used as a set of data necessary for the landscape generation process.

[0726] Step 5:

[0727] The server combines the user's emotional state, recognized by the emotion engine, with the music analysis results, and inputs this data into the generative AI. The generative AI then generates an appropriate scene in real time.

[0728] Step 6:

[0729] The generating AI adjusts the color scheme and composition of the landscape to match the user's emotions, creating a landscape with dynamic elements. Simultaneously, the server generates a character synchronized with the music, ready to interact with the user.

[0730] Step 7:

[0731] The server streams landscape data and character dialogue scripts generated by the server to the terminal. The terminal displays these on its screen, ensuring that music and video are seamlessly integrated.

[0732] Step 8:

[0733] Users utilize a feature to provide feedback during the experience. The device sends the entered feedback to a server for recording. This feedback is used to improve future experiences.

[0734] Step 9:

[0735] After the user finishes the experience, they can adjust the scenery and dialogue using customization options. The device sends the customization information to the server, and it is reflected in the next experience.

[0736] (Example 2)

[0737] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0738] In modern society, customized entertainment experiences tailored to the emotional state of users are limited. In particular, there is a lack of means to create immersive experiences where music and visual content are closely integrated. Furthermore, there is a need for mechanisms that incorporate user feedback into the system to improve individual experiences.

[0739] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0740] In this invention, the server includes means for providing a display device for the user to select the type of experience, means for acquiring and analyzing sound information, and means for utilizing a generative model that generates a landscape based on the analysis results and the user's emotional information. This makes it possible for music and visual content to be synchronized according to the user's emotions, providing a personalized and immersive experience.

[0741] "User" refers to an individual or group that operates the system and selects the type of experience.

[0742] "Experience type" refers to the category or theme of the content that the user wants to experience.

[0743] A "display device" refers to a device or software that provides an interface used by users to select the type of experience they are experiencing.

[0744] "Audio information" refers to data that a system processes, including music and other sound data.

[0745] "Analysis" refers to the process of processing and calculating acquired sound and emotional information.

[0746] "Emotional information" refers to the emotional state estimated from facial expressions and voices acquired through the user's camera and microphone.

[0747] A "generative model" refers to a system that uses artificial intelligence to generate new landscapes and content based on analysis results and emotional information.

[0748] A "virtual character" is a generated character that interacts with sound information and enables dialogue with the user.

[0749] "Visual information" refers to generated video data such as landscapes, which is presented to the user in conjunction with audio information.

[0750] This invention is a system that enables users to enjoy a customized entertainment experience tailored to their emotions in real time. This system mainly consists of three main elements: a terminal, a server, and an audio service.

[0751] First, users can select their preferred content category via a display on their device that allows them to choose the type of experience. This selection information is then sent from the device to the server. The device uses its built-in emotion engine to analyze emotional information from the user's facial expressions and voice captured through the camera and microphone. This information indicates the user's current emotions and is used to customize the content provided by the system.

[0752] The server uses the selection and emotion information received from the user to retrieve sound information from the audio service. The retrieved sound information is then analyzed on the server, and features such as rhythm and tempo are identified. Based on these features and emotion information, the server generates prompts and inputs them into a generative AI model. This model generates visual information such as landscapes and characters based on the specified information. For example, if the user selects "I want to relax," the generative AI model will depict calm natural scenery and friendly characters.

[0753] The generated visual information and acquired audio information are transmitted from the server to the terminal in real time. The terminal integrates these and presents them seamlessly to the user, providing an immersive experience. For example, if the user selects "I want to relax" and the system recognizes that the user is in a calm state based on emotional information, calming music and visual natural scenery appropriate to that situation will be presented.

[0754] An example of a prompt message would be: "The user has selected the category 'Natural Landscapes' and their emotional state is 'Calm.' Based on this, please generate a relaxing landscape and corresponding music." In this way, the system can provide an entertainment experience that matches the user's individual emotions.

[0755] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0756] Step 1:

[0757] The user launches the application on their device and selects the type of content they want to experience. This input information is saved on the device and then transferred to a server via the network. Users can choose categories such as "natural landscapes" or "urban scenery."

[0758] Step 2:

[0759] The device operates a built-in emotion engine and uses the user's camera and microphone to collect facial expressions and voice. This collected data is then analyzed by an emotion analysis algorithm to identify the user's emotional state (e.g., "relaxed," "excited," etc.). The analysis output is sent to the server as the user's emotion information.

[0760] Step 3:

[0761] The server receives selection and emotion information from the user. Based on this information, the server connects to an audio service and obtains appropriate sound information. Here, the server requests the optimal music data according to the user's selection and emotion, and the server analyzes the characteristics of the obtained music, such as rhythm, tempo, and melody. The output of this analysis becomes part of the generation prompt.

[0762] Step 4:

[0763] The server generates prompts for visual information based on the analyzed sound and emotion information. These prompts include information about the sound rhythm, the user's emotional state, and the landscape according to the selected category. For example, a specific prompt statement such as "The user's selected category is 'Natural Landscape', and their emotional state is 'Calm'" is passed to the generating AI model.

[0764] Step 5:

[0765] The generation AI model begins generating landscapes and characters based on prompts received from the server. The model adjusts color tones and behavioral elements based on emotional information, creating, for example, a forest landscape with gentle colors or a character with a friendly expression. The generated video data is then returned to the server.

[0766] Step 6:

[0767] The server streams the generated visual information and analyzed audio information to the device. The server transfers data in real time, and the device adjusts to seamlessly present the experience to the user. Through this integrated view, the user can experience an immersive experience.

[0768] (Application Example 2)

[0769] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0770] Analyzing users' emotions in real time and providing a personalized experience that integrates music and visuals has been difficult with conventional technology. Rapidly generating content suited to the user's emotions and creating an immersive experience through interaction is a challenge that needs to be addressed technologically.

[0771] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0772] In this invention, the server includes means for analyzing the user's emotional information to personalize the experience, means for analyzing and generating music and video information, and means for presenting the generated content to the user in real time. This enables the user to quickly enjoy a personalized experience tailored to their emotions.

[0773] "User" refers to a person or entity that uses the system.

[0774] "User interface" refers to the screens and input devices used by users to interact with a system.

[0775] "Music information" refers to audio data and related metadata, which constitute the audio content that the system analyzes.

[0776] "Data analysis" refers to the process of processing received information based on algorithms to obtain the intended result.

[0777] "Generative AI" refers to artificial intelligence technology that automatically generates new content based on received data.

[0778] "Visual information" refers to visual elements such as images and screen displays presented to the user.

[0779] "Emotional information" is data that represents the user's current emotional state and is acquired through facial expressions and voice.

[0780] "Personalization" refers to optimizing content and services to suit the individual user's preferences and circumstances.

[0781] The system for realizing this invention analyzes the user's emotional information and provides personalized visual information and music to the user using generative AI. The system consists of a user interface, an emotional analysis engine, generative AI, and servers and terminals responsible for data synchronization.

[0782] The server performs data analysis to integrate music and visual information based on the user's selected experience topic and emotional information. For this analysis, a camera and microphone for acquiring emotional data must be installed on the terminal. Software such as "Affectiva" or "Microsoft Azure Emotion API" are used for analyzing emotional information. Generative AI such as "OpenAI's GPT-4" or "Midjourney" are used to generate visual information, and music information is acquired and analyzed in real time from music streaming services.

[0783] In this invention, the terminal synchronizes generated visual information and music transmitted from the server, presenting it to the user as an immersive experience. The terminal also receives feedback from the user, transmits it to the server, and stores it as data for future experience improvements.

[0784] As a concrete example, if a user who has finished a stressful day selects "relaxation" as the theme in the application, the generating AI will provide a tranquil forest scene and soothing piano music. In this case, an example of a prompt message would be: "The user's emotions are calm, and we want to generate a relaxing scene. Specifically, we will provide a combination of a quiet forest scene and gentle piano music."

[0785] This system allows users to instantly enjoy the experience that best suits their emotions at that moment.

[0786] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0787] Step 1:

[0788] The user launches the application and selects a theme to determine the subject they wish to experience. The user's theme selection information is provided as input and sent to the system via the user interface. The selected theme is then sent to the server as output.

[0789] Step 2:

[0790] The device uses a camera and microphone to record facial expressions and voice in order to acquire user emotional information. The input includes real-time video and audio of the user. An emotion engine analyzes this data and outputs the user's current emotional state. This emotional information is sent to a server.

[0791] Step 3:

[0792] Based on the subject selection information and sentiment information received by the server, music information is retrieved from a music streaming service. Subject and sentiment information are used as input, and the corresponding music tracks are output. This music is analyzed by the server, and its characteristics are extracted.

[0793] Step 4:

[0794] The server generates visual information using a generative AI model based on the analyzed musical characteristics and emotional information. The input for this step is musical characteristics and emotional information, and the output is a video tailored to the user. The generative AI constructs the specific video using prompts.

[0795] Step 5:

[0796] Visual information and music generated from the server are streamed to the terminal. The generated data is used as input, and the terminal receives and synchronizes it to present it to the user. The output is a harmonious presentation of visuals and music experienced by the user.

[0797] Step 6:

[0798] Users provide feedback on the experience they are given. As input, the user's feedback information is recorded on the device. As output, this data is sent to a server and stored as learning data for improving future experiences.

[0799] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0800] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include those described above. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions shown by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0801] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0802] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0803] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0804] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0805] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0806] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0807] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0808] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0809] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0810] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0811] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0812] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0813] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0814] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0815] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0816] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0817] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0818] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0819] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted as being incorporated by reference.

[0820] The following is further disclosed regarding the embodiments described above.

[0821] (Claim 1)

[0822] A means of providing an interface for users to select an experience category,

[0823] A means of receiving and analyzing music data,

[0824] A method that utilizes a generative AI to generate landscapes based on analysis results,

[0825] A means of generating and interacting with characters synchronized with music,

[0826] A system that includes means for presenting generated video and music to the user in sync.

[0827] (Claim 2)

[0828] The system according to claim 1, wherein the user provides feedback, which is then stored as training data.

[0829] (Claim 3)

[0830] The system according to claim 1, which provides the user with the option to customize the scenery and dialogue settings.

[0831] "Example 1"

[0832] (Claim 1)

[0833] A means of providing a user screen for users to select the type of experience,

[0834] A means of receiving and analyzing audio data,

[0835] A means of using an artificial intelligence model that generates landscapes based on analysis results,

[0836] A means of generating and interacting with virtual characters that are synchronized with voice,

[0837] A system that includes means for synchronizing and providing generated video and audio to the user.

[0838] (Claim 2)

[0839] The system according to claim 1, wherein a user provides evaluation information, which is stored as information for learning.

[0840] (Claim 3)

[0841] The system according to claim 1, which provides users with options to adjust landscape and interaction settings.

[0842] "Application Example 1"

[0843] (Claim 1)

[0844] A means of providing a control panel for the user to select the type of experience,

[0845] A means of receiving and analyzing acoustic signals,

[0846] A method that utilizes a generative AI to generate scenes based on analysis results,

[0847] A means of generating and interacting with an object that moves in conjunction with sound,

[0848] A means of harmonizing the generated video and audio and displaying them to the user,

[0849] A means for generating a landscape and making moving objects appear in real time based on the sound selected by the user,

[0850] A system that includes this.

[0851] (Claim 2)

[0852] The system according to claim 1, wherein the user provides feedback after use, and this is recorded as learning information.

[0853] (Claim 3)

[0854] The system according to claim 1, which provides the user with options to adjust the setting of the scene and conversation.

[0855] "Example 2 of combining an emotion engine"

[0856] (Claim 1)

[0857] A means for providing a display device for users to select the type of experience,

[0858] A means of acquiring and analyzing sound information,

[0859] A means of utilizing a generative model that generates landscapes based on analysis results and user sentiment information,

[0860] A means of generating a virtual character that interacts with sound information and engaging in conversation,

[0861] A system that includes means for synchronizing and presenting generated visual and auditory information to the user.

[0862] (Claim 2)

[0863] The system according to claim 1, wherein the user performs an evaluation and stores this as training information.

[0864] (Claim 3)

[0865] The system according to claim 1, which provides the option for the user to individually configure the scenery and conversation settings.

[0866] "Application example 2 when combining with an emotional engine"

[0867] (Claim 1)

[0868] A means of providing a user interface for users to select the subject they want to experience,

[0869] A means of receiving music information and performing data analysis,

[0870] A means of using a generative AI that generates images based on analyzed information,

[0871] A means of creating characters linked to music and conducting dialogue,

[0872] A means of presenting generated visual information and music to the user in sync,

[0873] A system that includes means to analyze emotional information and personalize the experience.

[0874] (Claim 2)

[0875] The system according to claim 1, wherein users provide feedback, which is then saved and utilized as learning data.

[0876] (Claim 3)

[0877] The system according to claim 1, which provides users with personalized options for visual information and interaction settings. [Explanation of Symbols]

[0878] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. A means of providing an interface for users to select an experience category, A means of receiving and analyzing music data, A method that utilizes a generative AI to generate landscapes based on analysis results, A means of generating and interacting with characters synchronized with music, A system that includes means for presenting generated video and music to the user in sync.

2. The system according to claim 1, wherein the user provides feedback, which is then saved as training data.

3. The system according to claim 1, which provides the user with the option to customize the scenery and dialogue settings.