system
The system addresses the challenge of creating emotionally rich, customized music by performing sentiment analysis and using generative AI to generate music, record professional performances, and manage rights, ensuring users can easily create and share music for special occasions.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-10-11
- Publication Date
- 2026-04-23
AI Technical Summary
Users lack the musical skills and techniques to create emotionally rich and individually customized music for special occasions, and existing systems fail to manage rights and revenue sharing effectively for generated audio content.
A system that performs sentiment analysis on user input, generates lyrics and melodies using generative AI, records professional performances, manages rights, and facilitates revenue sharing, all while allowing users to easily create and share emotionally rich music.
Enables users to generate and experience emotionally rich, customized music without specialized knowledge, while managing rights and revenue, enhancing the emotional impact of special events.
Smart Images

Figure 2026069021000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] On special anniversaries or events, there is an increasing need to create emotionally rich and individually customized music. However, in general, users often lack sufficient musical skills and techniques to create such music. Therefore, there is a demand for a system that can easily generate and experience original music reflecting individual emotions. This invention aims to solve the problem of providing users with emotionally rich music on special occasions through emotion analysis technology, generative AI, and the cooperation of professional musicians.
Means for Solving the Problems
[0005] The present invention provides a means for performing sentiment analysis on data received via a user interface as an information input means, and a generation means for generating lyrics and melodies based on the analysis results. The system also includes a performance means for generating performance data based on the generated lyrics and melodies, and further comprises a means for recording this performance data, generating it as an audio file, and providing it to the user. Furthermore, the present invention has a function for managing rights related to the secondary use of the audio file and for revenue sharing, and by having a function for extracting keywords and phrases from the data using natural language processing as the sentiment analysis means, it effectively solves specific problems.
[0006] An "information input means" is an interface for receiving data from users, including event information and emotions.
[0007] "Methods for analyzing emotions" refer to functions that analyze received data and extract the emotions and themes contained within it.
[0008] A "generative means" is a device that has the function of creating corresponding lyrics and melodies based on analyzed emotions and themes.
[0009] "Performance means" refers to a process or apparatus for generating performance data based on the generated lyrics and melody.
[0010] An "audio file" is a music file that contains recorded performance data saved in digital format.
[0011] "Means for managing rights related to secondary use" refers to functions that manage rights related to the commercial use and licensing of sound source files.
[0012] "Means of revenue distribution" refers to a function for appropriately distributing the revenue obtained from the use of sound files to relevant parties.
[0013] "Natural language processing" is a technology that extracts keywords and phrases from text data and analyzes emotions and themes. [Brief explanation of the drawing]
[0014] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] This is a sequence diagram showing the processing flow of the data processing system in Example 2, which incorporates an emotion engine. [Figure 14]A sequence diagram showing the processing flow of a data processing system in Application Example 2 when combined with an emotion engine. [[ID=application mode for implementing the invention]]
[0015] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0016] First, the terms used in the following description will be explained.
[0017] In the following embodiments, a labeled processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), APU (Accelerated Processing Unit), etc.
[0018] In the following embodiments, a labeled RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0019] In the following embodiments, a labeled storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disk (e.g., hard disk), or magnetic tape, etc.
[0020] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0021] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0022] [First Embodiment]
[0023] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0024] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0025] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0026] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0027] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0028] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0029] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0030] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0031] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0032] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0033] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0034] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0035] This invention is a system that allows users to easily generate and experience emotionally rich custom music for special anniversaries and events. The system receives event information and emotional data provided by the user and generates music based on this data. Specifically, it is implemented as follows:
[0036] First, users input event information and the emotions they wish to convey through a dedicated application or web platform. This information is then processed on their device and sent to the server. This data can include a wide range of information, such as the type of event, participant information, and desired emotions or themes.
[0037] Next, the terminal sends the received information to the server, which then passes that information to the sentiment analysis module. This module uses natural language processing to extract important emotions and related themes from the data. Based on these results, a generative AI module automatically generates the lyrics and melody of the song.
[0038] The generated lyrics and melody are sent from the server to professional musicians. The musicians perform based on these and record the result. The recorded audio data is uploaded to the server as a music file.
[0039] Ultimately, the server sends the audio file to the user's device. The user can download this file and play it during a special event. This system is easy for users to use, yet technically, it utilizes advanced emotion analysis and generative AI to provide individually tailored, emotionally rich music.
[0040] For example, when generating a song for a wedding anniversary, the user inputs memories from their married life and feelings of gratitude they want to convey. The resulting song will express love and gratitude, making the wedding anniversary a more emotionally moving experience.
[0041] This system also includes functions for rights management and revenue sharing related to the secondary use of sound files, and can support procedures for commercial use. This allows users to use music for various purposes with peace of mind.
[0042] The following describes the processing flow.
[0043] Step 1:
[0044] Users input event information and the emotions they wish to convey using a dedicated application or web platform. This information includes the date, location, participants, and the event's theme. This prepares the system to provide the necessary information.
[0045] Step 2:
[0046] The terminal validates the user's input data, checking for any errors or missing information. Once validation is complete, it converts the information into a specified format and sends it to the server. This conversion is crucial for efficient data processing.
[0047] Step 3:
[0048] The server stores the received data and transfers it to the sentiment analysis module. The sentiment analysis module uses natural language processing techniques to extract key emotions and related themes from the input information. This prepares the data for the next processing step.
[0049] Step 4:
[0050] The server calls a generation AI module based on the results of emotion analysis to generate lyrics and melody. This process uses an algorithm designed to reflect the emotions and themes entered by the user. The generated song will be in line with the user's desires.
[0051] Step 5:
[0052] The server sends the generated lyrics and melody to professional musicians. The musicians record the song using instruments and vocals based on the instructions. This recording is a crucial step towards the final version of the song.
[0053] Step 6:
[0054] Musicians upload their completed performance data to a server. The server receives this data and, if necessary, adjusts the sound quality or converts it to the correct format. This ensures that the music meets the user's expectations in terms of quality.
[0055] Step 7:
[0056] The server sends the final audio file to the user's device. The user downloads the provided music and is ready to play it at the event. Through this process, the user can obtain customized music content.
[0057] (Example 1)
[0058] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0059] In modern times, people desire to easily create and experience music that richly expresses their emotions for special events and anniversaries. However, creating music requires specialized knowledge and skills, making it difficult for ordinary people to easily create satisfying music. Furthermore, the management of rights regarding the secondary use of created sound sources is often unclear, creating problems that prevent people from using them with peace of mind.
[0060] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0061] In this invention, the server includes means for performing sentiment analysis on data received via a user interface acting as an information input device, means for securely transmitting data from a user terminal to the server via digital data communication, and means for generating a download link for the generated sound source data and transmitting it to the user terminal. This makes it possible for ordinary users to securely generate and use emotionally rich music tailored to their individual needs without requiring specialized knowledge.
[0062] An "information input device" is a device that provides an interface for users to input data, and can receive input such as emotions and event information.
[0063] "Sentiment analysis" is a process that uses natural language processing techniques to extract emotions and related themes from input text data.
[0064] A "generation method" refers to a device or program for generating lyrics and melodies based on the results of emotion analysis.
[0065] "Performance control means" refers to a device or program for generating performance data based on the generated lyrics and melody, and for controlling that performance.
[0066] "Audio data" refers to digital music files of recorded performances, and is the final music content provided to the user.
[0067] A "download link" is an access link that sends the generated audio data to the user's device, allowing the user to receive and download the audio data.
[0068] "Digital data communication" refers to a communication method used to transmit data between a user terminal and a server, and is carried out using protocols that ensure security and confidentiality.
[0069] This invention is a system that allows users to easily generate and experience emotionally rich custom music for special anniversaries and events. Since this system creates music based on event information and emotionally related data provided by the user, a specific embodiment of this system is described below.
[0070] Users input information from their PCs, smartphones, or other devices using a dedicated application or web platform. This information includes the type and details of the event the user desires, as well as the emotions or themes they wish to convey. This input information is transmitted to the server as encrypted digital data.
[0071] The server uses an emotion analysis module to process the received data. This module incorporates natural language processing technology to extract emotions and related themes from the data. Next, the server uses a generative AI model to automatically generate lyrics and melodies based on the extracted emotion data. At this time, it is given a specific prompt message such as, "Generate an upbeat melody that expresses happiness, and lyrics that convey gratitude."
[0072] The generated lyrics and melody are sent from the server to professional musicians, who then perform the song based on them. The performance is recorded and saved as audio data. The recorded audio data is finally stored on the server and provided to the user's device as a download link.
[0073] For example, if a user requests a song themed around emotions such as "love" or "memories" for their wedding anniversary, the system will generate a customized song based on the user's detailed memories and emotions, which can then be used to enhance the special anniversary celebration. This system provides a way for users to easily generate emotionally rich songs without requiring specialized knowledge, ensuring a worry-free experience.
[0074] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0075] Step 1:
[0076] Users input event information and desired emotions using a dedicated application or web platform. This input includes event type, participant information, and specific themes or keywords. This information is formatted as digital data on the user's device and sent to the server.
[0077] Step 2:
[0078] The terminal sends user input information to the server using a secure communication protocol, and the server receives it. The entered data is encrypted, ensuring security and privacy.
[0079] Step 3:
[0080] The server passes the received data to the sentiment analysis module. The module uses natural language processing to extract important emotions and themes from the provided data. This analysis converts the extracted sentiment data into a new digital data format.
[0081] Step 4:
[0082] Based on the results of the emotion analysis, the server uses a generative AI model to automatically generate lyrics and melodies. During this process, prompts such as "Please generate a gentle melody and lyrics that convey love and gratitude" are input to the AI. The AI then outputs song data that matches the analysis results, following these prompts.
[0083] Step 5:
[0084] The server sends the generated lyrics and melody to professional musicians. The musicians perform the song based on the provided information and record the performance in a studio in high quality. After the recorded data is edited, it is converted into a digital file as audio data.
[0085] Step 6:
[0086] The server uploads the audio data received from the musicians to its storage system and provides a download link to the user's device. The user then obtains the audio data via this link and is ready to play it at the event as needed.
[0087] (Application Example 1)
[0088] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0089] Traditionally, generating emotionally rich custom music required specialized knowledge and skills, making it difficult for the average user to access. Furthermore, there were few easy ways to share the generated music with others, limiting its use to special occasions and events. Additionally, insufficient rights management regarding the secondary use of the generated audio data presented challenges for commercial use. There is a need for a system that solves these problems and allows users to easily generate and share individually tailored music.
[0090] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0091] In this invention, the server includes means for emotionally analyzing data received via a user interface as an information input means; means for generating lyrics and melodies based on the emotional analysis; means for generating performance data based on the lyrics and melody; means for recording the performance data and generating it as an acoustic signal; means for providing the acoustic signal to a mobile terminal; means for sharing the acoustic signal using a communication function; and control for managing rights related to the secondary use of the sound source data and for revenue sharing. As a result, users can generate emotionally rich music without requiring special skills, easily share it with others, and manage the commercial use of the generated music.
[0092] An "information input method" is an interface for users to input information about special anniversaries or events.
[0093] A "means for analyzing emotions" refers to a mechanism that uses natural language processing technology to analyze emotions and themes based on information input by the user.
[0094] "Generation method" refers to technology for automatically generating lyrics and melodies based on the results of emotion analysis.
[0095] "Performance method" refers to a system for creating performance data based on generated lyrics and melodies.
[0096] "Means of generating as an acoustic signal" refers to the process of recording performance data and outputting it as a music file.
[0097] "Means of provision" refers to a method for transmitting the generated acoustic signal to the user's terminal.
[0098] "Means of sharing" refers to a system that allows users to easily share audio signals with others using communication functions.
[0099] "Control for managing rights and distributing revenue" refers to a management function that appropriately manages the rights related to the use of sound source data and distributes revenue in commercial use.
[0100] The system implementing this invention allows users to generate original music tailored to special anniversaries and events, providing an emotionally rich experience. Users input event information and emotional data via an application on their smartphone or other device. The input data is transmitted from the device to a server. The server analyzes the emotions and themes in the input data using an emotion analysis means employing natural language processing technology.
[0101] Based on these analysis results, a generation AI model on the server automatically generates lyrics and melody. This generated music data is converted into performance data, which is then performed by professional musicians and recorded as an audio signal. The recorded audio signal is provided to the user's terminal by the server, and the user can receive, play, and share it.
[0102] Furthermore, the server manages the rights and controls revenue sharing related to the secondary use of audio signals, supporting the commercial use of the generated music. A key feature of this system is that users do not need any special technical knowledge to use it.
[0103] As a concrete example, consider a scenario where a user requests the creation of a song expressing gratitude for Mother's Day. The user inputs a message or feelings they want to convey to their mother, and the server generates a song expressing gratitude, such as "Thank you for everything." The generated song can then be played at a Mother's Day event where the family gathers, creating a special and moving experience.
[0104] As an example of a prompt in a generative AI model, the following text is used: "Today is Mother's Day. Last year I gave my mother her favorite flowers. This year I would like to give her a song to express my gratitude, so please generate a gentle and moving melody and lyrics. It would be desirable for the content to convey gratitude and happiness."
[0105] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0106] Step 1:
[0107] Users input event information and emotional data via a smartphone or device application. This input data includes details about specific anniversaries or events, as well as the emotions or themes they wish to convey. This data is then transmitted to the server via the user interface.
[0108] Step 2:
[0109] The server passes the received user data through a sentiment analysis system. This analysis system uses natural language processing techniques to extract keywords and phrases from the data and identify emotions and themes. The input is user data, and the output is the analyzed sentiment information.
[0110] Step 3:
[0111] The server uses a generative AI model to generate lyrics and melodies based on the acquired emotional information. In this process, the AI uses patterns learned from a vast amount of past data to create lyrics and melodies appropriate to the input emotion. The output is the generated song data.
[0112] Step 4:
[0113] Based on the generated music data, the server creates performance data. This performance data is used to give instructions to professional musicians to actually perform the music.
[0114] Step 5:
[0115] Professional musicians perform a song according to performance data, and the resulting audio signals are recorded. The recorded audio signals are then saved to a server as new music files.
[0116] Step 6:
[0117] The server provides the generated audio signal to the user's terminal. This music file can be downloaded by the user for use or playback at specific events.
[0118] Step 7:
[0119] Users play the generated audio signal on their device and share it with others as needed. Sharing can also include surveys and messages, making it a way to express gratitude to friends and family.
[0120] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0121] This invention relates to a system for generating and providing emotionally rich custom music for special anniversaries and events. A key feature of this system is its ability to recognize the user's emotions in real time by incorporating an emotion engine, and to utilize this information in the music generation process. The invention is embodied as follows:
[0122] First, the user inputs information about the event type, date and time, and the emotions and themes they wish to convey through a dedicated application or web platform. The user interface also accepts voice and facial expression input, which is used for analysis. In addition, an emotion engine for real-time emotion recognition begins to operate.
[0123] The terminal sends the input information and the analysis results from the emotion engine to the server. The emotion engine identifies emotions from voice and facial expression data and incorporates this into the system's overall data analysis process. As a result, the user's current emotional state is taken into consideration as part of the music generation process.
[0124] Next, the server uses an emotion analysis module to integrate the data sent from the user with the output of the emotion engine, extracting key emotions and themes from the text. The supplementary data from the emotion engine plays a particularly important role in reflecting subtle emotions and nuances in greater detail. As a result, the generated lyrics and melodies are more empathetic to the user's actual emotions.
[0125] The generated song information is sent from the server to professional musicians. The musicians perform and record the song in accordance with the emotions and themes. This recording data is returned to the server and saved as an audio file. The server also checks the quality of the audio file and makes adjustments as needed.
[0126] Ultimately, the server provides the completed audio file to the user's device. The user can download this audio and use it at their desired event. The system also facilitates rights management for secondary use of the audio and manages revenue sharing.
[0127] For example, if a user requests music to be played at a birthday party, their emotions specific to that day (e.g., surprise, joy) are recorded in real time through an emotion engine, and these emotions are reflected in the generated music. In this way, the present invention realizes a system that enriches individual user experiences and provides music closely related to emotions.
[0128] The following describes the processing flow.
[0129] Step 1:
[0130] Users access a dedicated application or web platform and enter event details. This includes the type of event, date and time, and participants, as well as the emotions or themes they want to reflect in the music, and relevant background information. In addition, users provide real-time audio or facial expression data to the system via webcam and microphone.
[0131] Step 2:
[0132] The device verifies the information received from the user and passes it to the emotion engine for analysis of audio and image data. The emotion engine uses audio and visual analysis algorithms to identify and quantify the current emotional state (e.g., joy, surprise, sadness). This information serves as the basis for the system to generate emotion-based music.
[0133] Step 3:
[0134] The terminal sends user data and sentiment analysis results from the sentiment engine to the server. The server receives and integrates this information and performs a detailed analysis through the sentiment analysis module. This analysis establishes key emotions and themes based on the user's input data and real-time recognized sentiment data.
[0135] Step 4:
[0136] The server uses the analysis results described above to run a generation AI module, generating lyrics and melodies that match the user's specific emotions. The AI takes into account the user's experiences and desired emotions, automatically creating a personalized song base. This process is supported by an algorithm that includes numerous parameters.
[0137] Step 5:
[0138] The server sends the generated lyrics and melody to professional musicians. The musicians perform the song based on the specified emotions and themes, and record the performance. High sound quality is ensured by using professional equipment. The performed song is recorded as the final audio data.
[0139] Step 6:
[0140] Musicians upload their completed recordings to a server. The server receives the data and verifies the sound quality and format. If necessary, the engine automatically adds effects and optimizes the quality before generating the final audio file.
[0141] Step 7:
[0142] The server sends the completed audio file to the user's terminal. The user can download this audio file and play it at the scheduled event. The system also provides the user with options for secondary use of the audio file and revenue sharing, and handles appropriate rights management if commercial use is desired.
[0143] (Example 2)
[0144] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0145] For special anniversaries and events, there is a demand for custom music that deeply resonates with the emotions of users. Conventional systems have struggled to reflect users' actual emotional states in real time and generate music based on those emotions. Therefore, there is a need to develop technology that accurately recognizes users' emotions and quickly generates music that is appropriate for them.
[0146] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0147] In this invention, the server includes a function to analyze emotions from data acquired through a user interface acting as an information receiving device, a function to recognize the user's emotional state by analyzing voice and facial expressions in real time, and an integration function to integrate the user's input data and the output from the emotion engine. This makes it possible to generate custom music that aligns with the user's emotions in real time.
[0148] "Information receiving device" is a general term for systems and devices used to acquire data through a user interface.
[0149] "Sentimental analysis" refers to the process of analyzing acquired data to identify the emotional state of users.
[0150] "Creation function" refers to a system that has the ability to generate text and music based on sentiment surveys.
[0151] The "performance function" refers to the process of creating musical information based on generated text and music.
[0152] An "audio file" refers to a file that stores recorded audio data in digital format.
[0153] A "music professional" refers to an individual or group that possesses the skills to perform and record music based on generated musical information.
[0154] "Language processing technology" is a general term for technologies that analyze human language and process it using computers.
[0155] This invention relates to a system that generates and provides custom music for special anniversaries and events based on the user's emotions. The invention utilizes dedicated hardware and software. Specifically, it employs an application or web platform that acts as a user interface for the user to input their emotions, and an emotion engine that analyzes emotions in real time.
[0156] Users input event information and emotions through applications or web platforms. This information may also include voice input and facial expression data, which are used to more accurately understand the user's emotional state.
[0157] The device analyzes data received from the user and uses an emotion engine to recognize the emotional state in real time. This analysis information is then sent to the server. Voice analysis and facial recognition technologies are used for emotion recognition, enabling highly accurate emotion analysis.
[0158] The server generates music using a generative AI model based on the received data. During the generation process, natural language processing techniques are used to extract emotional and thematic keywords from the text. The generated music information is then sent to music professionals for actual performance and recording.
[0159] For example, if a user requests music to be played at a birthday party, they might enter a prompt such as, "Please generate a custom song for my birthday party. It should be fast-paced and have themes of joy and surprise." This allows the system to generate an original song in real time that reflects the user's emotions.
[0160] Ultimately, the generated music files are provided from the server to the user's terminal, allowing the user to utilize them at events. The system also manages the rights and revenue sharing for these music files, supporting secondary use of the music.
[0161] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0162] Step 1:
[0163] Users input the type of special anniversary or event, date and time, and the emotions and themes they wish to convey, using a dedicated application or web platform. Voice and facial expression input are also available, allowing users to express their emotions more accurately. The input in this step, including event information and voice and facial expression data, serves as the initial data for emotion analysis.
[0164] Step 2:
[0165] The terminal begins processing the input data received from the user. Using voice analysis software and facial recognition technology, it analyzes the user's emotions in real time. The input voice and image data are converted into numerical data to identify the user's emotions. This process outputs the analyzed emotion data, which is then sent to the server.
[0166] Step 3:
[0167] The server receives emotion data and theme information sent from the terminal as input. First, it activates the emotion analysis module and uses natural language processing techniques to extract keywords and major emotions from the text. This process generates structured data that matches the user's theme and emotions. The output consists of prompt text and emotion theme data necessary for music generation.
[0168] Step 4:
[0169] The server takes structured emotional theme data as input and generates music data using a generative AI model. During this generation process, the lyrics and melody of a custom song are designed based on extracted keywords and emotional states. The output is digital data containing music information, which is then sent to music professionals.
[0170] Step 5:
[0171] The server provides the generated music information to a music professional. The music professional performs and records the music according to this information. The music information is received as input, and an audio file is sent back to the server as output. This recorded data is saved as a high-quality audio file.
[0172] Step 6:
[0173] The server verifies the final audio file and prepares it for delivery to the user's terminal. If sound quality adjustments are needed, these are made possible, and the file becomes easily accessible to the user afterward. The user downloads the completed audio file and makes it available for use in a specific event. The output of this step is the completed music file that the user receives.
[0174] (Application Example 2)
[0175] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0176] Traditional in-store customer experiences are uniform, making it difficult to provide personalized experiences tailored to the emotional state of individual customers. As a result, there were limitations to improving customer purchasing intent and satisfaction. This invention solves the problem of providing an optimal atmosphere tailored to each individual customer by analyzing the customer's current emotional state in real time and dynamically changing the music in the store based on that analysis.
[0177] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0178] In this invention, the server includes means for performing sentiment analysis on data received via a user interface as an information input means, means for analyzing the customer's emotional state and providing music data corresponding to that state, and means for transmitting the music data to an audio playback device and playing the music in real time. This makes it possible to appropriately adjust the atmosphere in the store according to the customer's emotional state and provide an optimal shopping environment for each individual customer.
[0179] An "information input means" is an interface for receiving and processing data from a user.
[0180] "Emotion analysis means" refers to technology that identifies a user's emotional state based on received data.
[0181] The "generation method" refers to the function that creates appropriate lyrics and melodies based on the results of emotion analysis.
[0182] "Performance method" refers to a system that creates specific performance data based on the generated lyrics and melody.
[0183] "Means of generating as an audio file" refers to a function for recording performance data and saving it in a file format.
[0184] "Means of providing to users" refers to the means of distributing the generated audio files in a format that users can use.
[0185] "Means of providing music data" refers to the function of selecting and preparing music data that matches the analyzed emotional state.
[0186] "A means of transmitting to an audio playback device and playing music in real time" refers to a technology for transmitting prepared music data to a playback device and playing music immediately.
[0187] The system for implementing this invention utilizes specific hardware and software to provide music based on the emotional state of customers in physical stores.
[0188] The server first receives data transmitted from the terminal and analyzes the customer's current emotional state using emotion analysis tools. This analysis utilizes OpenCV for facial recognition and Google® Cloud Speech-to-Text for speech analysis. The analyzed data is then used with a generative AI model to clearly determine the emotional state.
[0189] Subsequently, the server generates appropriate music data to be played in the store based on the generated emotional information. The music data is selected and adjusted using an emotional prediction model based on TENSORFLOW®. The selected music data is transmitted to the audio playback device and played in the store in real time.
[0190] This system makes it possible to provide an optimized music experience for each individual customer in physical stores, and to appropriately adjust the store's atmosphere to match the customer's emotions.
[0191] For example, playing relaxing jazz music when customers visit during busy weekday evenings can create a more comfortable shopping environment.
[0192] An example of a prompt message would be, "Calculate the customer's current emotional state based on their facial expression analysis and voice tone, and then select and play background music that reflects that state." This message specifically instructs the system to perform certain actions.
[0193] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0194] Step 1:
[0195] The device collects the user's facial expressions and voice using its camera and microphone. This provides raw data to understand the user's current emotional state. This data is stored on the device as image and audio data.
[0196] Step 2:
[0197] The device analyzes the collected facial expression data using OpenCV and quantifies indicators of specific emotions (e.g., joy, surprise). This process extracts emotion values from the input image data.
[0198] Step 3:
[0199] Simultaneously, the device converts the audio data into text using Google Cloud Speech-to-Text and analyzes the tone of speech and important keywords using natural language processing. This extracts textual information related to emotions from the input audio data.
[0200] Step 4:
[0201] The terminal sends the analysis results to the server, which uses a generative AI model to integrate the analyzed image and text data to determine the overall emotional state. Here, we perform an integrated analysis of the output from OpenCV and Google Cloud Speech-to-Text.
[0202] Step 5:
[0203] The server selects appropriate music data based on the emotional state. This selection process uses TensorFlow to generate a list of the most suitable music based on the analysis results. This results in music data that corresponds to the user's emotions.
[0204] Step 6:
[0205] The server transmits the selected music data to the terminal or Bluetooth-enabled audio playback device and issues a real-time instruction to play the music. As a result, music that reflects the inputted emotional state is played in the store.
[0206] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0207] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0208] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0209] [Second Embodiment]
[0210] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0211] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0212] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0213] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0214] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0215] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0216] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0217] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0218] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0219] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0220] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0221] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0222] This invention is a system that allows users to easily generate and experience emotionally rich custom music for special anniversaries and events. The system receives event information and emotional data provided by the user and generates music based on this data. Specifically, it is implemented as follows:
[0223] First, users input event information and the emotions they wish to convey through a dedicated application or web platform. This information is then processed on their device and sent to the server. This data can include a wide range of information, such as the type of event, participant information, and desired emotions or themes.
[0224] Next, the terminal sends the received information to the server, which then passes that information to the sentiment analysis module. This module uses natural language processing to extract important emotions and related themes from the data. Based on these results, a generative AI module automatically generates the lyrics and melody of the song.
[0225] The generated lyrics and melody are sent from the server to professional musicians. The musicians perform based on these and record the result. The recorded audio data is uploaded to the server as a music file.
[0226] Ultimately, the server sends the audio file to the user's device. The user can download this file and play it during a special event. This system is easy for users to use, yet technically, it utilizes advanced emotion analysis and generative AI to provide individually tailored, emotionally rich music.
[0227] For example, when generating a song for a wedding anniversary, the user inputs memories from their married life and feelings of gratitude they want to convey. The resulting song will express love and gratitude, making the wedding anniversary a more emotionally moving experience.
[0228] This system also includes functions for rights management and revenue sharing related to the secondary use of sound files, and can support procedures for commercial use. This allows users to use music for various purposes with peace of mind.
[0229] The following describes the processing flow.
[0230] Step 1:
[0231] Users input event information and the emotions they wish to convey using a dedicated application or web platform. This information includes the date, location, participants, and the event's theme. This prepares the system to provide the necessary information.
[0232] Step 2:
[0233] The terminal validates the user's input data, checking for any errors or missing information. Once validation is complete, it converts the information into a specified format and sends it to the server. This conversion is crucial for efficient data processing.
[0234] Step 3:
[0235] The server stores the received data and transfers it to the sentiment analysis module. The sentiment analysis module uses natural language processing techniques to extract key emotions and related themes from the input information. This prepares the data for the next processing step.
[0236] Step 4:
[0237] The server calls a generation AI module based on the results of emotion analysis to generate lyrics and melody. This process uses an algorithm designed to reflect the emotions and themes entered by the user. The generated song will be in line with the user's desires.
[0238] Step 5:
[0239] The server sends the generated lyrics and melody to professional musicians. The musicians record the song using instruments and vocals based on the instructions. This recording is a crucial step towards the final version of the song.
[0240] Step 6:
[0241] Musicians upload their completed performance data to a server. The server receives this data and, if necessary, adjusts the sound quality or converts it to the correct format. This ensures that the music meets the user's expectations in terms of quality.
[0242] Step 7:
[0243] The server sends the final audio file to the user's device. The user downloads the provided music and is ready to play it at the event. Through this process, the user can obtain customized music content.
[0244] (Example 1)
[0245] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0246] In modern times, people desire to easily create and experience music that richly expresses their emotions for special events and anniversaries. However, creating music requires specialized knowledge and skills, making it difficult for ordinary people to easily create satisfying music. Furthermore, the management of rights regarding the secondary use of created sound sources is often unclear, creating problems that prevent people from using them with peace of mind.
[0247] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0248] In this invention, the server includes means for performing sentiment analysis on data received via a user interface acting as an information input device, means for securely transmitting data from a user terminal to the server via digital data communication, and means for generating a download link for the generated sound source data and transmitting it to the user terminal. This makes it possible for ordinary users to securely generate and use emotionally rich music tailored to their individual needs without requiring specialized knowledge.
[0249] An "information input device" is a device that provides an interface for users to input data, and can receive input such as emotions and event information.
[0250] "Sentiment analysis" is a process that uses natural language processing techniques to extract emotions and related themes from input text data.
[0251] A "generation method" refers to a device or program for generating lyrics and melodies based on the results of emotion analysis.
[0252] "Performance control means" refers to a device or program for generating performance data based on the generated lyrics and melody, and for controlling that performance.
[0253] "Audio data" refers to digital music files of recorded performances, and is the final music content provided to the user.
[0254] A "download link" is an access link that sends the generated audio data to the user's device, allowing the user to receive and download the audio data.
[0255] "Digital data communication" refers to a communication method used to transmit data between a user terminal and a server, and is carried out using protocols that ensure security and confidentiality.
[0256] This invention is a system that allows users to easily generate and experience emotionally rich custom music for special anniversaries and events. Since this system creates music based on event information and emotionally related data provided by the user, a specific embodiment of this system is described below.
[0257] Users input information from their PCs, smartphones, or other devices using a dedicated application or web platform. This information includes the type and details of the event the user desires, as well as the emotions or themes they wish to convey. This input information is transmitted to the server as encrypted digital data.
[0258] The server uses an emotion analysis module to process the received data. This module incorporates natural language processing technology to extract emotions and related themes from the data. Next, the server uses a generative AI model to automatically generate lyrics and melodies based on the extracted emotion data. At this time, it is given a specific prompt message such as, "Generate an upbeat melody that expresses happiness, and lyrics that convey gratitude."
[0259] The generated lyrics and melody are sent from the server to professional musicians, who then perform the song based on them. The performance is recorded and saved as audio data. The recorded audio data is finally stored on the server and provided to the user's device as a download link.
[0260] For example, if a user requests a song themed around emotions such as "love" or "memories" for their wedding anniversary, the system will generate a customized song based on the user's detailed memories and emotions, which can then be used to enhance the special anniversary celebration. This system provides a way for users to easily generate emotionally rich songs without requiring specialized knowledge, ensuring a worry-free experience.
[0261] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0262] Step 1:
[0263] Users input event information and desired emotions using a dedicated application or web platform. This input includes event type, participant information, and specific themes or keywords. This information is formatted as digital data on the user's device and sent to the server.
[0264] Step 2:
[0265] The terminal sends user input information to the server using a secure communication protocol, and the server receives it. The entered data is encrypted, ensuring security and privacy.
[0266] Step 3:
[0267] The server passes the received data to the sentiment analysis module. The module uses natural language processing to extract important emotions and themes from the provided data. This analysis converts the extracted sentiment data into a new digital data format.
[0268] Step 4:
[0269] Based on the results of the emotion analysis, the server uses a generative AI model to automatically generate lyrics and melodies. During this process, prompts such as "Please generate a gentle melody and lyrics that convey love and gratitude" are input to the AI. The AI then outputs song data that matches the analysis results, following these prompts.
[0270] Step 5:
[0271] The server sends the generated lyrics and melody to professional musicians. The musicians perform the song based on the provided information and record the performance in a studio in high quality. After the recorded data is edited, it is converted into a digital file as audio data.
[0272] Step 6:
[0273] The server uploads the audio data received from the musicians to its storage system and provides a download link to the user's device. The user then obtains the audio data via this link and is ready to play it at the event as needed.
[0274] (Application Example 1)
[0275] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0276] Traditionally, generating emotionally rich custom music required specialized knowledge and skills, making it difficult for the average user to access. Furthermore, there were few easy ways to share the generated music with others, limiting its use to special occasions and events. Additionally, insufficient rights management regarding the secondary use of the generated audio data presented challenges for commercial use. There is a need for a system that solves these problems and allows users to easily generate and share individually tailored music.
[0277] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0278] In this invention, the server includes means for emotionally analyzing data received via a user interface as an information input means; means for generating lyrics and melodies based on the emotional analysis; means for generating performance data based on the lyrics and melody; means for recording the performance data and generating it as an acoustic signal; means for providing the acoustic signal to a mobile terminal; means for sharing the acoustic signal using a communication function; and control for managing rights related to the secondary use of the sound source data and for revenue sharing. As a result, users can generate emotionally rich music without requiring special skills, easily share it with others, and manage the commercial use of the generated music.
[0279] The "information input means" is an interface for a user to input information related to special anniversaries or events.
[0280] The "means for sentiment analysis" is a mechanism having a function of analyzing sentiment and theme using natural language processing technology based on the information input by the user.
[0281] The "generation means" is a technology for automatically generating lyrics and melody based on the result of sentiment analysis.
[0282] The "performance means" is a system for creating performance data based on the generated lyrics and melody.
[0283] The "means for generating as an acoustic signal" is a process for recording performance data and outputting it as a music file.
[0284] The "providing means" is a method for transmitting the generated acoustic signal to the user's terminal.
[0285] The "sharing means" is a mechanism for a user to easily share an acoustic signal with others using a communication function.
[0286] The "control for managing rights and performing profit distribution" is a management function for appropriately managing the rights related to the use of sound source data and distributing profits in commercial use.
[0287] The system for implementing this invention generates an original music piece according to a special anniversary or event by the user and provides a rich emotional experience. The user inputs event information and emotional data via an application on a smartphone or other terminal. The input data is transmitted from the terminal to the server. The server analyzes the sentiment and theme in the input data by a sentiment analysis means using natural language processing technology.
[0288] Based on these analysis results, a generation AI model on the server automatically generates lyrics and melody. This generated music data is converted into performance data, which is then performed by professional musicians and recorded as an audio signal. The recorded audio signal is provided to the user's terminal by the server, and the user can receive, play, and share it.
[0289] Furthermore, the server manages the rights and controls revenue sharing related to the secondary use of audio signals, supporting the commercial use of the generated music. A key feature of this system is that users do not need any special technical knowledge to use it.
[0290] As a concrete example, consider a scenario where a user requests the creation of a song expressing gratitude for Mother's Day. The user inputs a message or feelings they want to convey to their mother, and the server generates a song expressing gratitude, such as "Thank you for everything." The generated song can then be played at a Mother's Day event where the family gathers, creating a special and moving experience.
[0291] As an example of a prompt in a generative AI model, the following text is used: "Today is Mother's Day. Last year I gave my mother her favorite flowers. This year I would like to give her a song to express my gratitude, so please generate a gentle and moving melody and lyrics. It would be desirable for the content to convey gratitude and happiness."
[0292] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0293] Step 1:
[0294] Users input event information and emotional data via a smartphone or device application. This input data includes details about specific anniversaries or events, as well as the emotions or themes they wish to convey. This data is then transmitted to the server via the user interface.
[0295] Step 2:
[0296] The server passes the received user data through a sentiment analysis system. This analysis system uses natural language processing techniques to extract keywords and phrases from the data and identify emotions and themes. The input is user data, and the output is the analyzed sentiment information.
[0297] Step 3:
[0298] The server uses a generative AI model to generate lyrics and melodies based on the acquired emotional information. In this process, the AI uses patterns learned from a vast amount of past data to create lyrics and melodies appropriate to the input emotion. The output is the generated song data.
[0299] Step 4:
[0300] Based on the generated music data, the server creates performance data. This performance data is used to give instructions to professional musicians to actually perform the music.
[0301] Step 5:
[0302] Professional musicians perform a song according to performance data, and the resulting audio signals are recorded. The recorded audio signals are then saved to a server as new music files.
[0303] Step 6:
[0304] The server provides the generated audio signal to the user's terminal. This music file can be downloaded by the user for use or playback at specific events.
[0305] Step 7:
[0306] Users play the generated audio signal on their device and share it with others as needed. Sharing can also include surveys and messages, making it a way to express gratitude to friends and family.
[0307] Furthermore, an emotion engine for estimating the user's emotion may be combined. That is, the specific processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform specific processing using the user's emotion.
[0308] The present invention relates to a system for generating and providing emotionally rich custom music on special commemorative days or events. This system is characterized by combining an emotion engine to recognize the user's emotion in real time and utilize it in the music generation process. The present invention is embodied as follows.
[0309] First, the user inputs information regarding the type and time of the event, the emotion and theme to be conveyed, using a dedicated application or web platform. At this time, the user interface also accepts input of voice and expressions and utilizes this for analysis. In addition, an emotion engine for performing real-time emotion recognition starts operating.
[0310] The terminal transmits the input information and the analysis result by the emotion engine to the server. The emotion engine identifies emotion from voice and expression data and reflects this in the data analysis process of the entire system. As a result, the user's current emotional state is considered as part of music generation.
[0311] Next, the server uses an emotion analysis module to integrate the data sent from the user and the output of the emotion engine, and extracts the main emotions and themes from the text. The complementary data by the emotion engine serves to reflect particularly delicate emotions and nuances in more detail. Thereby, the generated lyrics and melody enhance the empathy with the user's actual emotion.
[0312] The generated music information is transmitted from the server to professional musicians. The musicians perform the music in a form along with the emotion and theme and record it. This recording data is returned to the server and stored as a sound source file. The server also checks the quality of the sound source file and makes appropriate adjustments if necessary.
[0313] Ultimately, the server provides the completed audio file to the user's device. The user can download this audio and use it at their desired event. The system also facilitates rights management for secondary use of the audio and manages revenue sharing.
[0314] For example, if a user requests music to be played at a birthday party, their emotions specific to that day (e.g., surprise, joy) are recorded in real time through an emotion engine, and these emotions are reflected in the generated music. In this way, the present invention realizes a system that enriches individual user experiences and provides music closely related to emotions.
[0315] The following describes the processing flow.
[0316] Step 1:
[0317] Users access a dedicated application or web platform and enter event details. This includes the type of event, date and time, and participants, as well as the emotions or themes they want to reflect in the music, and relevant background information. In addition, users provide real-time audio or facial expression data to the system via webcam and microphone.
[0318] Step 2:
[0319] The device verifies the information received from the user and passes it to the emotion engine for analysis of audio and image data. The emotion engine uses audio and visual analysis algorithms to identify and quantify the current emotional state (e.g., joy, surprise, sadness). This information serves as the basis for the system to generate emotion-based music.
[0320] Step 3:
[0321] The terminal sends user data and sentiment analysis results from the sentiment engine to the server. The server receives and integrates this information and performs a detailed analysis through the sentiment analysis module. This analysis establishes key emotions and themes based on the user's input data and real-time recognized sentiment data.
[0322] Step 4:
[0323] The server uses the analysis results described above to run a generation AI module, generating lyrics and melodies that match the user's specific emotions. The AI takes into account the user's experiences and desired emotions, automatically creating a personalized song base. This process is supported by an algorithm that includes numerous parameters.
[0324] Step 5:
[0325] The server sends the generated lyrics and melody to professional musicians. The musicians perform the song based on the specified emotions and themes, and record the performance. High sound quality is ensured by using professional equipment. The performed song is recorded as the final audio data.
[0326] Step 6:
[0327] Musicians upload their completed recordings to a server. The server receives the data and verifies the sound quality and format. If necessary, the engine automatically adds effects and optimizes the quality before generating the final audio file.
[0328] Step 7:
[0329] The server sends the completed audio file to the user's terminal. The user can download this audio file and play it at the scheduled event. The system also provides the user with options for secondary use of the audio file and revenue sharing, and handles appropriate rights management if commercial use is desired.
[0330] (Example 2)
[0331] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0332] For special anniversaries and events, there is a demand for custom music that deeply resonates with the emotions of users. Conventional systems have struggled to reflect users' actual emotional states in real time and generate music based on those emotions. Therefore, there is a need to develop technology that accurately recognizes users' emotions and quickly generates music that is appropriate for them.
[0333] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0334] In this invention, the server includes a function to analyze emotions from data acquired through a user interface acting as an information receiving device, a function to recognize the user's emotional state by analyzing voice and facial expressions in real time, and an integration function to integrate the user's input data and the output from the emotion engine. This makes it possible to generate custom music that aligns with the user's emotions in real time.
[0335] "Information receiving device" is a general term for systems and devices used to acquire data through a user interface.
[0336] "Sentimental analysis" refers to the process of analyzing acquired data to identify the emotional state of users.
[0337] "Creation function" refers to a system that has the ability to generate text and music based on sentiment surveys.
[0338] The "performance function" refers to the process of creating musical information based on generated text and music.
[0339] An "audio file" refers to a file that stores recorded audio data in digital format.
[0340] A "music professional" refers to an individual or group that possesses the skills to perform and record music based on generated musical information.
[0341] "Language processing technology" is a general term for technologies that analyze human language and process it using computers.
[0342] This invention relates to a system that generates and provides custom music for special anniversaries and events based on the user's emotions. The invention utilizes dedicated hardware and software. Specifically, it employs an application or web platform that acts as a user interface for the user to input their emotions, and an emotion engine that analyzes emotions in real time.
[0343] Users input event information and emotions through applications or web platforms. This information may also include voice input and facial expression data, which are used to more accurately understand the user's emotional state.
[0344] The device analyzes data received from the user and uses an emotion engine to recognize the emotional state in real time. This analysis information is then sent to the server. Voice analysis and facial recognition technologies are used for emotion recognition, enabling highly accurate emotion analysis.
[0345] The server generates music using a generative AI model based on the received data. During the generation process, natural language processing techniques are used to extract emotional and thematic keywords from the text. The generated music information is then sent to music professionals for actual performance and recording.
[0346] For example, if a user requests music to be played at a birthday party, they might enter a prompt such as, "Please generate a custom song for my birthday party. It should be fast-paced and have themes of joy and surprise." This allows the system to generate an original song in real time that reflects the user's emotions.
[0347] Ultimately, the generated music files are provided from the server to the user's terminal, allowing the user to utilize them at events. The system also manages the rights and revenue sharing for these music files, supporting secondary use of the music.
[0348] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0349] Step 1:
[0350] Users input the type of special anniversary or event, date and time, and the emotions and themes they wish to convey, using a dedicated application or web platform. Voice and facial expression input are also available, allowing users to express their emotions more accurately. The input in this step, including event information and voice and facial expression data, serves as the initial data for emotion analysis.
[0351] Step 2:
[0352] The terminal begins processing the input data received from the user. Using voice analysis software and facial recognition technology, it analyzes the user's emotions in real time. The input voice and image data are converted into numerical data to identify the user's emotions. This process outputs the analyzed emotion data, which is then sent to the server.
[0353] Step 3:
[0354] The server receives emotion data and theme information sent from the terminal as input. First, it activates the emotion analysis module and uses natural language processing techniques to extract keywords and major emotions from the text. This process generates structured data that matches the user's theme and emotions. The output consists of prompt text and emotion theme data necessary for music generation.
[0355] Step 4:
[0356] The server takes structured emotional theme data as input and generates music data using a generative AI model. During this generation process, the lyrics and melody of a custom song are designed based on extracted keywords and emotional states. The output is digital data containing music information, which is then sent to music professionals.
[0357] Step 5:
[0358] The server provides the generated music information to a music professional. The music professional performs and records the music according to this information. The music information is received as input, and an audio file is sent back to the server as output. This recorded data is saved as a high-quality audio file.
[0359] Step 6:
[0360] The server verifies the final audio file and prepares it for delivery to the user's terminal. If sound quality adjustments are needed, these are made possible, and the file becomes easily accessible to the user afterward. The user downloads the completed audio file and makes it available for use in a specific event. The output of this step is the completed music file that the user receives.
[0361] (Application Example 2)
[0362] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the smart glasses 214 as the "terminal".
[0363] Traditional in-store customer experiences are uniform, making it difficult to provide personalized experiences tailored to the emotional state of individual customers. As a result, there were limitations to improving customer purchasing intent and satisfaction. This invention solves the problem of providing an optimal atmosphere tailored to each individual customer by analyzing the customer's current emotional state in real time and dynamically changing the music in the store based on that analysis.
[0364] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0365] In this invention, the server includes means for performing sentiment analysis on data received via a user interface as an information input means, means for analyzing the customer's emotional state and providing music data corresponding to that state, and means for transmitting the music data to an audio playback device and playing the music in real time. This makes it possible to appropriately adjust the atmosphere in the store according to the customer's emotional state and provide an optimal shopping environment for each individual customer.
[0366] An "information input means" is an interface for receiving and processing data from a user.
[0367] "Emotion analysis means" refers to technology that identifies a user's emotional state based on received data.
[0368] The "generation method" refers to the function that creates appropriate lyrics and melodies based on the results of emotion analysis.
[0369] "Performance method" refers to a system that creates specific performance data based on the generated lyrics and melody.
[0370] "Means of generating as an audio file" refers to a function for recording performance data and saving it in a file format.
[0371] "Means of providing to users" refers to the means of distributing the generated audio files in a format that users can use.
[0372] "Means of providing music data" refers to the function of selecting and preparing music data that matches the analyzed emotional state.
[0373] "A means of transmitting to an audio playback device and playing music in real time" refers to a technology for transmitting prepared music data to a playback device and playing music immediately.
[0374] The system for implementing this invention utilizes specific hardware and software to provide music based on the emotional state of customers in physical stores.
[0375] The server first receives data transmitted from the terminal and analyzes the customer's current emotional state using emotion analysis tools. This analysis utilizes OpenCV for facial recognition and Google Cloud Speech-to-Text for speech analysis. The analyzed data is then used with a generative AI model to clearly determine the emotional state.
[0376] Subsequently, the server generates appropriate music data to be played in the store based on the generated sentiment information. The music data is selected and adjusted by a sentiment prediction model using TensorFlow. The selected music data is sent to the audio playback device and played in the store in real time.
[0377] This system makes it possible to provide an optimized music experience for each individual customer in physical stores, and to appropriately adjust the store's atmosphere to match the customer's emotions.
[0378] For example, playing relaxing jazz music when customers visit during busy weekday evenings can create a more comfortable shopping environment.
[0379] An example of a prompt message would be, "Calculate the customer's current emotional state based on their facial expression analysis and voice tone, and then select and play background music that reflects that state." This message specifically instructs the system to perform certain actions.
[0380] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0381] Step 1:
[0382] The device collects the user's facial expressions and voice using its camera and microphone. This provides raw data to understand the user's current emotional state. This data is stored on the device as image and audio data.
[0383] Step 2:
[0384] The device analyzes the collected facial expression data using OpenCV and quantifies indicators of specific emotions (e.g., joy, surprise). This process extracts emotion values from the input image data.
[0385] Step 3:
[0386] Simultaneously, the device converts the audio data into text using Google Cloud Speech-to-Text and analyzes the tone of speech and important keywords using natural language processing. This extracts textual information related to emotions from the input audio data.
[0387] Step 4:
[0388] The terminal sends the analysis results to the server, which uses a generative AI model to integrate the analyzed image and text data to determine the overall emotional state. Here, we perform an integrated analysis of the output from OpenCV and Google Cloud Speech-to-Text.
[0389] Step 5:
[0390] The server selects appropriate music data based on the emotional state. This selection process uses TensorFlow to generate a list of the most suitable music based on the analysis results. This results in music data that corresponds to the user's emotions.
[0391] Step 6:
[0392] The server transmits the selected music data to the terminal or Bluetooth-enabled audio playback device and issues a real-time instruction to play the music. As a result, music that reflects the inputted emotional state is played in the store.
[0393] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0394] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0395] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0396] [Third Embodiment]
[0397] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0398] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0399] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0400] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0401] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0402] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0403] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0404] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0405] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0406] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0407] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0408] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0409] This invention is a system that allows users to easily generate and experience emotionally rich custom music for special anniversaries and events. The system receives event information and emotional data provided by the user and generates music based on this data. Specifically, it is implemented as follows:
[0410] First, users input event information and the emotions they wish to convey through a dedicated application or web platform. This information is then processed on their device and sent to the server. This data can include a wide range of information, such as the type of event, participant information, and desired emotions or themes.
[0411] Next, the terminal sends the received information to the server, which then passes that information to the sentiment analysis module. This module uses natural language processing to extract important emotions and related themes from the data. Based on these results, a generative AI module automatically generates the lyrics and melody of the song.
[0412] The generated lyrics and melody are sent from the server to professional musicians. The musicians perform based on these and record the result. The recorded audio data is uploaded to the server as a music file.
[0413] Ultimately, the server sends the audio file to the user's device. The user can download this file and play it during a special event. This system is easy for users to use, yet technically, it utilizes advanced emotion analysis and generative AI to provide individually tailored, emotionally rich music.
[0414] For example, when generating a song for a wedding anniversary, the user inputs memories from their married life and feelings of gratitude they want to convey. The resulting song will express love and gratitude, making the wedding anniversary a more emotionally moving experience.
[0415] This system also includes functions for rights management and revenue sharing related to the secondary use of sound files, and can support procedures for commercial use. This allows users to use music for various purposes with peace of mind.
[0416] The following describes the processing flow.
[0417] Step 1:
[0418] Users input event information and the emotions they wish to convey using a dedicated application or web platform. This information includes the date, location, participants, and the event's theme. This prepares the system to provide the necessary information.
[0419] Step 2:
[0420] The terminal validates the user's input data, checking for any errors or missing information. Once validation is complete, it converts the information into a specified format and sends it to the server. This conversion is crucial for efficient data processing.
[0421] Step 3:
[0422] The server stores the received data and transfers it to the sentiment analysis module. The sentiment analysis module uses natural language processing techniques to extract key emotions and related themes from the input information. This prepares the data for the next processing step.
[0423] Step 4:
[0424] The server calls a generation AI module based on the results of emotion analysis to generate lyrics and melody. This process uses an algorithm designed to reflect the emotions and themes entered by the user. The generated song will be in line with the user's desires.
[0425] Step 5:
[0426] The server sends the generated lyrics and melody to professional musicians. The musicians record the song using instruments and vocals based on the instructions. This recording is a crucial step towards the final version of the song.
[0427] Step 6:
[0428] Musicians upload their completed performance data to a server. The server receives this data and, if necessary, adjusts the sound quality or converts it to the correct format. This ensures that the music meets the user's expectations in terms of quality.
[0429] Step 7:
[0430] The server sends the final audio file to the user's device. The user downloads the provided music and is ready to play it at the event. Through this process, the user can obtain customized music content.
[0431] (Example 1)
[0432] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0433] In modern times, people desire to easily create and experience music that richly expresses their emotions for special events and anniversaries. However, creating music requires specialized knowledge and skills, making it difficult for ordinary people to easily create satisfying music. Furthermore, the management of rights regarding the secondary use of created sound sources is often unclear, creating problems that prevent people from using them with peace of mind.
[0434] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0435] In this invention, the server includes means for performing sentiment analysis on data received via a user interface acting as an information input device, means for securely transmitting data from a user terminal to the server via digital data communication, and means for generating a download link for the generated sound source data and transmitting it to the user terminal. This makes it possible for ordinary users to securely generate and use emotionally rich music tailored to their individual needs without requiring specialized knowledge.
[0436] An "information input device" is a device that provides an interface for users to input data, and can receive input such as emotions and event information.
[0437] "Sentiment analysis" is a process that uses natural language processing techniques to extract emotions and related themes from input text data.
[0438] A "generation method" refers to a device or program for generating lyrics and melodies based on the results of emotion analysis.
[0439] "Performance control means" refers to a device or program for generating performance data based on the generated lyrics and melody, and for controlling that performance.
[0440] "Audio data" refers to digital music files of recorded performances, and is the final music content provided to the user.
[0441] A "download link" is an access link that sends the generated audio data to the user's device, allowing the user to receive and download the audio data.
[0442] "Digital data communication" refers to a communication method used to transmit data between a user terminal and a server, and is carried out using protocols that ensure security and confidentiality.
[0443] This invention is a system that allows users to easily generate and experience emotionally rich custom music for special anniversaries and events. Since this system creates music based on event information and emotionally related data provided by the user, a specific embodiment of this system is described below.
[0444] Users input information from their PCs, smartphones, or other devices using a dedicated application or web platform. This information includes the type and details of the event the user desires, as well as the emotions or themes they wish to convey. This input information is transmitted to the server as encrypted digital data.
[0445] The server uses an emotion analysis module to process the received data. This module incorporates natural language processing technology to extract emotions and related themes from the data. Next, the server uses a generative AI model to automatically generate lyrics and melodies based on the extracted emotion data. At this time, it is given a specific prompt message such as, "Generate an upbeat melody that expresses happiness, and lyrics that convey gratitude."
[0446] The generated lyrics and melody are sent from the server to professional musicians, who then perform the song based on them. The performance is recorded and saved as audio data. The recorded audio data is finally stored on the server and provided to the user's device as a download link.
[0447] For example, if a user requests a song themed around emotions such as "love" or "memories" for their wedding anniversary, the system will generate a customized song based on the user's detailed memories and emotions, which can then be used to enhance the special anniversary celebration. This system provides a way for users to easily generate emotionally rich songs without requiring specialized knowledge, ensuring a worry-free experience.
[0448] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0449] Step 1:
[0450] Users input event information and desired emotions using a dedicated application or web platform. This input includes event type, participant information, and specific themes or keywords. This information is formatted as digital data on the user's device and sent to the server.
[0451] Step 2:
[0452] The terminal sends user input information to the server using a secure communication protocol, and the server receives it. The entered data is encrypted, ensuring security and privacy.
[0453] Step 3:
[0454] The server passes the received data to the sentiment analysis module. The module uses natural language processing to extract important emotions and themes from the provided data. This analysis converts the extracted sentiment data into a new digital data format.
[0455] Step 4:
[0456] Based on the results of the emotion analysis, the server uses a generative AI model to automatically generate lyrics and melodies. During this process, prompts such as "Please generate a gentle melody and lyrics that convey love and gratitude" are input to the AI. The AI then outputs song data that matches the analysis results, following these prompts.
[0457] Step 5:
[0458] The server sends the generated lyrics and melody to professional musicians. The musicians perform the song based on the provided information and record the performance in a studio in high quality. After the recorded data is edited, it is converted into a digital file as audio data.
[0459] Step 6:
[0460] The server uploads the audio data received from the musicians to its storage system and provides a download link to the user's device. The user then obtains the audio data via this link and is ready to play it at the event as needed.
[0461] (Application Example 1)
[0462] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0463] Traditionally, generating emotionally rich custom music required specialized knowledge and skills, making it difficult for the average user to access. Furthermore, there were few easy ways to share the generated music with others, limiting its use to special occasions and events. Additionally, insufficient rights management regarding the secondary use of the generated audio data presented challenges for commercial use. There is a need for a system that solves these problems and allows users to easily generate and share individually tailored music.
[0464] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0465] In this invention, the server includes means for emotionally analyzing data received via a user interface as an information input means; means for generating lyrics and melodies based on the emotional analysis; means for generating performance data based on the lyrics and melody; means for recording the performance data and generating it as an acoustic signal; means for providing the acoustic signal to a mobile terminal; means for sharing the acoustic signal using a communication function; and control for managing rights related to the secondary use of the sound source data and for revenue sharing. As a result, users can generate emotionally rich music without requiring special skills, easily share it with others, and manage the commercial use of the generated music.
[0466] An "information input method" is an interface for users to input information about special anniversaries or events.
[0467] A "means for analyzing emotions" refers to a mechanism that uses natural language processing technology to analyze emotions and themes based on information input by the user.
[0468] "Generation method" refers to technology for automatically generating lyrics and melodies based on the results of emotion analysis.
[0469] "Performance method" refers to a system for creating performance data based on generated lyrics and melodies.
[0470] "Means of generating as an acoustic signal" refers to the process of recording performance data and outputting it as a music file.
[0471] "Means of provision" refers to a method for transmitting the generated acoustic signal to the user's terminal.
[0472] "Means of sharing" refers to a system that allows users to easily share audio signals with others using communication functions.
[0473] "Control for managing rights and distributing revenue" refers to a management function that appropriately manages the rights related to the use of sound source data and distributes revenue in commercial use.
[0474] The system implementing this invention allows users to generate original music tailored to special anniversaries and events, providing an emotionally rich experience. Users input event information and emotional data via an application on their smartphone or other device. The input data is transmitted from the device to a server. The server analyzes the emotions and themes in the input data using an emotion analysis means employing natural language processing technology.
[0475] Based on these analysis results, a generation AI model on the server automatically generates lyrics and melody. This generated music data is converted into performance data, which is then performed by professional musicians and recorded as an audio signal. The recorded audio signal is provided to the user's terminal by the server, and the user can receive, play, and share it.
[0476] Furthermore, the server manages the rights and controls revenue sharing related to the secondary use of audio signals, supporting the commercial use of the generated music. A key feature of this system is that users do not need any special technical knowledge to use it.
[0477] As a concrete example, consider a scenario where a user requests the creation of a song expressing gratitude for Mother's Day. The user inputs a message or feelings they want to convey to their mother, and the server generates a song expressing gratitude, such as "Thank you for everything." The generated song can then be played at a Mother's Day event where the family gathers, creating a special and moving experience.
[0478] As an example of a prompt in a generative AI model, the following text is used: "Today is Mother's Day. Last year I gave my mother her favorite flowers. This year I would like to give her a song to express my gratitude, so please generate a gentle and moving melody and lyrics. It would be desirable for the content to convey gratitude and happiness."
[0479] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0480] Step 1:
[0481] Users input event information and emotional data via a smartphone or device application. This input data includes details about specific anniversaries or events, as well as the emotions or themes they wish to convey. This data is then transmitted to the server via the user interface.
[0482] Step 2:
[0483] The server passes the received user data through a sentiment analysis system. This analysis system uses natural language processing techniques to extract keywords and phrases from the data and identify emotions and themes. The input is user data, and the output is the analyzed sentiment information.
[0484] Step 3:
[0485] The server uses a generative AI model to generate lyrics and melodies based on the acquired emotional information. In this process, the AI uses patterns learned from a vast amount of past data to create lyrics and melodies appropriate to the input emotion. The output is the generated song data.
[0486] Step 4:
[0487] Based on the generated music data, the server creates performance data. This performance data is used to give instructions to professional musicians to actually perform the music.
[0488] Step 5:
[0489] Professional musicians perform a song according to performance data, and the resulting audio signals are recorded. The recorded audio signals are then saved to a server as new music files.
[0490] Step 6:
[0491] The server provides the generated audio signal to the user's terminal. This music file can be downloaded by the user for use or playback at specific events.
[0492] Step 7:
[0493] Users play the generated audio signal on their device and share it with others as needed. Sharing can also include surveys and messages, making it a way to express gratitude to friends and family.
[0494] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0495] This invention relates to a system for generating and providing emotionally rich custom music for special anniversaries and events. A key feature of this system is its ability to recognize the user's emotions in real time by incorporating an emotion engine, and to utilize this information in the music generation process. The invention is embodied as follows:
[0496] First, the user inputs information about the event type, date and time, and the emotions and themes they wish to convey through a dedicated application or web platform. The user interface also accepts voice and facial expression input, which is used for analysis. In addition, an emotion engine for real-time emotion recognition begins to operate.
[0497] The terminal sends the input information and the analysis results from the emotion engine to the server. The emotion engine identifies emotions from voice and facial expression data and incorporates this into the system's overall data analysis process. As a result, the user's current emotional state is taken into consideration as part of the music generation process.
[0498] Next, the server uses an emotion analysis module to integrate the data sent from the user with the output of the emotion engine, extracting key emotions and themes from the text. The supplementary data from the emotion engine plays a particularly important role in reflecting subtle emotions and nuances in greater detail. As a result, the generated lyrics and melodies are more empathetic to the user's actual emotions.
[0499] The generated song information is sent from the server to professional musicians. The musicians perform and record the song in accordance with the emotions and themes. This recording data is returned to the server and saved as an audio file. The server also checks the quality of the audio file and makes adjustments as needed.
[0500] Ultimately, the server provides the completed audio file to the user's device. The user can download this audio and use it at their desired event. The system also facilitates rights management for secondary use of the audio and manages revenue sharing.
[0501] For example, if a user requests music to be played at a birthday party, their emotions specific to that day (e.g., surprise, joy) are recorded in real time through an emotion engine, and these emotions are reflected in the generated music. In this way, the present invention realizes a system that enriches individual user experiences and provides music closely related to emotions.
[0502] The following describes the processing flow.
[0503] Step 1:
[0504] Users access a dedicated application or web platform and enter event details. This includes the type of event, date and time, and participants, as well as the emotions or themes they want to reflect in the music, and relevant background information. In addition, users provide real-time audio or facial expression data to the system via webcam and microphone.
[0505] Step 2:
[0506] The device verifies the information received from the user and passes it to the emotion engine for analysis of audio and image data. The emotion engine uses audio and visual analysis algorithms to identify and quantify the current emotional state (e.g., joy, surprise, sadness). This information serves as the basis for the system to generate emotion-based music.
[0507] Step 3:
[0508] The terminal sends user data and sentiment analysis results from the sentiment engine to the server. The server receives and integrates this information and performs a detailed analysis through the sentiment analysis module. This analysis establishes key emotions and themes based on the user's input data and real-time recognized sentiment data.
[0509] Step 4:
[0510] The server uses the analysis results described above to run a generation AI module, generating lyrics and melodies that match the user's specific emotions. The AI takes into account the user's experiences and desired emotions, automatically creating a personalized song base. This process is supported by an algorithm that includes numerous parameters.
[0511] Step 5:
[0512] The server sends the generated lyrics and melody to professional musicians. The musicians perform the song based on the specified emotions and themes, and record the performance. High sound quality is ensured by using professional equipment. The performed song is recorded as the final audio data.
[0513] Step 6:
[0514] Musicians upload their completed recordings to a server. The server receives the data and verifies the sound quality and format. If necessary, the engine automatically adds effects and optimizes the quality before generating the final audio file.
[0515] Step 7:
[0516] The server sends the completed audio file to the user's terminal. The user can download this audio file and play it at the scheduled event. The system also provides the user with options for secondary use of the audio file and revenue sharing, and handles appropriate rights management if commercial use is desired.
[0517] (Example 2)
[0518] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0519] For special anniversaries and events, there is a demand for custom music that deeply resonates with the emotions of users. Conventional systems have struggled to reflect users' actual emotional states in real time and generate music based on those emotions. Therefore, there is a need to develop technology that accurately recognizes users' emotions and quickly generates music that is appropriate for them.
[0520] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0521] In this invention, the server includes a function to analyze emotions from data acquired through a user interface acting as an information receiving device, a function to recognize the user's emotional state by analyzing voice and facial expressions in real time, and an integration function to integrate the user's input data and the output from the emotion engine. This makes it possible to generate custom music that aligns with the user's emotions in real time.
[0522] "Information receiving device" is a general term for systems and devices used to acquire data through a user interface.
[0523] "Sentimental analysis" refers to the process of analyzing acquired data to identify the emotional state of users.
[0524] "Creation function" refers to a system that has the ability to generate text and music based on sentiment surveys.
[0525] The "performance function" refers to the process of creating musical information based on generated text and music.
[0526] An "audio file" refers to a file that stores recorded audio data in digital format.
[0527] A "music professional" refers to an individual or group that possesses the skills to perform and record music based on generated musical information.
[0528] "Language processing technology" is a general term for technologies that analyze human language and process it using computers.
[0529] This invention relates to a system that generates and provides custom music for special anniversaries and events based on the user's emotions. The invention utilizes dedicated hardware and software. Specifically, it employs an application or web platform that acts as a user interface for the user to input their emotions, and an emotion engine that analyzes emotions in real time.
[0530] Users input event information and emotions through applications or web platforms. This information may also include voice input and facial expression data, which are used to more accurately understand the user's emotional state.
[0531] The device analyzes data received from the user and uses an emotion engine to recognize the emotional state in real time. This analysis information is then sent to the server. Voice analysis and facial recognition technologies are used for emotion recognition, enabling highly accurate emotion analysis.
[0532] The server generates music using a generative AI model based on the received data. During the generation process, natural language processing techniques are used to extract emotional and thematic keywords from the text. The generated music information is then sent to music professionals for actual performance and recording.
[0533] For example, if a user requests music to be played at a birthday party, they might enter a prompt such as, "Please generate a custom song for my birthday party. It should be fast-paced and have themes of joy and surprise." This allows the system to generate an original song in real time that reflects the user's emotions.
[0534] Ultimately, the generated music files are provided from the server to the user's terminal, allowing the user to utilize them at events. The system also manages the rights and revenue sharing for these music files, supporting secondary use of the music.
[0535] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0536] Step 1:
[0537] Users input the type of special anniversary or event, date and time, and the emotions and themes they wish to convey, using a dedicated application or web platform. Voice and facial expression input are also available, allowing users to express their emotions more accurately. The input in this step, including event information and voice and facial expression data, serves as the initial data for emotion analysis.
[0538] Step 2:
[0539] The terminal begins processing the input data received from the user. Using voice analysis software and facial recognition technology, it analyzes the user's emotions in real time. The input voice and image data are converted into numerical data to identify the user's emotions. This process outputs the analyzed emotion data, which is then sent to the server.
[0540] Step 3:
[0541] The server receives emotion data and theme information sent from the terminal as input. First, it activates the emotion analysis module and uses natural language processing techniques to extract keywords and major emotions from the text. This process generates structured data that matches the user's theme and emotions. The output consists of prompt text and emotion theme data necessary for music generation.
[0542] Step 4:
[0543] The server takes structured emotional theme data as input and generates music data using a generative AI model. During this generation process, the lyrics and melody of a custom song are designed based on extracted keywords and emotional states. The output is digital data containing music information, which is then sent to music professionals.
[0544] Step 5:
[0545] The server provides the generated music information to a music professional. The music professional performs and records the music according to this information. The music information is received as input, and an audio file is sent back to the server as output. This recorded data is saved as a high-quality audio file.
[0546] Step 6:
[0547] The server verifies the final audio file and prepares it for delivery to the user's terminal. If sound quality adjustments are needed, these are made possible, and the file becomes easily accessible to the user afterward. The user downloads the completed audio file and makes it available for use in a specific event. The output of this step is the completed music file that the user receives.
[0548] (Application Example 2)
[0549] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0550] Traditional in-store customer experiences are uniform, making it difficult to provide personalized experiences tailored to the emotional state of individual customers. As a result, there were limitations to improving customer purchasing intent and satisfaction. This invention solves the problem of providing an optimal atmosphere tailored to each individual customer by analyzing the customer's current emotional state in real time and dynamically changing the music in the store based on that analysis.
[0551] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0552] In this invention, the server includes means for performing sentiment analysis on data received via a user interface as an information input means, means for analyzing the customer's emotional state and providing music data corresponding to that state, and means for transmitting the music data to an audio playback device and playing the music in real time. This makes it possible to appropriately adjust the atmosphere in the store according to the customer's emotional state and provide an optimal shopping environment for each individual customer.
[0553] An "information input means" is an interface for receiving and processing data from a user.
[0554] "Emotion analysis means" refers to technology that identifies a user's emotional state based on received data.
[0555] The "generation method" refers to the function that creates appropriate lyrics and melodies based on the results of emotion analysis.
[0556] "Performance method" refers to a system that creates specific performance data based on the generated lyrics and melody.
[0557] "Means of generating as an audio file" refers to a function for recording performance data and saving it in a file format.
[0558] "Means of providing to users" refers to the means of distributing the generated audio files in a format that users can use.
[0559] "Means of providing music data" refers to the function of selecting and preparing music data that matches the analyzed emotional state.
[0560] "A means of transmitting to an audio playback device and playing music in real time" refers to a technology for transmitting prepared music data to a playback device and playing music immediately.
[0561] The system for implementing this invention utilizes specific hardware and software to provide music based on the emotional state of customers in physical stores.
[0562] The server first receives data transmitted from the terminal and analyzes the customer's current emotional state using emotion analysis tools. This analysis utilizes OpenCV for facial recognition and Google Cloud Speech-to-Text for speech analysis. The analyzed data is then used with a generative AI model to clearly determine the emotional state.
[0563] Subsequently, the server generates appropriate music data to be played in the store based on the generated sentiment information. The music data is selected and adjusted by a sentiment prediction model using TensorFlow. The selected music data is sent to the audio playback device and played in the store in real time.
[0564] This system makes it possible to provide an optimized music experience for each individual customer in physical stores, and to appropriately adjust the store's atmosphere to match the customer's emotions.
[0565] For example, playing relaxing jazz music when customers visit during busy weekday evenings can create a more comfortable shopping environment.
[0566] An example of a prompt message would be, "Calculate the customer's current emotional state based on their facial expression analysis and voice tone, and then select and play background music that reflects that state." This message specifically instructs the system to perform certain actions.
[0567] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0568] Step 1:
[0569] The device collects the user's facial expressions and voice using its camera and microphone. This provides raw data to understand the user's current emotional state. This data is stored on the device as image and audio data.
[0570] Step 2:
[0571] The device analyzes the collected facial expression data using OpenCV and quantifies indicators of specific emotions (e.g., joy, surprise). This process extracts emotion values from the input image data.
[0572] Step 3:
[0573] Simultaneously, the device converts the audio data into text using Google Cloud Speech-to-Text and analyzes the tone of speech and important keywords using natural language processing. This extracts textual information related to emotions from the input audio data.
[0574] Step 4:
[0575] The terminal sends the analysis results to the server, which uses a generative AI model to integrate the analyzed image and text data to determine the overall emotional state. Here, we perform an integrated analysis of the output from OpenCV and Google Cloud Speech-to-Text.
[0576] Step 5:
[0577] The server selects appropriate music data based on the emotional state. This selection process uses TensorFlow to generate a list of the most suitable music based on the analysis results. This results in music data that corresponds to the user's emotions.
[0578] Step 6:
[0579] The server transmits the selected music data to the terminal or Bluetooth-enabled audio playback device and issues a real-time instruction to play the music. As a result, music that reflects the inputted emotional state is played in the store.
[0580] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0581] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0582] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[0583] [Fourth Embodiment]
[0584] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0585] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0586] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0587] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0588] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0589] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0590] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0591] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0592] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0593] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0594] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0595] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0596] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0597] This invention is a system that allows users to easily generate and experience emotionally rich custom music for special anniversaries and events. The system receives event information and emotional data provided by the user and generates music based on this data. Specifically, it is implemented as follows:
[0598] First, users input event information and the emotions they wish to convey through a dedicated application or web platform. This information is then processed on their device and sent to the server. This data can include a wide range of information, such as the type of event, participant information, and desired emotions or themes.
[0599] Next, the terminal sends the received information to the server, which then passes that information to the sentiment analysis module. This module uses natural language processing to extract important emotions and related themes from the data. Based on these results, a generative AI module automatically generates the lyrics and melody of the song.
[0600] The generated lyrics and melody are sent from the server to professional musicians. The musicians perform based on these and record the result. The recorded audio data is uploaded to the server as a music file.
[0601] Ultimately, the server sends the audio file to the user's device. The user can download this file and play it during a special event. This system is easy for users to use, yet technically, it utilizes advanced emotion analysis and generative AI to provide individually tailored, emotionally rich music.
[0602] For example, when generating a song for a wedding anniversary, the user inputs memories from their married life and feelings of gratitude they want to convey. The resulting song will express love and gratitude, making the wedding anniversary a more emotionally moving experience.
[0603] This system also includes functions for rights management and revenue sharing related to the secondary use of sound files, and can support procedures for commercial use. This allows users to use music for various purposes with peace of mind.
[0604] The following describes the processing flow.
[0605] Step 1:
[0606] Users input event information and the emotions they wish to convey using a dedicated application or web platform. This information includes the date, location, participants, and the event's theme. This prepares the system to provide the necessary information.
[0607] Step 2:
[0608] The terminal validates the user's input data, checking for any errors or missing information. Once validation is complete, it converts the information into a specified format and sends it to the server. This conversion is crucial for efficient data processing.
[0609] Step 3:
[0610] The server stores the received data and transfers it to the sentiment analysis module. The sentiment analysis module uses natural language processing techniques to extract key emotions and related themes from the input information. This prepares the data for the next processing step.
[0611] Step 4:
[0612] The server calls a generation AI module based on the results of emotion analysis to generate lyrics and melody. This process uses an algorithm designed to reflect the emotions and themes entered by the user. The generated song will be in line with the user's desires.
[0613] Step 5:
[0614] The server sends the generated lyrics and melody to professional musicians. The musicians record the song using instruments and vocals based on the instructions. This recording is a crucial step towards the final version of the song.
[0615] Step 6:
[0616] Musicians upload their completed performance data to a server. The server receives this data and, if necessary, adjusts the sound quality or converts it to the correct format. This ensures that the music meets the user's expectations in terms of quality.
[0617] Step 7:
[0618] The server sends the final audio file to the user's device. The user downloads the provided music and is ready to play it at the event. Through this process, the user can obtain customized music content.
[0619] (Example 1)
[0620] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0621] In modern times, people desire to easily create and experience music that richly expresses their emotions for special events and anniversaries. However, creating music requires specialized knowledge and skills, making it difficult for ordinary people to easily create satisfying music. Furthermore, the management of rights regarding the secondary use of created sound sources is often unclear, creating problems that prevent people from using them with peace of mind.
[0622] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0623] In this invention, the server includes means for performing sentiment analysis on data received via a user interface acting as an information input device, means for securely transmitting data from a user terminal to the server via digital data communication, and means for generating a download link for the generated sound source data and transmitting it to the user terminal. This makes it possible for ordinary users to securely generate and use emotionally rich music tailored to their individual needs without requiring specialized knowledge.
[0624] An "information input device" is a device that provides an interface for users to input data, and can receive input such as emotions and event information.
[0625] "Sentiment analysis" is a process that uses natural language processing techniques to extract emotions and related themes from input text data.
[0626] A "generation method" refers to a device or program for generating lyrics and melodies based on the results of emotion analysis.
[0627] "Performance control means" refers to a device or program for generating performance data based on the generated lyrics and melody, and for controlling that performance.
[0628] "Audio data" refers to digital music files of recorded performances, and is the final music content provided to the user.
[0629] A "download link" is an access link that sends the generated audio data to the user's device, allowing the user to receive and download the audio data.
[0630] "Digital data communication" refers to a communication method used to transmit data between a user terminal and a server, and is carried out using protocols that ensure security and confidentiality.
[0631] This invention is a system that allows users to easily generate and experience emotionally rich custom music for special anniversaries and events. Since this system creates music based on event information and emotionally related data provided by the user, a specific embodiment of this system is described below.
[0632] Users input information from their PCs, smartphones, or other devices using a dedicated application or web platform. This information includes the type and details of the event the user desires, as well as the emotions or themes they wish to convey. This input information is transmitted to the server as encrypted digital data.
[0633] The server uses an emotion analysis module to process the received data. This module incorporates natural language processing technology to extract emotions and related themes from the data. Next, the server uses a generative AI model to automatically generate lyrics and melodies based on the extracted emotion data. At this time, it is given a specific prompt message such as, "Generate an upbeat melody that expresses happiness, and lyrics that convey gratitude."
[0634] The generated lyrics and melody are sent from the server to professional musicians, who then perform the song based on them. The performance is recorded and saved as audio data. The recorded audio data is finally stored on the server and provided to the user's device as a download link.
[0635] For example, if a user requests a song themed around emotions such as "love" or "memories" for their wedding anniversary, the system will generate a customized song based on the user's detailed memories and emotions, which can then be used to enhance the special anniversary celebration. This system provides a way for users to easily generate emotionally rich songs without requiring specialized knowledge, ensuring a worry-free experience.
[0636] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0637] Step 1:
[0638] Users input event information and desired emotions using a dedicated application or web platform. This input includes event type, participant information, and specific themes or keywords. This information is formatted as digital data on the user's device and sent to the server.
[0639] Step 2:
[0640] The terminal sends user input information to the server using a secure communication protocol, and the server receives it. The entered data is encrypted, ensuring security and privacy.
[0641] Step 3:
[0642] The server passes the received data to the sentiment analysis module. The module uses natural language processing to extract important emotions and themes from the provided data. This analysis converts the extracted sentiment data into a new digital data format.
[0643] Step 4:
[0644] Based on the results of the emotion analysis, the server uses a generative AI model to automatically generate lyrics and melodies. During this process, prompts such as "Please generate a gentle melody and lyrics that convey love and gratitude" are input to the AI. The AI then outputs song data that matches the analysis results, following these prompts.
[0645] Step 5:
[0646] The server sends the generated lyrics and melody to professional musicians. The musicians perform the song based on the provided information and record the performance in a studio in high quality. After the recorded data is edited, it is converted into a digital file as audio data.
[0647] Step 6:
[0648] The server uploads the audio data received from the musicians to its storage system and provides a download link to the user's device. The user then obtains the audio data via this link and is ready to play it at the event as needed.
[0649] (Application Example 1)
[0650] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0651] Traditionally, generating emotionally rich custom music required specialized knowledge and skills, making it difficult for the average user to access. Furthermore, there were few easy ways to share the generated music with others, limiting its use to special occasions and events. Additionally, insufficient rights management regarding the secondary use of the generated audio data presented challenges for commercial use. There is a need for a system that solves these problems and allows users to easily generate and share individually tailored music.
[0652] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0653] In this invention, the server includes means for emotionally analyzing data received via a user interface as an information input means; means for generating lyrics and melodies based on the emotional analysis; means for generating performance data based on the lyrics and melody; means for recording the performance data and generating it as an acoustic signal; means for providing the acoustic signal to a mobile terminal; means for sharing the acoustic signal using a communication function; and control for managing rights related to the secondary use of the sound source data and for revenue sharing. As a result, users can generate emotionally rich music without requiring special skills, easily share it with others, and manage the commercial use of the generated music.
[0654] An "information input method" is an interface for users to input information about special anniversaries or events.
[0655] A "means for analyzing emotions" refers to a mechanism that uses natural language processing technology to analyze emotions and themes based on information input by the user.
[0656] "Generation method" refers to technology for automatically generating lyrics and melodies based on the results of emotion analysis.
[0657] "Performance method" refers to a system for creating performance data based on generated lyrics and melodies.
[0658] "Means of generating as an acoustic signal" refers to the process of recording performance data and outputting it as a music file.
[0659] "Means of provision" refers to a method for transmitting the generated acoustic signal to the user's terminal.
[0660] "Means of sharing" refers to a system that allows users to easily share audio signals with others using communication functions.
[0661] "Control for managing rights and distributing revenue" refers to a management function that appropriately manages the rights related to the use of sound source data and distributes revenue in commercial use.
[0662] The system implementing this invention allows users to generate original music tailored to special anniversaries and events, providing an emotionally rich experience. Users input event information and emotional data via an application on their smartphone or other device. The input data is transmitted from the device to a server. The server analyzes the emotions and themes in the input data using an emotion analysis means employing natural language processing technology.
[0663] Based on these analysis results, a generation AI model on the server automatically generates lyrics and melody. This generated music data is converted into performance data, which is then performed by professional musicians and recorded as an audio signal. The recorded audio signal is provided to the user's terminal by the server, and the user can receive, play, and share it.
[0664] Furthermore, the server manages the rights and controls revenue sharing related to the secondary use of audio signals, supporting the commercial use of the generated music. A key feature of this system is that users do not need any special technical knowledge to use it.
[0665] As a concrete example, consider a scenario where a user requests the creation of a song expressing gratitude for Mother's Day. The user inputs a message or feelings they want to convey to their mother, and the server generates a song expressing gratitude, such as "Thank you for everything." The generated song can then be played at a Mother's Day event where the family gathers, creating a special and moving experience.
[0666] As an example of a prompt in a generative AI model, the following text is used: "Today is Mother's Day. Last year I gave my mother her favorite flowers. This year I would like to give her a song to express my gratitude, so please generate a gentle and moving melody and lyrics. It would be desirable for the content to convey gratitude and happiness."
[0667] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0668] Step 1:
[0669] Users input event information and emotional data via a smartphone or device application. This input data includes details about specific anniversaries or events, as well as the emotions or themes they wish to convey. This data is then transmitted to the server via the user interface.
[0670] Step 2:
[0671] The server passes the received user data through a sentiment analysis system. This analysis system uses natural language processing techniques to extract keywords and phrases from the data and identify emotions and themes. The input is user data, and the output is the analyzed sentiment information.
[0672] Step 3:
[0673] The server uses a generative AI model to generate lyrics and melodies based on the acquired emotional information. In this process, the AI uses patterns learned from a vast amount of past data to create lyrics and melodies appropriate to the input emotion. The output is the generated song data.
[0674] Step 4:
[0675] Based on the generated music data, the server creates performance data. This performance data is used to give instructions to professional musicians to actually perform the music.
[0676] Step 5:
[0677] Professional musicians perform a song according to performance data, and the resulting audio signals are recorded. The recorded audio signals are then saved to a server as new music files.
[0678] Step 6:
[0679] The server provides the generated audio signal to the user's terminal. This music file can be downloaded by the user for use or playback at specific events.
[0680] Step 7:
[0681] Users play the generated audio signal on their device and share it with others as needed. Sharing can also include surveys and messages, making it a way to express gratitude to friends and family.
[0682] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0683] This invention relates to a system for generating and providing emotionally rich custom music for special anniversaries and events. A key feature of this system is its ability to recognize the user's emotions in real time by incorporating an emotion engine, and to utilize this information in the music generation process. The invention is embodied as follows:
[0684] First, the user inputs information about the event type, date and time, and the emotions and themes they wish to convey through a dedicated application or web platform. The user interface also accepts voice and facial expression input, which is used for analysis. In addition, an emotion engine for real-time emotion recognition begins to operate.
[0685] The terminal sends the input information and the analysis results from the emotion engine to the server. The emotion engine identifies emotions from voice and facial expression data and incorporates this into the system's overall data analysis process. As a result, the user's current emotional state is taken into consideration as part of the music generation process.
[0686] Next, the server uses an emotion analysis module to integrate the data sent from the user with the output of the emotion engine, extracting key emotions and themes from the text. The supplementary data from the emotion engine plays a particularly important role in reflecting subtle emotions and nuances in greater detail. As a result, the generated lyrics and melodies are more empathetic to the user's actual emotions.
[0687] The generated song information is sent from the server to professional musicians. The musicians perform and record the song in accordance with the emotions and themes. This recording data is returned to the server and saved as an audio file. The server also checks the quality of the audio file and makes adjustments as needed.
[0688] Ultimately, the server provides the completed audio file to the user's device. The user can download this audio and use it at their desired event. The system also facilitates rights management for secondary use of the audio and manages revenue sharing.
[0689] For example, if a user requests music to be played at a birthday party, their emotions specific to that day (e.g., surprise, joy) are recorded in real time through an emotion engine, and these emotions are reflected in the generated music. In this way, the present invention realizes a system that enriches individual user experiences and provides music closely related to emotions.
[0690] The following describes the processing flow.
[0691] Step 1:
[0692] Users access a dedicated application or web platform and enter event details. This includes the type of event, date and time, and participants, as well as the emotions or themes they want to reflect in the music, and relevant background information. In addition, users provide real-time audio or facial expression data to the system via webcam and microphone.
[0693] Step 2:
[0694] The device verifies the information received from the user and passes it to the emotion engine for analysis of audio and image data. The emotion engine uses audio and visual analysis algorithms to identify and quantify the current emotional state (e.g., joy, surprise, sadness). This information serves as the basis for the system to generate emotion-based music.
[0695] Step 3:
[0696] The terminal sends user data and sentiment analysis results from the sentiment engine to the server. The server receives and integrates this information and performs a detailed analysis through the sentiment analysis module. This analysis establishes key emotions and themes based on the user's input data and real-time recognized sentiment data.
[0697] Step 4:
[0698] The server uses the analysis results described above to run a generation AI module, generating lyrics and melodies that match the user's specific emotions. The AI takes into account the user's experiences and desired emotions, automatically creating a personalized song base. This process is supported by an algorithm that includes numerous parameters.
[0699] Step 5:
[0700] The server sends the generated lyrics and melody to professional musicians. The musicians perform the song based on the specified emotions and themes, and record the performance. High sound quality is ensured by using professional equipment. The performed song is recorded as the final audio data.
[0701] Step 6:
[0702] Musicians upload their completed recordings to a server. The server receives the data and verifies the sound quality and format. If necessary, the engine automatically adds effects and optimizes the quality before generating the final audio file.
[0703] Step 7:
[0704] The server sends the completed audio file to the user's terminal. The user can download this audio file and play it at the scheduled event. The system also provides the user with options for secondary use of the audio file and revenue sharing, and handles appropriate rights management if commercial use is desired.
[0705] (Example 2)
[0706] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0707] For special anniversaries and events, there is a demand for custom music that deeply resonates with the emotions of users. Conventional systems have struggled to reflect users' actual emotional states in real time and generate music based on those emotions. Therefore, there is a need to develop technology that accurately recognizes users' emotions and quickly generates music that is appropriate for them.
[0708] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0709] In this invention, the server includes a function to analyze emotions from data acquired through a user interface acting as an information receiving device, a function to recognize the user's emotional state by analyzing voice and facial expressions in real time, and an integration function to integrate the user's input data and the output from the emotion engine. This makes it possible to generate custom music that aligns with the user's emotions in real time.
[0710] "Information receiving device" is a general term for systems and devices used to acquire data through a user interface.
[0711] "Sentimental analysis" refers to the process of analyzing acquired data to identify the emotional state of users.
[0712] "Creation function" refers to a system that has the ability to generate text and music based on sentiment surveys.
[0713] The "performance function" refers to the process of creating musical information based on generated text and music.
[0714] An "audio file" refers to a file that stores recorded audio data in digital format.
[0715] A "music professional" refers to an individual or group that possesses the skills to perform and record music based on generated musical information.
[0716] "Language processing technology" is a general term for technologies that analyze human language and process it using computers.
[0717] This invention relates to a system that generates and provides custom music for special anniversaries and events based on the user's emotions. The invention utilizes dedicated hardware and software. Specifically, it employs an application or web platform that acts as a user interface for the user to input their emotions, and an emotion engine that analyzes emotions in real time.
[0718] Users input event information and emotions through applications or web platforms. This information may also include voice input and facial expression data, which are used to more accurately understand the user's emotional state.
[0719] The device analyzes data received from the user and uses an emotion engine to recognize the emotional state in real time. This analysis information is then sent to the server. Voice analysis and facial recognition technologies are used for emotion recognition, enabling highly accurate emotion analysis.
[0720] The server generates music using a generative AI model based on the received data. During the generation process, natural language processing techniques are used to extract emotional and thematic keywords from the text. The generated music information is then sent to music professionals for actual performance and recording.
[0721] For example, if a user requests music to be played at a birthday party, they might enter a prompt such as, "Please generate a custom song for my birthday party. It should be fast-paced and have themes of joy and surprise." This allows the system to generate an original song in real time that reflects the user's emotions.
[0722] Ultimately, the generated music files are provided from the server to the user's terminal, allowing the user to utilize them at events. The system also manages the rights and revenue sharing for these music files, supporting secondary use of the music.
[0723] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0724] Step 1:
[0725] Users input the type of special anniversary or event, date and time, and the emotions and themes they wish to convey, using a dedicated application or web platform. Voice and facial expression input are also available, allowing users to express their emotions more accurately. The input in this step, including event information and voice and facial expression data, serves as the initial data for emotion analysis.
[0726] Step 2:
[0727] The terminal begins processing the input data received from the user. Using voice analysis software and facial recognition technology, it analyzes the user's emotions in real time. The input voice and image data are converted into numerical data to identify the user's emotions. This process outputs the analyzed emotion data, which is then sent to the server.
[0728] Step 3:
[0729] The server receives emotion data and theme information sent from the terminal as input. First, it activates the emotion analysis module and uses natural language processing techniques to extract keywords and major emotions from the text. This process generates structured data that matches the user's theme and emotions. The output consists of prompt text and emotion theme data necessary for music generation.
[0730] Step 4:
[0731] The server takes structured emotional theme data as input and generates music data using a generative AI model. During this generation process, the lyrics and melody of a custom song are designed based on extracted keywords and emotional states. The output is digital data containing music information, which is then sent to music professionals.
[0732] Step 5:
[0733] The server provides the generated music information to a music professional. The music professional performs and records the music according to this information. The music information is received as input, and an audio file is sent back to the server as output. This recorded data is saved as a high-quality audio file.
[0734] Step 6:
[0735] The server verifies the final audio file and prepares it for delivery to the user's terminal. If sound quality adjustments are needed, these are made possible, and the file becomes easily accessible to the user afterward. The user downloads the completed audio file and makes it available for use in a specific event. The output of this step is the completed music file that the user receives.
[0736] (Application Example 2)
[0737] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0738] Traditional in-store customer experiences are uniform, making it difficult to provide personalized experiences tailored to the emotional state of individual customers. As a result, there were limitations to improving customer purchasing intent and satisfaction. This invention solves the problem of providing an optimal atmosphere tailored to each individual customer by analyzing the customer's current emotional state in real time and dynamically changing the music in the store based on that analysis.
[0739] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0740] In this invention, the server includes means for performing sentiment analysis on data received via a user interface as an information input means, means for analyzing the customer's emotional state and providing music data corresponding to that state, and means for transmitting the music data to an audio playback device and playing the music in real time. This makes it possible to appropriately adjust the atmosphere in the store according to the customer's emotional state and provide an optimal shopping environment for each individual customer.
[0741] An "information input means" is an interface for receiving and processing data from a user.
[0742] "Emotion analysis means" refers to technology that identifies a user's emotional state based on received data.
[0743] The "generation method" refers to the function that creates appropriate lyrics and melodies based on the results of emotion analysis.
[0744] "Performance method" refers to a system that creates specific performance data based on the generated lyrics and melody.
[0745] "Means of generating as an audio file" refers to a function for recording performance data and saving it in a file format.
[0746] "Means of providing to users" refers to the means of distributing the generated audio files in a format that users can use.
[0747] "Means of providing music data" refers to the function of selecting and preparing music data that matches the analyzed emotional state.
[0748] "A means of transmitting to an audio playback device and playing music in real time" refers to a technology for transmitting prepared music data to a playback device and playing music immediately.
[0749] The system for implementing this invention utilizes specific hardware and software to provide music based on the emotional state of customers in physical stores.
[0750] The server first receives data transmitted from the terminal and analyzes the customer's current emotional state using emotion analysis tools. This analysis utilizes OpenCV for facial recognition and Google Cloud Speech-to-Text for speech analysis. The analyzed data is then used with a generative AI model to clearly determine the emotional state.
[0751] Subsequently, the server generates appropriate music data to be played in the store based on the generated sentiment information. The music data is selected and adjusted by a sentiment prediction model using TensorFlow. The selected music data is sent to the audio playback device and played in the store in real time.
[0752] This system makes it possible to provide an optimized music experience for each individual customer in physical stores, and to appropriately adjust the store's atmosphere to match the customer's emotions.
[0753] For example, playing relaxing jazz music when customers visit during busy weekday evenings can create a more comfortable shopping environment.
[0754] An example of a prompt message would be, "Calculate the customer's current emotional state based on their facial expression analysis and voice tone, and then select and play background music that reflects that state." This message specifically instructs the system to perform certain actions.
[0755] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0756] Step 1:
[0757] The device collects the user's facial expressions and voice using its camera and microphone. This provides raw data to understand the user's current emotional state. This data is stored on the device as image and audio data.
[0758] Step 2:
[0759] The device analyzes the collected facial expression data using OpenCV and quantifies indicators of specific emotions (e.g., joy, surprise). This process extracts emotion values from the input image data.
[0760] Step 3:
[0761] Simultaneously, the device converts the audio data into text using Google Cloud Speech-to-Text and analyzes the tone of speech and important keywords using natural language processing. This extracts textual information related to emotions from the input audio data.
[0762] Step 4:
[0763] The terminal sends the analysis results to the server, which uses a generative AI model to integrate the analyzed image and text data to determine the overall emotional state. Here, we perform an integrated analysis of the output from OpenCV and Google Cloud Speech-to-Text.
[0764] Step 5:
[0765] The server selects appropriate music data based on the emotional state. This selection process uses TensorFlow to generate a list of the most suitable music based on the analysis results. This results in music data that corresponds to the user's emotions.
[0766] Step 6:
[0767] The server transmits the selected music data to the terminal or Bluetooth-enabled audio playback device and issues a real-time instruction to play the music. As a result, music that reflects the inputted emotional state is played in the store.
[0768] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0769] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0770] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[0771] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0772] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0773] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0774] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0775] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0776] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0777] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0778] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[0779] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[0780] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0781] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0782] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0783] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0784] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0785] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0786] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0787] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0788] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted as being incorporated by reference.
[0789] The following is further disclosed regarding the embodiments described above.
[0790] (Claim 1)
[0791] A means for performing sentiment analysis on data received via a user interface as an information input means,
[0792] A generation means for generating lyrics and melodies based on the aforementioned emotion analysis,
[0793] A performance means that generates performance data based on the aforementioned lyrics and melody,
[0794] A means for recording the aforementioned performance data and generating it as an audio file,
[0795] Means for providing the aforementioned sound source file to the user,
[0796] A system that includes this.
[0797] (Claim 2)
[0798] The system according to claim 1, further comprising means for managing rights relating to secondary use of the aforementioned sound source files and for distributing revenue.
[0799] (Claim 3)
[0800] The system according to claim 1, wherein the sentiment analysis means has a function of extracting keywords and phrases from the data using natural language processing.
[0801] "Example 1"
[0802] (Claim 1)
[0803] A means for performing sentiment analysis on data received via a user interface as an information input device,
[0804] A generation means for generating lyrics and melodies based on the aforementioned emotion analysis,
[0805] A performance control means that generates performance data based on the lyrics and melody,
[0806] A means for recording the aforementioned performance data and generating it as sound source data,
[0807] A means for providing the aforementioned sound source data to the user,
[0808] A means of securely transmitting data from a user terminal to a server via digital data communication,
[0809] A means for generating a download link for the generated audio data and sending it to the user's terminal,
[0810] A system that includes this.
[0811] (Claim 2)
[0812] The system according to claim 1, further comprising means for managing rights relating to the secondary use of the aforementioned sound source data and for distributing revenue.
[0813] (Claim 3)
[0814] The system according to claim 1, wherein the sentiment analysis means has the function of extracting keywords and related themes from the data using natural language processing technology.
[0815] "Application Example 1"
[0816] (Claim 1)
[0817] A means for performing sentiment analysis on data received via a user interface as an information input means,
[0818] A generation means for generating lyrics and melodies based on the aforementioned emotion analysis,
[0819] A performance means that generates performance data based on the aforementioned lyrics and melody,
[0820] A means for recording the aforementioned performance data and generating it as an acoustic signal,
[0821] Means for providing the aforementioned acoustic signal to a mobile terminal,
[0822] Means for sharing the aforementioned acoustic signal using a communication function,
[0823] A system that includes this.
[0824] (Claim 2)
[0825] The system according to claim 1, further comprising controls for managing rights relating to the secondary use of the aforementioned sound source data and for revenue sharing.
[0826] (Claim 3)
[0827] The system according to claim 1, wherein the emotion analysis means has a function of extracting relevant words and phrases from the data using language processing technology.
[0828] "Example 2 of combining an emotion engine"
[0829] (Claim 1)
[0830] A function to conduct sentiment surveys on data acquired through a user interface acting as an information receiving device,
[0831] A creation function that generates text and music based on the aforementioned sentiment survey,
[0832] A performance function that generates musical information based on the aforementioned text and music,
[0833] A function to record the aforementioned music information and generate it as an audio file,
[0834] A means for supplying the aforementioned audio file to the user,
[0835] A function that analyzes voice and facial expressions in real time to recognize the user's emotional state,
[0836] An integration function that combines user input data and output from the emotion engine,
[0837] It has a function to send the generated song information to music professionals for performance and recording,
[0838] A system that includes this.
[0839] (Claim 2)
[0840] The system according to claim 1, further comprising means for adjusting rights regarding secondary use of the aforementioned audio files and for revenue sharing.
[0841] (Claim 3)
[0842] The system according to claim 1, wherein the sentiment survey function has a function to extract specific words or phrases from the data using language processing technology.
[0843] "Application example 2 of combining emotional engines"
[0844] (Claim 1)
[0845] A means for performing sentiment analysis on data received via a user interface as an information input means,
[0846] A generation means for generating lyrics and melodies based on the aforementioned emotion analysis,
[0847] A performance means that generates performance data based on the aforementioned lyrics and melody,
[0848] A means for recording the aforementioned performance data and generating it as an audio file,
[0849] Means for providing the aforementioned sound source file to the user,
[0850] A means of analyzing the emotional state of a customer and providing music data appropriate to that state,
[0851] A means for transmitting the aforementioned music data to an audio playback device and playing the music in real time,
[0852] A system that includes this.
[0853] (Claim 2)
[0854] The system according to claim 1, further comprising means for managing rights relating to secondary use of the aforementioned sound source files and for distributing revenue.
[0855] (Claim 3)
[0856] The system according to claim 1, characterized in that the emotion analysis means has a function to extract keywords and phrases in the data using natural language processing, and the music data is provided in real time. [Explanation of Symbols]
[0857] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. A means of performing sentiment analysis on data received via a user interface as an information input means, A generation means for generating lyrics and melodies based on the aforementioned emotion analysis, A performance means that generates performance data based on the aforementioned lyrics and melody, A means for recording the aforementioned performance data and generating it as an audio file, Means for providing the aforementioned sound source file to the user, A system that includes this.
2. The system according to claim 1, further comprising means for managing rights relating to secondary use of the aforementioned sound source files and for distributing revenue.
3. The system according to claim 1, wherein the emotion analysis means has a function of extracting keywords and phrases from the data using natural language processing.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A