System
The system allows users to freely select and share singer-song combinations in real time, addressing the limitations of conventional music applications by enabling flexible music experiences through voice synthesis and sharing capabilities.
Patent Information
- Application Number
- JP2024121509
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-26
- Publication Date
- 2026-02-05
AI Technical Summary
Conventional music applications limit users to a fixed list of songs, restricting the flexibility of their music experience and making it difficult to freely select singer-song combinations and share them in real time.
A system that allows users to select singer and song combinations, synthesizes music using voice synthesis engines, and enables real-time playback and sharing through a user terminal, server, and database, with features for creating and sharing playlists.
Enables users to freely select and share singer-song combinations in real time, providing a flexible and convenient music experience.
Smart Images

Figure 2026019761000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] Conventional music applications have not provided users with the experience of listening to different songs with the voice of a specific singer. As a result, users are limited to selecting from a limited list of songs, limiting the flexibility of their music experience. The present invention aims to solve this problem by providing a system that allows users to freely select singer-song combinations and play and share them in real time. [Means for solving the problem]
[0005] The system according to the present invention includes the following means: First, a means for processing singer and song selection information received from a user terminal is provided; Based on this selection information, a means for retrieving singer voice data and song score data from a database is provided; Next, a means for synthesizing a song using a voice synthesis engine using the retrieved voice data and song score data is provided; Finally, a means for transmitting the synthesized voice data to the user terminal is provided; Furthermore, a means for providing a search interface that allows the user to select a singer and song and for transmitting input from the interface to a server is provided; Also, a means for providing the user with a link to the synthesized voice based on the received selection results is provided; Finally, a means for creating a playlist that allows the user to save songs and share them on social media, a means for saving the playlist in a database and generating a dedicated link, and a means for providing the generated link to the user terminal. In this way, a system is realized that provides users with a free music experience and opportunities to share.
[0006] A "user terminal" is a device used by a user for access, and includes a smartphone, tablet, PC, etc.
[0007] "Singer and song selection information" is data that allows the user to specify a combination of a specific singer and song.
[0008] The "search interface" is a user interface that allows a user to search for singers and songs, and includes a text box and a search button.
[0009] A "server" is a computer system that receives and processes requests from user terminals.
[0010] A "database" is an information storage system that stores singers' voice data, music score data, and so on.
[0011] A "voice synthesis engine" is a software program that uses the voice data of a specific singer to generate music based on sheet music data.
[0012] "Audio data" refers to a file that digitally stores the voice of a particular singer.
[0013] "Music score data" is a file that digitally represents the melody and rhythm of a song.
[0014] "Synthesized audio data" is a new music file generated using a speech synthesis engine.
[0015] A "link" is reference information such as a URL for accessing specified digital content.
[0016] A "playlist" is a digital collection of songs that can be saved and shared by users.
[0017] "Means for saving" refers to the process for storing the selected songs and playlists in a database or the like.
[0018] "Means for sharing" is a function for sharing playlists and song links created by a user with other users. [Brief explanation of the drawings]
[0019] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram illustrating a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0020] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0021] First, the terms used in the following description will be explained.
[0022] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0023] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0024] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0025] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0026] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0027] [First embodiment]
[0028] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0029] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0030] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0031] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0032] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0033] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0034] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0035] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0036] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0037] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0038] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0039] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0040] The present invention is a system that allows users to freely select singers and song combinations, play them in real time, and share them. This system is composed of a user terminal, a server, a database, and a voice synthesis engine.
[0041] Overall system configuration
[0042] 1. User Device
[0043] The user terminal is a device such as a smartphone, tablet, or PC, and provides a user interface through which the user can search for and select singers and songs.
[0044] 2. Server
[0045] The server is responsible for processing the singer and song selection information received from the user device. The server analyzes the received data, retrieves the singer's voice data and song score data from the database, sends a request to the voice synthesis engine, and transmits the synthesized voice data to the user device.
[0046] 3. Database
[0047] The database stores the singer's voice data and the music score data. After the server receives the request, it retrieves the necessary data from the database.
[0048] 4. Speech synthesis engine
[0049] The voice synthesis engine uses the acquired singer's voice data and the music score data to synthesize a song according to the specified conditions. The synthesized voice data is sent to the user's device via the server.
[0050] Program processing and specific examples
[0051] Program processing
[0052] The program supports users in a series of operations from logging in to playing and sharing songs. The specific processing steps of the program are as follows:
[0053] 1. User login
[0054] The user enters the information required to log in to the app (user ID, password, etc.) and is authenticated. If authentication is successful, the user is redirected to the home screen.
[0055] 2. Search and select an artist and song
[0056] The user searches for singers and songs using keywords in the search interface and selects from the corresponding candidates. The selected information is sent to the server.
[0057] 3. Data Acquisition and Speech Synthesis
[0058] The server retrieves the singer's voice data and the song's score data from a database based on the received selection information. The retrieved data is passed to a voice synthesis engine, which synthesizes the song according to the specified conditions. The resulting voice data is returned to the server and sent to the user's device.
[0059] 4. Playing Music and Creating Playlists
[0060] Users can play the synthesized music using the link sent to them, and can also save multiple songs as a playlist and share it on social media.
[0061] Specific examples
[0062] For example, suppose a user wants to listen to "Song B" in the voice of "Singer A."
[0063] 1. Log in
[0064] A user logs in to the app.
[0065] Enter your username and password and the server will authenticate you.
[0066] If the authentication is successful, the user will be taken to the home screen.
[0067] 2. Selection of singers and songs
[0068] The user enters "singer A" in the search bar and selects "singer A" from the list of singers.
[0069] Next, enter "Song B" in the song search bar and select "Song B" from the list of songs.
[0070] 3. Data Acquisition and Speech Synthesis
[0071] The server processes the combination of "singer A" and "song B" and retrieves the respective data from the database.
[0072] The server passes this data to a voice synthesis engine, which synthesizes music according to the specified conditions.
[0073] 4. Playing songs
[0074] The server transmits the synthesized voice data to the user terminal.
[0075] The user plays "Song B" with the voice of "Singer A" on the device and enjoys the music.
[0076] 5. Create and share playlists
[0077] Users can save the songs they play as a playlist and share it on social media.
[0078] Users can easily share music with friends and family using the links of their saved playlists.
[0079] The present invention is a system that can provide users with a new way of enjoying music.
[0080] The processing flow will be explained below.
[0081] Step 1:
[0082] A user logs in to the app
[0083] The user launches the app on the device and accesses the login screen.
[0084] The terminal receives the user ID and password entered by the user.
[0085] The terminal sends this login information to the server.
[0086] The server checks the received user ID and password in a database and returns the authentication result to the terminal.
[0087] If the authentication is successful, the terminal transitions the user to the home screen.
[0088] Step 2:
[0089] User searches and selects singer
[0090] The user enters the singer's name into the device's search bar.
[0091] The terminal sends the entered name to the server.
[0092] The server searches the database for a list of singers based on the received singer name and returns the list to the terminal.
[0093] The user selects a particular singer from a list of singers displayed on the terminal.
[0094] The terminal transmits information about the selected singer to the server.
[0095] Step 3:
[0096] User searches and selects a song
[0097] The user enters the title of the song into the device's search bar.
[0098] The terminal transmits the input song title to the server.
[0099] The server searches the database for a list of songs based on the received song title and returns it to the terminal.
[0100] The user selects a specific song from the song list displayed on the terminal.
[0101] The terminal transmits information about the selected song to the server.
[0102] Step 4:
[0103] The server retrieves the data
[0104] The server analyzes the received selection information (singer name and song title).
[0105] The server retrieves the singer's voice data and the music score data from the database.
[0106] The acquired data is temporarily stored on the server.
[0107] Step 5:
[0108] Perform speech synthesis
[0109] The server passes the acquired voice data and musical score data to the voice synthesis engine.
[0110] The speech synthesis engine generates synthetic music based on this data.
[0111] The synthesized music file is sent back to the server.
[0112] Step 6:
[0113] Sending synthesized voice data
[0114] The server transmits the synthesized voice data to the user terminal as streaming or a download link.
[0115] The terminal receives this and makes it available for playback by the user.
[0116] Step 7:
[0117] User plays music
[0118] The user clicks on the link provided by the device and plays the music.
[0119] The device streams the audio data and plays it back in real time.
[0120] Step 8:
[0121] Create and share playlists
[0122] After listening to a number of songs, the user accesses a playlist creation screen.
[0123] The terminal sends a request to the server to save the list of songs selected by the user as a playlist.
[0124] The server stores the playlist information in a database and generates a playlist ID.
[0125] A link is created based on the generated playlist ID and sent to the device.
[0126] Users can share this link on social media and with other users.
[0127] Example 1
[0128] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0129] With conventional systems, it was difficult for users to play back their desired singer-song combinations in real time and easily share them. Furthermore, the process of synthesizing audio data and sheet music data was complex and time-consuming, which sometimes compromised the user experience. This resulted in a lack of flexibility and convenience for proposing new ways to enjoy music.
[0130] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0131] In this invention, the server includes means for processing selection information of voice data and music data received from a user terminal, means for acquiring voice data and music score data from a storage device based on the received selection information, means for synthesizing music using the acquired voice data and music score data with a voice synthesizer, and means for transmitting the synthesized voice data to the user terminal, thereby enabling users to freely select combinations of singers and music to be played in real time and easily shared.
[0132] A "user terminal" is a device that allows a user to interact with the system through an interface and perform various operations, and specifically refers to a smartphone, tablet, PC, etc.
[0133] "Selection Information" refers to data about a particular artist and song that a user enters through the system.
[0134] A "server" refers to a computer system that receives requests from user terminals and processes data and provides services in response to those requests.
[0135] "Audio data" refers to acoustic information that is a digital recording of the voice of a particular singer.
[0136] "Musical score data" refers to digital data that records musical information such as the melody, rhythm, and chords of a song.
[0137] "Storage device" refers to hardware and software for storing digital information such as audio data, sheet music data, and song selection lists.
[0138] A "speech synthesizer" refers to hardware or software that uses analog or digital voice data to generate new voices.
[0139] "Synthesized voice data" refers to music data generated by a voice synthesizer based on the voice data of a singer selected by the user and music score data.
[0140] A "search screen" refers to a user interface that provides an interface for a user to search for a specific singer or song.
[0141] "Song selection list" refers to a data list for managing multiple songs selected and saved by the user.
[0142] "Information sharing service" means an online platform that enables users to share information with other users via the Internet.
[0143] "Dedicated link" refers to a URL or URI generated to directly access a specific song selection list or voice synthesis data.
[0144] The present invention provides a system that allows users to freely select and share combinations of voice data and music data in real time. This system is comprised of a user terminal, a server, a storage device, and a voice synthesizer.
[0145] Overall system configuration
[0146] 1. User Device
[0147] The user terminal is a device such as a smartphone, tablet, or PC, and provides a user interface. The user can search and select audio data and music data through a search screen, allowing the user to easily find the music they want.
[0148] 2. Server
[0149] The server is responsible for processing the voice data and music data selection information received from the user terminal. The server analyzes the received data and retrieves the voice data and music score data from the storage device. It also sends a request to the voice synthesizer and sends the synthesized voice data to the user terminal. This allows the user to generate the music they want in real time.
[0150] 3. Storage device
[0151] The storage device stores the voice data and the musical score data. After the server receives the request, it retrieves the necessary data from the storage device. This data is provided to the voice synthesizer and used to synthesize the music.
[0152] 4. Speech synthesizer
[0153] The speech synthesizer uses the acquired voice data and musical score data to synthesize music under specified conditions. The synthesized voice data is sent to the user's terminal via a server. The speech synthesizer can generate music in real time.
[0154] Specific examples
[0155] Below is a specific example where a user wants to listen to "Song B" in the voice of "Singer A."
[0156] 1. Log in
[0157] A user logs in to a smartphone app.
[0158] Enter your username and password and the server will authenticate you.
[0159] If the authentication is successful, the user will be taken to the home screen.
[0160] 2. Selecting audio and music data
[0161] The user enters "singer A" in the search bar and selects "singer A" from the list of audio data.
[0162] Next, enter "Song B" in the song search bar and select "Song B" from the list of song data.
[0163] 3. Data Acquisition and Speech Synthesis
[0164] The server processes the combination of "singer A" and "song B" and retrieves the respective data from the storage device.
[0165] The server passes this data to a voice synthesizer, which synthesizes music according to the specified conditions.
[0166] 4. Playing songs
[0167] The server transmits the synthesized voice data to the user terminal.
[0168] The user plays "Song B" with the voice of "Singer A" on the device and enjoys the music.
[0169] 5. Create and share song lists
[0170] The songs played by the user are saved as a song selection list and shared via an information sharing service.
[0171] Users can easily share music with friends and family using links to their saved selections.
[0172] Prompt Sentence Examples
[0173] This is a system that generates music in real time using "Singer A" and "Song B." Specifically, a user logs into the app and enters "Singer A" and "Song B" in the search bar. The data for the selected singer and song is sent to a server, which then retrieves the necessary voice data and musical score data from a storage device and passes them to a voice synthesizer to synthesize the song. The synthesized song is sent to the user's device, where the user can play it.
[0174] With the above configuration, this system can provide users with a new way of enjoying music.
[0175] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0176] Step 1: The user launches the app and enters their user ID and password on the login screen.
[0177] Input: User ID, Password
[0178] Output: Authentication request
[0179] Specific operation: When the user taps the login button, the entered information is sent to the server.
[0180] Step 2: The server processes the received authentication information and checks it against a database.
[0181] Input: User ID, Password
[0182] Output: Authentication result
[0183] Specific operation: The server checks the user ID and password against the database, and if authentication is successful, generates a response indicating successful authentication.
[0184] Step 3: If the authentication is successful, the server sends an instruction to the user terminal to transition to the home screen.
[0185] Input: Authentication result (success)
[0186] Output: Home screen display instruction
[0187] Specific operation: Information indicating access rights to the home screen is sent to the user terminal, and the home screen is displayed on the user terminal.
[0188] Step 4: The user enters keywords for audio data (e.g., "singer A") and song data (e.g., "song B") into the search bar.
[0189] Input: Keywords for audio data and music data
[0190] Output: Search request
[0191] Specific operation: When the user taps the search button, the entered keywords are sent to the server.
[0192] Step 5: The server searches the storage device for related audio data and music data based on the received search keyword.
[0193] Input: Search keywords for audio data and music data
[0194] Output: Search results
[0195] Specific operation: The server searches the storage device and obtains the corresponding audio data and music data.
[0196] Step 6: The server sends the search results to the user terminal, and the user selects the appropriate singer and song.
[0197] Input: Search results
[0198] Output: User's choice
[0199] Specific operation: The user taps to select the desired artist and song from the list of search results.
[0200] Step 7: The server receives the user's selection information and retrieves the corresponding audio data and music score data from the storage device.
[0201] Input: User selection information
[0202] Output: Audio data and sheet music data
[0203] Specific operation: The server accesses the storage device and acquires the selected audio data and music score data.
[0204] Step 8: The server passes the acquired data to a voice synthesizer, which synthesizes music according to the specified conditions.
[0205] Input: Audio data and sheet music data
[0206] Output: Synthesized voice data
[0207] Specific operation: The voice synthesizer synthesizes music according to the specified conditions, and the generated voice data is sent back to the server.
[0208] Step 9: The server transmits the synthesized voice data to the user terminal.
[0209] Input: Synthesized voice data
[0210] Output: Playback link for audio data
[0211] Specific operation: The synthesized voice data is sent to the user's terminal and a playback link is displayed.
[0212] Step 10: The user taps the play link to play the synthesized voice data.
[0213] Input: Audio data playback link
[0214] Output: Playing music
[0215] Specific operation: The user device opens the link and plays the synthesized voice data.
[0216] Step 11: The user selects multiple songs, creates a song selection list, and shares it via an information sharing service.
[0217] Input: Song list
[0218] Output: Link to song selection list
[0219] Specific operation: The song list is saved in a storage device, and a link is generated and displayed on the user's device. The user can then share the link on social media.
[0220] (Application example 1)
[0221] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0222] Conventional music distribution services have difficulty generating and playing combinations of singers and songs freely selected by users in real time, limiting the variety of music experiences they can offer. Furthermore, there is a lack of systems that allow users to easily share the music they create through social networking sites or messaging applications. A system that can solve these issues and provide users with a new music experience is needed.
[0223] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0224] In this invention, the server includes means for processing singer and song selection information received from a user terminal, means for retrieving singer voice data and song score data from a database based on the received selection information, means for synthesizing a song using a voice synthesis engine using the retrieved voice data and song score data, means for transmitting the synthesized voice data to the user terminal, and means for users to play and share the song using a link to the synthesized voice data. This allows users to freely create and play back combinations of singers and songs in real time, and to easily share the created songs via social networking sites and messaging applications.
[0225] A "user terminal" is a device used by a user to operate the system, and includes smartphones, tablets, PCs, etc.
[0226] "Singer and song selection information" is information input by the user to combine a specific singer with a song.
[0227] "Audio data" refers to data that digitally records the voice of a particular singer.
[0228] "Musical score data" refers to data in which the notes, chords, rhythms, etc. of a piece of music are recorded in digital format.
[0229] "Database" means a data storage system for storing and managing audio data and musical score data.
[0230] A "voice synthesis engine" is software that generates music by combining voice data and musical score data.
[0231] "Synthetic voice data" is data of a song sung in the voice of a specific singer, generated by a voice synthesis engine.
[0232] A "search interface" is a screen or input device that allows a user to search for and select an artist and song.
[0233] A "server" is a computer system that processes requests from user terminals and provides information in cooperation with a database and a speech synthesis engine.
[0234] A "link" is a URL or URI that a user uses to play back synthesized speech data.
[0235] "SNS" is an abbreviation for social networking service, a platform for users to share information online.
[0236] A "messaging application" is software that allows users to exchange messages and information in real time.
[0237] A "playlist" is a collection of songs saved by a user and organized for easy playback and sharing.
[0238] The present invention is a system that receives information on singer and song selection from a user terminal, synthesizes the song in real time using a specific voice synthesis engine, and provides the generated voice data to the user. Specifically, the present invention is implemented using a user terminal, a server, a database, and a voice synthesis engine.
[0239] User terminal
[0240] The user terminal consists of a smartphone, tablet, PC, etc., and is the device through which the user operates the system. The user terminal is provided with a search interface for searching and selecting singers and songs. It also has the functionality to receive a link to the synthesized voice data and play and share it.
[0241] server
[0242] The server processes the selection information received from the user's device and retrieves the necessary audio data and music score data from the database. Specifically, the server was built using the Python Flask framework and receives input from the search interface. The server then works with a speech synthesis engine to generate music using the retrieved data. It is then responsible for sending a link to the generated audio data to the user's device.
[0243] Database
[0244] The database is a data storage system for storing and managing singers' voice data and music score data. The server quickly retrieves the necessary data from the database based on the selection information received.
[0245] Text-to-speech engine
[0246] A speech synthesis engine is software that synthesizes music based on specified conditions using acquired voice data and musical score data. Specifically, it can generate music in real time using external services such as speech synthesis APIs.
[0247] Specific examples of user operations
[0248] The user launches the app and logs in. They enter their favorite singer and song in the search bar and select from the search results. For example, if they want to listen to "Song B" in the voice of "Singer A," they enter and select this. The server retrieves the necessary data from the database based on the selection and passes it to the speech synthesis engine. A link to the synthesized voice data is sent to the user's device, and the user can use this link to play and share the song.
[0249] Prompt Sentence Examples
[0250] User: Please create song B using singer A's voice.
[0251] System: Synthesized audio data URL: http: / / example.com / audio / 1234
[0252] This invention allows users to freely select singers and song combinations, generate and play them in real time, and easily share them via social networking sites and messaging applications, providing a new musical experience.
[0253] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0254] Step 1:
[0255] The user starts the application on their device, such as a smartphone or PC, and goes to the login screen. The user enters their user ID and password and presses the login button.
[0256] Input: User ID, Password
[0257] Action: User device sends authentication information to server
[0258] Output: Authentication result from the server (success or failure)
[0259] Step 2:
[0260] The server authenticates the user based on the received authentication information. If authentication is successful, the user is redirected to the home screen and the search interface is displayed.
[0261] Input: Authentication information (user ID, password)
[0262] How it works: The server checks the authentication information against a database and sends the result back to the user's device.
[0263] Output: Authentication result (home screen if successful, error message if unsuccessful)
[0264] Step 3:
[0265] The user searches for and selects an artist and song in the search interface, for example, by entering "artist A" and "song B" in the search bar and selecting from the matching results.
[0266] Input: Search keyword (singer name, song name)
[0267] Action: Sends selections from the search interface to the server
[0268] Output: The server receives the selection information and displays the results on the user's terminal.
[0269] Step 4:
[0270] The server acquires the singer's voice data and the musical score data of the song from the database based on the received selection information.
[0271] Input: Selection information (singer name, song name)
[0272] How it works: The server sends a request to the database and retrieves the required data.
[0273] Output: Audio data, music score data
[0274] Step 5:
[0275] The server passes the acquired voice data and musical score data to a voice synthesis engine, which synthesizes music according to the specified conditions.
[0276] Input: Audio data, music score data
[0277] How it works: The server calls the speech synthesis engine and sends a synthesis request.
[0278] Output: Synthesized audio data
[0279] Step 6:
[0280] The server transmits the synthesized voice data to the user terminal and provides the user with a link to the synthesized voice data.
[0281] Input: Synthetic speech data
[0282] Operation: The server stores the synthesized voice data, generates a link to it, and sends it to the user's device.
[0283] Output: Link to the synthesized speech data
[0284] Step 7:
[0285] Users can use the provided link to play the synthesized voice and share it via social media or messaging applications.
[0286] Input: Synthetic speech data link
[0287] What happens: The user receives the link, plays the audio, and offers sharing options.
[0288] Output: Play audio, share
[0289] By following these steps, users can freely create and play combinations of singers and songs in real time, and can easily share the created songs via social media or messaging applications.
[0290] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0291] The present invention is a system that plays back singer-music combinations freely selected by the user in real time, and also recognizes the user's emotions to provide recommended content. This system is composed of a user terminal, a server, a database, a speech synthesis engine, and an emotion engine.
[0292] Overall system configuration
[0293] 1. User Device
[0294] The user terminal is a device such as a smartphone, tablet, or PC, and provides a user interface. From the user interface, users can search for and select singers and songs. It also has the function of sending the user's voice and text input to the emotion engine.
[0295] 2. Server
[0296] The server is responsible for processing the singer and song selection information received from the user terminal, as well as the emotion data from the emotion engine. The server analyzes the received data and retrieves the singer's voice data and song score data from the database. It also sends requests to the voice synthesis engine and transmits the synthesized voice data to the user terminal. It also has the function of generating recommended content based on the emotion data.
[0297] 3. Database
[0298] The database stores the singer's voice data, the music score data, and the user's emotion history data. After the server receives the request, it retrieves the required data from the database.
[0299] 4. Speech synthesis engine
[0300] The voice synthesis engine uses the acquired singer's voice data and the music score data to synthesize a song according to the specified conditions. The synthesized voice data is sent to the user's device via the server.
[0301] 5. Emotion Engine
[0302] The emotion engine recognizes emotions from the user's voice or text and sends the emotion data to the server. The recognized emotion data is stored in a database along with the user's emotion history data and used to generate recommended content.
[0303] Program processing and specific examples
[0304] Program processing
[0305] The program supports users in a series of operations, from logging in to playing songs, recognizing emotions, providing recommended content, and sharing. The specific processing steps of the program are as follows:
[0306] 1. User login
[0307] The user enters the information required to log in to the app (user ID, password, etc.) and is authenticated. If authentication is successful, the user is redirected to the home screen.
[0308] 2. Search and select an artist and song
[0309] The user searches for singers and songs using keywords in the search interface and selects from the corresponding candidates. The selected information is sent to the server.
[0310] 3. Data Acquisition and Speech Synthesis
[0311] The server analyzes the received selection information (singer name and song title) and retrieves the singer's voice data and the song's score data from the database. The retrieved data is passed to a voice synthesis engine, which synthesizes the song according to the specified conditions. The resulting voice data is returned to the server and sent to the user's device.
[0312] 4. Music playback and emotion recognition
[0313] The user can use the sent link to play the synthesized music, while the user's voice and text input are analyzed by the emotion engine.
[0314] 5. Providing recommended content
[0315] Based on the emotional data analyzed by the emotion engine, the server presents recommended singers and songs to the user, allowing the user to have a music experience that best suits their current emotions.
[0316] 6. Create and share playlists
[0317] After listening to multiple songs, the user accesses the playlist creation screen. The device sends a request to the server to save the list of songs selected by the user as a playlist. The server saves the playlist information in a database and generates a playlist ID. A link is created based on the generated playlist ID and sent to the device. The user can share this link via social media or with other users.
[0318] Specific examples
[0319] For example, suppose a user wants to listen to "Song B" in the voice of "Singer A."
[0320] 1. Log in
[0321] A user logs in to the app.
[0322] Enter your username and password and the server will authenticate you.
[0323] If the authentication is successful, the user will be taken to the home screen.
[0324] 2. Selection of singers and songs
[0325] The user enters "singer A" in the search bar and selects "singer A" from the list of singers.
[0326] Next, enter "Song B" in the song search bar and select "Song B" from the list of songs.
[0327] 3. Data Acquisition and Speech Synthesis
[0328] The server processes the combination of "singer A" and "song B" and retrieves the respective data from the database.
[0329] The server passes this data to a voice synthesis engine, which synthesizes music according to the specified conditions.
[0330] 4. Music playback and emotion recognition
[0331] The server transmits the synthesized voice data to the user terminal.
[0332] The user clicks on the link provided by the device and plays the music.
[0333] During playback, the user's voice and text input are analyzed by the emotion engine.
[0334] 5. Providing recommended content
[0335] The emotion data analyzed by the emotion engine is sent to the server.
[0336] The server recommends singers and songs that suit the user based on the emotion data.
[0337] 6. Create and share playlists
[0338] Users can save the songs they play as a playlist and share it on social media.
[0339] Users can easily share music with friends and family using the links of their saved playlists.
[0340] This invention is a system that provides users with a new way of enjoying music and also realizes a personalized music experience that matches the user's emotions.
[0341] The processing flow will be explained below.
[0342] Step 1:
[0343] A user logs in to the app
[0344] The user launches the app on the device and accesses the login screen.
[0345] The terminal receives the user ID and password entered by the user.
[0346] The terminal sends this login information to the server.
[0347] The server checks the received user ID and password in a database and returns the authentication result to the terminal.
[0348] If the authentication is successful, the terminal transitions the user to the home screen.
[0349] Step 2:
[0350] User searches and selects singer
[0351] The user enters the singer's name into the device's search bar.
[0352] The terminal sends the entered name to the server.
[0353] The server searches the database for a list of singers based on the received singer name and returns the list to the terminal.
[0354] The user selects a particular singer from a list of singers displayed on the terminal.
[0355] The terminal transmits information about the selected singer to the server.
[0356] Step 3:
[0357] User searches and selects a song
[0358] The user enters the title of the song into the device's search bar.
[0359] The terminal transmits the input song title to the server.
[0360] The server searches the database for a list of songs based on the received song title and returns it to the terminal.
[0361] The user selects a specific song from the song list displayed on the terminal.
[0362] The terminal transmits information about the selected song to the server.
[0363] Step 4:
[0364] The server retrieves the data
[0365] The server analyzes the received selection information (singer name and song title).
[0366] The server retrieves the singer's voice data and the music score data from the database.
[0367] The acquired data is temporarily stored on the server.
[0368] Step 5:
[0369] Perform speech synthesis
[0370] The server passes the acquired voice data and musical score data to the voice synthesis engine.
[0371] The speech synthesis engine generates synthetic music based on this data.
[0372] The synthesized music file is sent back to the server.
[0373] Step 6:
[0374] Sending synthesized voice data
[0375] The server transmits the synthesized voice data to the user terminal as streaming or a download link.
[0376] The terminal receives this and makes it available for playback by the user.
[0377] Step 7:
[0378] User plays music
[0379] The user clicks on the link provided by the device and plays the music.
[0380] The device streams the audio data and plays it back in real time.
[0381] Step 8:
[0382] Emotion recognition processing
[0383] Users can enter text and voice messages while listening to music.
[0384] The device sends the user's voice and text input to the emotion engine.
[0385] The emotion engine analyzes the user's input, recognizes the emotion, and sends the result to the server.
[0386] Step 9:
[0387] Providing recommended content
[0388] The server searches the database for singers and songs that suit the user's emotions based on the emotion data received from the emotion engine.
[0389] The server generates a list of relevant singers and songs and transmits it to the user terminal.
[0390] The user can check the recommended content displayed on the terminal and select new songs.
[0391] Step 10:
[0392] Create and share playlists
[0393] After playing a number of songs, the user accesses the playlist creation screen.
[0394] The terminal sends a request to the server to save the list of songs selected by the user as a playlist.
[0395] The server stores the playlist information in a database and generates a playlist ID.
[0396] A link is created based on the generated playlist ID and sent to the device.
[0397] Users can share this link on social media and with other users.
[0398] Example 2
[0399] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0400] Conventional music playback systems have difficulty in playing a combination of singers and songs freely selected by the user in real time. Furthermore, they lack the functionality to recognize the user's emotions and provide appropriate recommended content, making it difficult to provide a music experience that best suits the user's current emotions. To solve this problem, a system is needed that can efficiently process user selection information and emotional data, synthesize music in real time, and make appropriate recommendations.
[0401] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0402] In this invention, the server includes means for processing singer and song selection information received from a user terminal, means for retrieving singer voice data and song score data from a database based on the received selection information, means for synthesizing a song using a voice synthesis engine using the retrieved voice data and song score data, means for transmitting the user's voice and text input to an emotion engine, means for recognizing the user's emotion data and transmitting it to the server to generate recommended content, and means for transmitting the recommended content to the user terminal. This allows the user to freely select singers and songs to be played in real time, and further allows the user to enjoy a personalized music experience based on their emotions.
[0403] A "user terminal" is a device that allows a user to access and operate the system through an interface, and includes smartphones, tablets, PCs, etc.
[0404] A "server" is a computer system that processes data received from a user terminal, retrieves necessary data from a database based on selection information, and performs appropriate processing.
[0405] "Singer and song selection information" is data relating to the singer name and song title selected by the user through the system.
[0406] The "database" is a data storage system that stores necessary information such as singer's voice data, music score data, and user's emotional history data.
[0407] A "voice synthesis engine" is software or hardware that synthesizes music under specified conditions based on acquired voice data and musical score data.
[0408] An "emotion engine" is software or hardware that recognizes emotions from a user's voice or text input and generates emotion data.
[0409] "Recommended content" is information about singers and songs that the system determines to be appropriate based on the user's emotional data.
[0410] "Playlist" is a function for saving and playing a list of multiple songs selected by the user.
[0411] The present invention is a system that plays back singer-music combinations freely selected by the user in real time, and also recognizes the user's emotions to provide recommended content. This system is composed of a user terminal, a server, a database, a speech synthesis engine, and an emotion engine.
[0412] Overall system configuration
[0413] 1. User Device
[0414] The user terminal is a device such as a smartphone, tablet, or PC, and provides a user interface through which users can search for and select singers and songs. It also has the function of sending the user's voice and text input to the emotion engine.
[0415] 2. Server
[0416] The server is responsible for processing the singer and song selection information received from the user device, as well as the emotion data from the emotion engine. The server analyzes the received data and retrieves the singer's voice data and song score data from a database. It also sends a request to a speech synthesis engine (e.g., Google Cloud Text-to-Speech API or Amazon Polly) and sends the synthesized voice data to the user device. It also has the function of generating recommended content based on the emotion data.
[0417] 3. Database
[0418] The database stores the singer's voice data, the music score data, and the user's emotion history data. After the server receives the request, it retrieves the necessary data from the database.
[0419] 4. Speech synthesis engine
[0420] The voice synthesis engine uses the acquired singer's voice data and the music score data to synthesize a song according to the specified conditions. The synthesized voice data is sent to the user's device via the server.
[0421] 5. Emotion Engine
[0422] The emotion engine recognizes emotions from the user's voice or text and sends the emotion data to the server. The recognized emotion data is stored in a database along with the user's emotion history data and is used to generate recommended content.
[0423] Program processing and specific examples
[0424] The program helps users with a series of operations, from logging in to playing songs, recognizing emotions, providing recommended content, and sharing.
[0425] Specific examples
[0426] If a user wants to listen to "Song B" with the voice of "Singer A":
[0427] 1. User login
[0428] The user enters the information required to log in to the app (user ID, password, etc.) and is authenticated. If authentication is successful, the user is redirected to the home screen.
[0429] 2. Search and select an artist and song
[0430] The user enters "Singer A" in the search bar and selects "Singer A" from the displayed candidates. Next, the user enters "Song B" in the song search bar and selects "Song B" from the displayed candidates.
[0431] 3. Data Acquisition and Speech Synthesis
[0432] The server processes the combination of "Singer A" and "Song B" and retrieves the respective data from the database. This data is passed to the voice synthesis engine, which synthesizes the song according to the specified conditions. The resulting voice data is returned to the server and sent to the user's device.
[0433] 4. Music playback and emotion recognition
[0434] The user clicks on the link provided by the device to play the music, and during playback, the user's voice and text input are analyzed by the emotion engine.
[0435] 5. Providing recommended content
[0436] The server recommends singers and songs suitable for the user based on the emotion data received from the emotion engine.
[0437] 6. Create and share playlists
[0438] Users can add the songs they play to a playlist, save the playlist, and share the link to the playlist via social media or email.
[0439] This system allows users to enjoy new musical experiences and provides a personalized musical experience that is tailored to their emotions.
[0440] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0441] Step 1:
[0442] User Login
[0443] The user launches the app, enters their user ID and password, and clicks the "Login" button. The device sends this information to the server, which then authenticates them by checking it against the user information in its database. If authentication is successful, the server sends the home screen data to the device, and the device displays the home screen.
[0444] Input: User ID, Password
[0445] Output: Home screen display
[0446] Specific operation: The device sends input to the server, the server performs authentication, generates a home screen and sends it to the device, which then displays the home screen.
[0447] Step 2:
[0448] Search and select singers and songs
[0449] The user enters "Singer A" in the search bar, and the device sends this query to the server. The server retrieves related singer information from the database and sends it to the device. The user selects "Singer A" from the displayed candidates, then enters "Song B" in the song search bar to search and select a song in the same way. This selection information is sent to the server.
[0450] Input: Artist name, song name
[0451] Output: Artist and song selection information
[0452] Specific operation: The device sends a search query to the server, the server retrieves candidates from the database and sends them to the device, the user enters selection information, which is sent to the server.
[0453] Step 3:
[0454] Data Acquisition and Speech Synthesis
[0455] The server retrieves voice data and music score data from a database based on the received singer name and song title. This data is passed to a voice synthesis engine, which synthesizes music according to the specified conditions. The resulting voice data is returned to the server and sent to the user's device.
[0456] Input: Artist name, song name
[0457] Output: Synthesized speech data
[0458] Specific operation: The server retrieves the necessary data from the database, sends a request to the speech synthesis engine to synthesize speech, and sends the result to the user's device.
[0459] Step 4:
[0460] Music playback and emotion recognition
[0461] The user clicks on the music playback link from their device. The device plays the music data and sends the user's voice and text input during playback to the emotion engine. The emotion engine analyzes this and sends emotional data to the server.
[0462] Input: Music link, voice or text input
[0463] Output: Emotion data
[0464] Specific operation: The device plays music, sends the data entered by the user to the emotion engine, and sends the analysis results to the server.
[0465] Step 5:
[0466] Providing recommended content
[0467] The server references the database based on the emotion data received from the emotion engine to obtain suitable singers and song candidates. It then generates a list of recommended content and sends it to the user's device. The device then displays the recommended content.
[0468] Input: Emotion data
[0469] Output: Recommended content list
[0470] Specific operation: The server analyzes the emotion data, generates recommended content, and sends it to the device. The device then displays the recommended content.
[0471] Step 6:
[0472] Create and share playlists
[0473] The user adds the songs they played to a playlist and accesses the playlist creation screen. The device sends the list of songs selected by the user to the server, which then saves the playlist information in a database and generates a playlist ID. The generated link is provided to the user's device, and the user can share it on social media, etc.
[0474] Input: Song list
[0475] Output: Playlist link
[0476] Specific operation: The device sends the song list to the server, which stores it in a database, generates a link, and sends it to the device. The user can then use the link to share the playlist.
[0477] (Application example 2)
[0478] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0479] Current music distribution services limit users to simply selecting and playing songs, and do not adequately provide a personalized music experience based on their emotions. Furthermore, few services recognize users' input emotions and recommend content based on those emotions, making it difficult to provide a music experience that best suits the user's emotions. The objective of this invention is to provide a system that not only plays singers and songs freely selected by the user, but also recognizes the user's emotions in real time and provides recommended content based on those emotions.
[0480] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0481] In this invention, the server includes means for processing singer and song selection information received from a user terminal, means for retrieving singer voice data and song score data from a database based on the received selection information, means for synthesizing music using a voice synthesis engine using the retrieved voice data and song score data, means for transmitting the synthesized voice data to the user terminal, means for processing emotion data input by the user and recommending related information, means for recommending and providing content such as news and music based on the emotion data, and means for recognizing the emotion input by the user, thereby enabling a personalized music experience that matches the user's emotions.
[0482] "User terminal" means a device that a user accesses and uses to input and make selections.
[0483] "Selection information" is information about the singers and songs selected by the user.
[0484] "Audio data" refers to data that stores the voice of a particular singer in digital form.
[0485] "Music score data" refers to data that stores the melody and chords of a particular piece of music in digital format.
[0486] A "voice synthesis engine" is software or a system that generates new music using voice data and musical score data.
[0487] "Emotional data" refers to data about emotions derived from user input and actions.
[0488] "Recommended content" refers to music and information provided based on a user's emotional data.
[0489] A "search interface" is a screen or input field that allows a user to search for a particular artist or song.
[0490] A "playlist" is a function that organizes songs into a list, making it easy to play them consecutively and share them.
[0491] A "unique link" is a URL or hyperlink that provides access to a specific playlist or content.
[0492] This invention is a system that plays back singer-music combinations freely selected by the user in real time, and also recognizes the user's emotions to provide recommended content. The system is composed of a user terminal, a server, a database, a speech synthesis engine, and an emotion engine.
[0493] Overall system configuration
[0494] 1. User Device
[0495] The user terminal is a device such as a smartphone, tablet, or PC, and provides a user interface. From the user interface, users can search for and select singers and songs. It also has the function of sending the user's voice and text input to the emotion engine.
[0496] 2. Server
[0497] The server is responsible for processing the singer and song selection information received from the user device, as well as the emotion data from the emotion engine. The server analyzes the received data and retrieves the singer's voice data and song score data from the database. It also sends requests to the voice synthesis engine and transmits the synthesized voice data to the user device. It also has the function of generating recommended content based on the emotion data.
[0498] 3. Database
[0499] The database stores the singer's voice data, the music score data, and the user's emotion history data. After the server receives the request, it retrieves the necessary data from the database.
[0500] 4. Speech synthesis engine
[0501] The voice synthesis engine uses the acquired singer's voice data and the music score data to synthesize a song according to the specified conditions, and the synthesized voice data is sent to the user's device via the server.
[0502] 5. Emotion Engine
[0503] The emotion engine recognizes emotions from the user's voice or text and sends the emotion data to the server. The recognized emotion data is stored in a database along with the user's emotion history data and used to generate recommended content.
[0504] Program processing and specific examples
[0505] The program supports a series of operations, from user login to playing music, recognizing emotions, providing recommended content, and sharing. While the program's specific processing steps are not included, the following concrete examples are provided:
[0506] Specific examples
[0507] 1. Log in
[0508] The user logs in to the app. They enter their username and password, and the server authenticates them. If authentication is successful, the user is taken to the home screen.
[0509] 2. Selection of singers and songs
[0510] Users can enter a specific artist in the search bar and choose from a list of artists, then enter a song name in the song search bar and choose from a list of songs.
[0511] 3. Data Acquisition and Speech Synthesis
[0512] The server processes the selection information, retrieves the respective data from the database, and passes this data to the speech synthesis engine to synthesize the music.
[0513] 4. Music playback and emotion recognition
[0514] The server sends the synthesized voice data to the user's device. The user clicks the provided link and plays the music. During playback, the user's voice and text input are analyzed by the emotion engine.
[0515] 5. Providing recommended content
[0516] The emotion data analyzed by the emotion engine is sent to the server, which then recommends singers and songs that are suitable for the user based on the emotion data.
[0517] 6. Create and share playlists
[0518] Users can save the songs they play as playlists and share them on social media. Users can easily share music with friends and family using the playlist link.
[0519] Prompt Sentence Examples
[0520] "User text input: 'I'm feeling a bit down...'"
[0521] "The generative AI model's output: 'That's sad. Let's pick a song that's uplifting.'"
[0522] In this way, the present invention is a system that provides recommended content that corresponds to the user's emotions, thereby realizing a personalized music experience.
[0523] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0524] Step 1:
[0525] A user logs in to the app.
[0526] Input: User ID, Password
[0527] Output: Authentication result (success / failure), home screen
[0528] Specific operation: The user enters their user ID and password on the login screen and sends them to the server. The server retrieves the relevant information from the database and compares it with the entered information. If authentication is successful, the user is taken to the home screen; if it is unsuccessful, an error message is displayed.
[0529] Step 2:
[0530] The user searches for and selects an artist and song.
[0531] Input: Search keyword (singer name, song name)
[0532] Output: Search result list, selection information
[0533] Specific operation: The user enters the name of an artist or song in the search bar and presses the Enter key. The device sends the input information to the server, which retrieves the corresponding artist and song information from the database and sends a list of search results to the user's device. The user then selects a specific artist and song from the list.
[0534] Step 3:
[0535] The server retrieves the necessary data from the database based on the selection information.
[0536] Input: Selection information (singer name, song name)
[0537] Output: Singer's voice data, music score data
[0538] Specific operation: The server analyzes the selection information received from the user terminal and sends a request to the database, which then returns the corresponding singer's voice data and the music score data to the server.
[0539] Step 4:
[0540] The server synthesizes the music using a speech synthesis engine.
[0541] Input: Audio data, music score data
[0542] Output: Synthesized speech data
[0543] Specific operation: The server passes the acquired voice data and musical score data to the speech synthesis engine, which synthesizes the music. The speech synthesis engine processes the data and generates new synthetic voice data, which is then sent back to the server.
[0544] Step 5:
[0545] The synthesized speech data is transmitted to the user terminal.
[0546] Input: Synthetic speech data
[0547] Output: Music playback link
[0548] Specific operation: The server sends the synthesized voice data to the user's device and provides a music playback link. The user can click the link to play the synthesized music in real time.
[0549] Step 6:
[0550] The emotion engine analyzes the user's voice and text input.
[0551] Input: User voice or text input
[0552] Output: Emotion data
[0553] Specific operation: The user inputs voice or text into the emotion engine while playing music. The emotion engine receives the data, analyzes it, and generates emotion data as a result. The emotion data is then sent to the server.
[0554] Step 7:
[0555] Based on the emotion data, the server provides recommended content.
[0556] Input: Emotion data
[0557] Output: Recommended singer and song information
[0558] Specific operation: The server analyzes the emotion data and generates recommended content based on it. Specifically, it selects appropriate singers and songs from the database and sends that information to the user's device.
[0559] Step 8:
[0560] Users can save the songs they play as a playlist and share it on social media.
[0561] Input: Information about the song played
[0562] Output: Playlist link
[0563] Specific operation: The user selects the songs they played and saves them as a playlist. The server saves the playlist information in a database and generates a dedicated link. The generated link is sent to the user's device, and the user can share it via social media or with other users.
[0564] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0565] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0566] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0567] [Second embodiment]
[0568] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0569] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0570] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0571] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0572] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0573] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0574] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0575] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0576] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0577] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0578] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0579] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0580] The present invention is a system that allows users to freely select singers and song combinations, play them in real time, and share them. This system is composed of a user terminal, a server, a database, and a voice synthesis engine.
[0581] Overall system configuration
[0582] 1. User Device
[0583] The user terminal is a device such as a smartphone, tablet, or PC, and provides a user interface through which the user can search for and select singers and songs.
[0584] 2. Server
[0585] The server is responsible for processing the singer and song selection information received from the user device. The server analyzes the received data, retrieves the singer's voice data and song score data from the database, sends a request to the voice synthesis engine, and transmits the synthesized voice data to the user device.
[0586] 3. Database
[0587] The database stores the singer's voice data and the music score data. After the server receives the request, it retrieves the necessary data from the database.
[0588] 4. Speech synthesis engine
[0589] The voice synthesis engine uses the acquired singer's voice data and the music score data to synthesize a song according to the specified conditions. The synthesized voice data is sent to the user's device via the server.
[0590] Program processing and specific examples
[0591] Program processing
[0592] The program supports users in a series of operations from logging in to playing and sharing songs. The specific processing steps of the program are as follows:
[0593] 1. User login
[0594] The user enters the information required to log in to the app (user ID, password, etc.) and is authenticated. If authentication is successful, the user is redirected to the home screen.
[0595] 2. Search and select an artist and song
[0596] The user searches for singers and songs using keywords in the search interface and selects from the corresponding candidates. The selected information is sent to the server.
[0597] 3. Data Acquisition and Speech Synthesis
[0598] The server retrieves the singer's voice data and the song's score data from a database based on the received selection information. The retrieved data is passed to a voice synthesis engine, which synthesizes the song according to the specified conditions. The resulting voice data is returned to the server and sent to the user's device.
[0599] 4. Playing Music and Creating Playlists
[0600] Users can play the synthesized music using the link sent to them, and can also save multiple songs as a playlist and share it on social media.
[0601] Specific examples
[0602] For example, suppose a user wants to listen to "Song B" in the voice of "Singer A."
[0603] 1. Log in
[0604] A user logs in to the app.
[0605] Enter your username and password and the server will authenticate you.
[0606] If the authentication is successful, the user will be taken to the home screen.
[0607] 2. Selection of singers and songs
[0608] The user enters "singer A" in the search bar and selects "singer A" from the list of singers.
[0609] Next, enter "Song B" in the song search bar and select "Song B" from the list of songs.
[0610] 3. Data Acquisition and Speech Synthesis
[0611] The server processes the combination of "singer A" and "song B" and retrieves the respective data from the database.
[0612] The server passes this data to a voice synthesis engine, which synthesizes music according to the specified conditions.
[0613] 4. Playing songs
[0614] The server transmits the synthesized voice data to the user terminal.
[0615] The user plays "Song B" with the voice of "Singer A" on the device and enjoys the music.
[0616] 5. Create and share playlists
[0617] Users can save the songs they play as a playlist and share it on social media.
[0618] Users can easily share music with friends and family using the links of their saved playlists.
[0619] The present invention is a system that can provide users with a new way of enjoying music.
[0620] The processing flow will be explained below.
[0621] Step 1:
[0622] A user logs in to the app
[0623] The user launches the app on the device and accesses the login screen.
[0624] The terminal receives the user ID and password entered by the user.
[0625] The terminal sends this login information to the server.
[0626] The server checks the received user ID and password in a database and returns the authentication result to the terminal.
[0627] If the authentication is successful, the terminal transitions the user to the home screen.
[0628] Step 2:
[0629] User searches and selects singer
[0630] The user enters the singer's name into the device's search bar.
[0631] The terminal sends the entered name to the server.
[0632] The server searches the database for a list of singers based on the received singer name and returns the list to the terminal.
[0633] The user selects a particular singer from a list of singers displayed on the terminal.
[0634] The terminal transmits information about the selected singer to the server.
[0635] Step 3:
[0636] User searches and selects a song
[0637] The user enters the title of the song into the device's search bar.
[0638] The terminal transmits the input song title to the server.
[0639] The server searches the database for a list of songs based on the received song title and returns it to the terminal.
[0640] The user selects a specific song from the song list displayed on the terminal.
[0641] The terminal transmits information about the selected song to the server.
[0642] Step 4:
[0643] The server retrieves the data
[0644] The server analyzes the received selection information (singer name and song title).
[0645] The server retrieves the singer's voice data and the music score data from the database.
[0646] The acquired data is temporarily stored on the server.
[0647] Step 5:
[0648] Perform speech synthesis
[0649] The server passes the acquired voice data and musical score data to the voice synthesis engine.
[0650] The speech synthesis engine generates synthetic music based on this data.
[0651] The synthesized music file is sent back to the server.
[0652] Step 6:
[0653] Sending synthesized voice data
[0654] The server transmits the synthesized voice data to the user terminal as streaming or a download link.
[0655] The terminal receives this and makes it available for playback by the user.
[0656] Step 7:
[0657] User plays music
[0658] The user clicks on the link provided by the device and plays the music.
[0659] The device streams the audio data and plays it back in real time.
[0660] Step 8:
[0661] Create and share playlists
[0662] After listening to a number of songs, the user accesses a playlist creation screen.
[0663] The terminal sends a request to the server to save the list of songs selected by the user as a playlist.
[0664] The server stores the playlist information in a database and generates a playlist ID.
[0665] A link is created based on the generated playlist ID and sent to the device.
[0666] Users can share this link on social media and with other users.
[0667] Example 1
[0668] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0669] With conventional systems, it was difficult for users to play back their desired singer-song combinations in real time and easily share them. Furthermore, the process of synthesizing audio data and sheet music data was complex and time-consuming, which sometimes compromised the user experience. This resulted in a lack of flexibility and convenience for proposing new ways to enjoy music.
[0670] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0671] In this invention, the server includes means for processing selection information of voice data and music data received from a user terminal, means for acquiring voice data and music score data from a storage device based on the received selection information, means for synthesizing music using the acquired voice data and music score data with a voice synthesizer, and means for transmitting the synthesized voice data to the user terminal, thereby enabling users to freely select combinations of singers and music to be played in real time and easily shared.
[0672] A "user terminal" is a device that allows a user to interact with the system through an interface and perform various operations, and specifically refers to a smartphone, tablet, PC, etc.
[0673] "Selection Information" refers to data about a particular artist and song that a user enters through the system.
[0674] A "server" refers to a computer system that receives requests from user terminals and processes data and provides services in response to those requests.
[0675] "Audio data" refers to acoustic information that is a digital recording of the voice of a particular singer.
[0676] "Musical score data" refers to digital data that records musical information such as the melody, rhythm, and chords of a song.
[0677] "Storage device" refers to hardware and software for storing digital information such as audio data, sheet music data, and song selection lists.
[0678] A "speech synthesizer" refers to hardware or software that uses analog or digital voice data to generate new voices.
[0679] "Synthesized voice data" refers to music data generated by a voice synthesizer based on the voice data of a singer selected by the user and music score data.
[0680] A "search screen" refers to a user interface that provides an interface for a user to search for a specific singer or song.
[0681] "Song selection list" refers to a data list for managing multiple songs selected and saved by the user.
[0682] "Information sharing service" means an online platform that enables users to share information with other users via the Internet.
[0683] "Dedicated link" refers to a URL or URI generated to directly access a specific song selection list or voice synthesis data.
[0684] The present invention provides a system that allows users to freely select and share combinations of voice data and music data in real time. This system is comprised of a user terminal, a server, a storage device, and a voice synthesizer.
[0685] Overall system configuration
[0686] 1. User Device
[0687] The user terminal is a device such as a smartphone, tablet, or PC, and provides a user interface. The user can search and select audio data and music data through a search screen, allowing the user to easily find the music they want.
[0688] 2. Server
[0689] The server is responsible for processing the voice data and music data selection information received from the user terminal. The server analyzes the received data and retrieves the voice data and music score data from the storage device. It also sends a request to the voice synthesizer and sends the synthesized voice data to the user terminal. This allows the user to generate the music they want in real time.
[0690] 3. Storage device
[0691] The storage device stores the voice data and the musical score data. After the server receives the request, it retrieves the necessary data from the storage device. This data is provided to the voice synthesizer and used to synthesize the music.
[0692] 4. Speech synthesizer
[0693] The speech synthesizer uses the acquired voice data and musical score data to synthesize music under specified conditions. The synthesized voice data is sent to the user's terminal via a server. The speech synthesizer can generate music in real time.
[0694] Specific examples
[0695] Below is a specific example where a user wants to listen to "Song B" in the voice of "Singer A."
[0696] 1. Log in
[0697] A user logs in to a smartphone app.
[0698] Enter your username and password and the server will authenticate you.
[0699] If the authentication is successful, the user will be taken to the home screen.
[0700] 2. Selecting audio and music data
[0701] The user enters "singer A" in the search bar and selects "singer A" from the list of audio data.
[0702] Next, enter "Song B" in the song search bar and select "Song B" from the list of song data.
[0703] 3. Data Acquisition and Speech Synthesis
[0704] The server processes the combination of "singer A" and "song B" and retrieves the respective data from the storage device.
[0705] The server passes this data to a voice synthesizer, which synthesizes music according to the specified conditions.
[0706] 4. Playing songs
[0707] The server transmits the synthesized voice data to the user terminal.
[0708] The user plays "Song B" with the voice of "Singer A" on the device and enjoys the music.
[0709] 5. Create and share song lists
[0710] The songs played by the user are saved as a song selection list and shared via an information sharing service.
[0711] Users can easily share music with friends and family using links to their saved selections.
[0712] Prompt Sentence Examples
[0713] This is a system that generates music in real time using "Singer A" and "Song B." Specifically, a user logs into the app and enters "Singer A" and "Song B" in the search bar. The data for the selected singer and song is sent to a server, which then retrieves the necessary voice data and musical score data from a storage device and passes them to a voice synthesizer to synthesize the song. The synthesized song is sent to the user's device, where the user can play it.
[0714] With the above configuration, this system can provide users with a new way of enjoying music.
[0715] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0716] Step 1: The user launches the app and enters their user ID and password on the login screen.
[0717] Input: User ID, Password
[0718] Output: Authentication request
[0719] Specific operation: When the user taps the login button, the entered information is sent to the server.
[0720] Step 2: The server processes the received authentication information and checks it against a database.
[0721] Input: User ID, Password
[0722] Output: Authentication result
[0723] Specific operation: The server checks the user ID and password against the database, and if authentication is successful, generates a response indicating successful authentication.
[0724] Step 3: If the authentication is successful, the server sends an instruction to the user terminal to transition to the home screen.
[0725] Input: Authentication result (success)
[0726] Output: Home screen display instruction
[0727] Specific operation: Information indicating access rights to the home screen is sent to the user terminal, and the home screen is displayed on the user terminal.
[0728] Step 4: The user enters keywords for audio data (e.g., "singer A") and song data (e.g., "song B") into the search bar.
[0729] Input: Keywords for audio data and music data
[0730] Output: Search request
[0731] Specific operation: When the user taps the search button, the entered keywords are sent to the server.
[0732] Step 5: The server searches the storage device for related audio data and music data based on the received search keyword.
[0733] Input: Search keywords for audio data and music data
[0734] Output: Search results
[0735] Specific operation: The server searches the storage device and obtains the corresponding audio data and music data.
[0736] Step 6: The server sends the search results to the user terminal, and the user selects the appropriate singer and song.
[0737] Input: Search results
[0738] Output: User's choice
[0739] Specific operation: The user taps to select the desired artist and song from the list of search results.
[0740] Step 7: The server receives the user's selection information and retrieves the corresponding audio data and music score data from the storage device.
[0741] Input: User selection information
[0742] Output: Audio data and sheet music data
[0743] Specific operation: The server accesses the storage device and acquires the selected audio data and music score data.
[0744] Step 8: The server passes the acquired data to a voice synthesizer, which synthesizes music according to the specified conditions.
[0745] Input: Audio data and sheet music data
[0746] Output: Synthesized voice data
[0747] Specific operation: The voice synthesizer synthesizes music according to the specified conditions, and the generated voice data is sent back to the server.
[0748] Step 9: The server transmits the synthesized voice data to the user terminal.
[0749] Input: Synthesized voice data
[0750] Output: Playback link for audio data
[0751] Specific operation: The synthesized voice data is sent to the user's terminal and a playback link is displayed.
[0752] Step 10: The user taps the play link to play the synthesized voice data.
[0753] Input: Audio data playback link
[0754] Output: Playing music
[0755] Specific operation: The user device opens the link and plays the synthesized voice data.
[0756] Step 11: The user selects multiple songs, creates a song selection list, and shares it via an information sharing service.
[0757] Input: Song list
[0758] Output: Link to song selection list
[0759] Specific operation: The song list is saved in a storage device, and a link is generated and displayed on the user's device. The user can then share the link on social media.
[0760] (Application example 1)
[0761] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0762] Conventional music distribution services have difficulty generating and playing combinations of singers and songs freely selected by users in real time, limiting the variety of music experiences they can offer. Furthermore, there is a lack of systems that allow users to easily share the music they create through social networking sites or messaging applications. A system that can solve these issues and provide users with a new music experience is needed.
[0763] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0764] In this invention, the server includes means for processing singer and song selection information received from a user terminal, means for retrieving singer voice data and song score data from a database based on the received selection information, means for synthesizing a song using a voice synthesis engine using the retrieved voice data and song score data, means for transmitting the synthesized voice data to the user terminal, and means for users to play and share the song using a link to the synthesized voice data. This allows users to freely create and play back combinations of singers and songs in real time, and to easily share the created songs via social networking sites and messaging applications.
[0765] A "user terminal" is a device used by a user to operate the system, and includes smartphones, tablets, PCs, etc.
[0766] "Singer and song selection information" is information input by the user to combine a specific singer with a song.
[0767] "Audio data" refers to data that digitally records the voice of a particular singer.
[0768] "Musical score data" refers to data in which the notes, chords, rhythms, etc. of a piece of music are recorded in digital format.
[0769] "Database" means a data storage system for storing and managing audio data and musical score data.
[0770] A "voice synthesis engine" is software that generates music by combining voice data and musical score data.
[0771] "Synthetic voice data" is data of a song sung in the voice of a specific singer, generated by a voice synthesis engine.
[0772] A "search interface" is a screen or input device that allows a user to search for and select an artist and song.
[0773] A "server" is a computer system that processes requests from user terminals and provides information in cooperation with a database and a speech synthesis engine.
[0774] A "link" is a URL or URI that a user uses to play back synthesized speech data.
[0775] "SNS" is an abbreviation for social networking service, a platform for users to share information online.
[0776] A "messaging application" is software that allows users to exchange messages and information in real time.
[0777] A "playlist" is a collection of songs saved by a user and organized for easy playback and sharing.
[0778] The present invention is a system that receives information on singer and song selection from a user terminal, synthesizes the song in real time using a specific voice synthesis engine, and provides the generated voice data to the user. Specifically, the present invention is implemented using a user terminal, a server, a database, and a voice synthesis engine.
[0779] User terminal
[0780] The user terminal consists of a smartphone, tablet, PC, etc., and is the device through which the user operates the system. The user terminal is provided with a search interface for searching and selecting singers and songs. It also has the functionality to receive a link to the synthesized voice data and play and share it.
[0781] server
[0782] The server processes the selection information received from the user's device and retrieves the necessary audio data and music score data from the database. Specifically, the server was built using the Python Flask framework and receives input from the search interface. The server then works with a speech synthesis engine to generate music using the retrieved data. It is then responsible for sending a link to the generated audio data to the user's device.
[0783] Database
[0784] The database is a data storage system for storing and managing singers' voice data and music score data. The server quickly retrieves the necessary data from the database based on the selection information received.
[0785] Text-to-speech engine
[0786] A speech synthesis engine is software that synthesizes music based on specified conditions using acquired voice data and musical score data. Specifically, it can generate music in real time using external services such as speech synthesis APIs.
[0787] Specific examples of user operations
[0788] The user launches the app and logs in. They enter their favorite singer and song in the search bar and select from the search results. For example, if they want to listen to "Song B" in the voice of "Singer A," they enter and select this. The server retrieves the necessary data from the database based on the selection and passes it to the speech synthesis engine. A link to the synthesized voice data is sent to the user's device, and the user can use this link to play and share the song.
[0789] Prompt Sentence Examples
[0790] User: Please create song B using singer A's voice.
[0791] System: Synthesized audio data URL: http: / / example.com / audio / 1234
[0792] This invention allows users to freely select singers and song combinations, generate and play them in real time, and easily share them via social networking sites and messaging applications, providing a new musical experience.
[0793] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0794] Step 1:
[0795] The user starts the application on their device, such as a smartphone or PC, and goes to the login screen. The user enters their user ID and password and presses the login button.
[0796] Input: User ID, Password
[0797] Action: User device sends authentication information to server
[0798] Output: Authentication result from the server (success or failure)
[0799] Step 2:
[0800] The server authenticates the user based on the received authentication information. If authentication is successful, the user is redirected to the home screen and the search interface is displayed.
[0801] Input: Authentication information (user ID, password)
[0802] How it works: The server checks the authentication information against a database and sends the result back to the user's device.
[0803] Output: Authentication result (home screen if successful, error message if unsuccessful)
[0804] Step 3:
[0805] The user searches for and selects an artist and song in the search interface, for example, by entering "artist A" and "song B" in the search bar and selecting from the matching results.
[0806] Input: Search keyword (singer name, song name)
[0807] Action: Sends selections from the search interface to the server
[0808] Output: The server receives the selection information and displays the results on the user's terminal.
[0809] Step 4:
[0810] The server acquires the singer's voice data and the musical score data of the song from the database based on the received selection information.
[0811] Input: Selection information (singer name, song name)
[0812] How it works: The server sends a request to the database and retrieves the required data.
[0813] Output: Audio data, music score data
[0814] Step 5:
[0815] The server passes the acquired voice data and musical score data to a voice synthesis engine, which synthesizes music according to the specified conditions.
[0816] Input: Audio data, music score data
[0817] How it works: The server calls the speech synthesis engine and sends a synthesis request.
[0818] Output: Synthesized audio data
[0819] Step 6:
[0820] The server transmits the synthesized voice data to the user terminal and provides the user with a link to the synthesized voice data.
[0821] Input: Synthetic speech data
[0822] Operation: The server stores the synthesized voice data, generates a link to it, and sends it to the user's device.
[0823] Output: Link to the synthesized speech data
[0824] Step 7:
[0825] Users can use the provided link to play the synthesized voice and share it via social media or messaging applications.
[0826] Input: Synthetic speech data link
[0827] What happens: The user receives the link, plays the audio, and offers sharing options.
[0828] Output: Play audio, share
[0829] By following these steps, users can freely create and play combinations of singers and songs in real time, and can easily share the created songs via social media or messaging applications.
[0830] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0831] The present invention is a system that plays back singer-music combinations freely selected by the user in real time, and also recognizes the user's emotions to provide recommended content. This system is composed of a user terminal, a server, a database, a speech synthesis engine, and an emotion engine.
[0832] Overall system configuration
[0833] 1. User Device
[0834] The user terminal is a device such as a smartphone, tablet, or PC, and provides a user interface. From the user interface, users can search for and select singers and songs. It also has the function of sending the user's voice and text input to the emotion engine.
[0835] 2. Server
[0836] The server is responsible for processing the singer and song selection information received from the user terminal, as well as the emotion data from the emotion engine. The server analyzes the received data and retrieves the singer's voice data and song score data from the database. It also sends requests to the voice synthesis engine and transmits the synthesized voice data to the user terminal. It also has the function of generating recommended content based on the emotion data.
[0837] 3. Database
[0838] The database stores the singer's voice data, the music score data, and the user's emotion history data. After the server receives the request, it retrieves the required data from the database.
[0839] 4. Speech synthesis engine
[0840] The voice synthesis engine uses the acquired singer's voice data and the music score data to synthesize a song according to the specified conditions. The synthesized voice data is sent to the user's device via the server.
[0841] 5. Emotion Engine
[0842] The emotion engine recognizes emotions from the user's voice or text and sends the emotion data to the server. The recognized emotion data is stored in a database along with the user's emotion history data and used to generate recommended content.
[0843] Program processing and specific examples
[0844] Program processing
[0845] The program supports users in a series of operations, from logging in to playing songs, recognizing emotions, providing recommended content, and sharing. The specific processing steps of the program are as follows:
[0846] 1. User login
[0847] The user enters the information required to log in to the app (user ID, password, etc.) and is authenticated. If authentication is successful, the user is redirected to the home screen.
[0848] 2. Search and select an artist and song
[0849] The user searches for singers and songs using keywords in the search interface and selects from the corresponding candidates. The selected information is sent to the server.
[0850] 3. Data Acquisition and Speech Synthesis
[0851] The server analyzes the received selection information (singer name and song title) and retrieves the singer's voice data and the song's score data from the database. The retrieved data is passed to a voice synthesis engine, which synthesizes the song according to the specified conditions. The resulting voice data is returned to the server and sent to the user's device.
[0852] 4. Music playback and emotion recognition
[0853] The user can use the sent link to play the synthesized music, while the user's voice and text input are analyzed by the emotion engine.
[0854] 5. Providing recommended content
[0855] Based on the emotional data analyzed by the emotion engine, the server presents recommended singers and songs to the user, allowing the user to have a music experience that best suits their current emotions.
[0856] 6. Create and share playlists
[0857] After listening to multiple songs, the user accesses the playlist creation screen. The device sends a request to the server to save the list of songs selected by the user as a playlist. The server saves the playlist information in a database and generates a playlist ID. A link is created based on the generated playlist ID and sent to the device. The user can share this link via social media or with other users.
[0858] Specific examples
[0859] For example, suppose a user wants to listen to "Song B" in the voice of "Singer A."
[0860] 1. Log in
[0861] A user logs in to the app.
[0862] Enter your username and password and the server will authenticate you.
[0863] If the authentication is successful, the user will be taken to the home screen.
[0864] 2. Selection of singers and songs
[0865] The user enters "singer A" in the search bar and selects "singer A" from the list of singers.
[0866] Next, enter "Song B" in the song search bar and select "Song B" from the list of songs.
[0867] 3. Data Acquisition and Speech Synthesis
[0868] The server processes the combination of "singer A" and "song B" and retrieves the respective data from the database.
[0869] The server passes this data to a voice synthesis engine, which synthesizes music according to the specified conditions.
[0870] 4. Music playback and emotion recognition
[0871] The server transmits the synthesized voice data to the user terminal.
[0872] The user clicks on the link provided by the device and plays the music.
[0873] During playback, the user's voice and text input are analyzed by the emotion engine.
[0874] 5. Providing recommended content
[0875] The emotion data analyzed by the emotion engine is sent to the server.
[0876] The server recommends singers and songs that suit the user based on the emotion data.
[0877] 6. Create and share playlists
[0878] Users can save the songs they play as a playlist and share it on social media.
[0879] Users can easily share music with friends and family using the links of their saved playlists.
[0880] This invention is a system that provides users with a new way of enjoying music and also realizes a personalized music experience that matches the user's emotions.
[0881] The processing flow will be explained below.
[0882] Step 1:
[0883] A user logs in to the app
[0884] The user launches the app on the device and accesses the login screen.
[0885] The terminal receives the user ID and password entered by the user.
[0886] The terminal sends this login information to the server.
[0887] The server checks the received user ID and password in a database and returns the authentication result to the terminal.
[0888] If the authentication is successful, the terminal transitions the user to the home screen.
[0889] Step 2:
[0890] User searches and selects singer
[0891] The user enters the singer's name into the device's search bar.
[0892] The terminal sends the entered name to the server.
[0893] The server searches the database for a list of singers based on the received singer name and returns the list to the terminal.
[0894] The user selects a particular singer from a list of singers displayed on the terminal.
[0895] The terminal transmits information about the selected singer to the server.
[0896] Step 3:
[0897] User searches and selects a song
[0898] The user enters the title of the song into the device's search bar.
[0899] The terminal transmits the input song title to the server.
[0900] The server searches the database for a list of songs based on the received song title and returns it to the terminal.
[0901] The user selects a specific song from the song list displayed on the terminal.
[0902] The terminal transmits information about the selected song to the server.
[0903] Step 4:
[0904] The server retrieves the data
[0905] The server analyzes the received selection information (singer name and song title).
[0906] The server retrieves the singer's voice data and the music score data from the database.
[0907] The acquired data is temporarily stored on the server.
[0908] Step 5:
[0909] Perform speech synthesis
[0910] The server passes the acquired voice data and musical score data to the voice synthesis engine.
[0911] The speech synthesis engine generates synthetic music based on this data.
[0912] The synthesized music file is sent back to the server.
[0913] Step 6:
[0914] Sending synthesized voice data
[0915] The server transmits the synthesized voice data to the user terminal as streaming or a download link.
[0916] The terminal receives this and makes it available for playback by the user.
[0917] Step 7:
[0918] User plays music
[0919] The user clicks on the link provided by the device and plays the music.
[0920] The device streams the audio data and plays it back in real time.
[0921] Step 8:
[0922] Emotion recognition processing
[0923] Users can enter text and voice messages while listening to music.
[0924] The device sends the user's voice and text input to the emotion engine.
[0925] The emotion engine analyzes the user's input, recognizes the emotion, and sends the result to the server.
[0926] Step 9:
[0927] Providing recommended content
[0928] The server searches the database for singers and songs that suit the user's emotions based on the emotion data received from the emotion engine.
[0929] The server generates a list of relevant singers and songs and transmits it to the user terminal.
[0930] The user can check the recommended content displayed on the terminal and select new songs.
[0931] Step 10:
[0932] Create and share playlists
[0933] After playing a number of songs, the user accesses the playlist creation screen.
[0934] The terminal sends a request to the server to save the list of songs selected by the user as a playlist.
[0935] The server stores the playlist information in a database and generates a playlist ID.
[0936] A link is created based on the generated playlist ID and sent to the device.
[0937] Users can share this link on social media and with other users.
[0938] Example 2
[0939] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0940] Conventional music playback systems have difficulty in playing a combination of singers and songs freely selected by the user in real time. Furthermore, they lack the functionality to recognize the user's emotions and provide appropriate recommended content, making it difficult to provide a music experience that best suits the user's current emotions. To solve this problem, a system is needed that can efficiently process user selection information and emotional data, synthesize music in real time, and make appropriate recommendations.
[0941] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0942] In this invention, the server includes means for processing singer and song selection information received from a user terminal, means for retrieving singer voice data and song score data from a database based on the received selection information, means for synthesizing a song using a voice synthesis engine using the retrieved voice data and song score data, means for transmitting the user's voice and text input to an emotion engine, means for recognizing the user's emotion data and transmitting it to the server to generate recommended content, and means for transmitting the recommended content to the user terminal. This allows the user to freely select singers and songs to be played in real time, and further allows the user to enjoy a personalized music experience based on their emotions.
[0943] A "user terminal" is a device that allows a user to access and operate the system through an interface, and includes smartphones, tablets, PCs, etc.
[0944] A "server" is a computer system that processes data received from a user terminal, retrieves necessary data from a database based on selection information, and performs appropriate processing.
[0945] "Singer and song selection information" is data relating to the singer name and song title selected by the user through the system.
[0946] The "database" is a data storage system that stores necessary information such as singer's voice data, music score data, and user's emotional history data.
[0947] A "voice synthesis engine" is software or hardware that synthesizes music under specified conditions based on acquired voice data and musical score data.
[0948] An "emotion engine" is software or hardware that recognizes emotions from a user's voice or text input and generates emotion data.
[0949] "Recommended content" is information about singers and songs that the system determines to be appropriate based on the user's emotional data.
[0950] "Playlist" is a function for saving and playing a list of multiple songs selected by the user.
[0951] The present invention is a system that plays back singer-music combinations freely selected by the user in real time, and also recognizes the user's emotions to provide recommended content. This system is composed of a user terminal, a server, a database, a speech synthesis engine, and an emotion engine.
[0952] Overall system configuration
[0953] 1. User Device
[0954] The user terminal is a device such as a smartphone, tablet, or PC, and provides a user interface through which users can search for and select singers and songs. It also has the function of sending the user's voice and text input to the emotion engine.
[0955] 2. Server
[0956] The server is responsible for processing the singer and song selection information received from the user device, as well as the emotion data from the emotion engine. The server analyzes the received data and retrieves the singer's voice data and song score data from a database. It also sends a request to a speech synthesis engine (e.g., Google Cloud Text-to-Speech API or Amazon Polly) and sends the synthesized voice data to the user device. It also has the function of generating recommended content based on the emotion data.
[0957] 3. Database
[0958] The database stores the singer's voice data, the music score data, and the user's emotion history data. After the server receives the request, it retrieves the necessary data from the database.
[0959] 4. Speech synthesis engine
[0960] The voice synthesis engine uses the acquired singer's voice data and the music score data to synthesize a song according to the specified conditions. The synthesized voice data is sent to the user's device via the server.
[0961] 5. Emotion Engine
[0962] The emotion engine recognizes emotions from the user's voice or text and sends the emotion data to the server. The recognized emotion data is stored in a database along with the user's emotion history data and is used to generate recommended content.
[0963] Program processing and specific examples
[0964] The program helps users with a series of operations, from logging in to playing songs, recognizing emotions, providing recommended content, and sharing.
[0965] Specific examples
[0966] If a user wants to listen to "Song B" with the voice of "Singer A":
[0967] 1. User login
[0968] The user enters the information required to log in to the app (user ID, password, etc.) and is authenticated. If authentication is successful, the user is redirected to the home screen.
[0969] 2. Search and select an artist and song
[0970] The user enters "Singer A" in the search bar and selects "Singer A" from the displayed candidates. Next, the user enters "Song B" in the song search bar and selects "Song B" from the displayed candidates.
[0971] 3. Data Acquisition and Speech Synthesis
[0972] The server processes the combination of "Singer A" and "Song B" and retrieves the respective data from the database. This data is passed to the voice synthesis engine, which synthesizes the song according to the specified conditions. The resulting voice data is returned to the server and sent to the user's device.
[0973] 4. Music playback and emotion recognition
[0974] The user clicks on the link provided by the device to play the music, and during playback, the user's voice and text input are analyzed by the emotion engine.
[0975] 5. Providing recommended content
[0976] The server recommends singers and songs suitable for the user based on the emotion data received from the emotion engine.
[0977] 6. Create and share playlists
[0978] Users can add the songs they play to a playlist, save the playlist, and share the link to the playlist via social media or email.
[0979] This system allows users to enjoy new musical experiences and provides a personalized musical experience that is tailored to their emotions.
[0980] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0981] Step 1:
[0982] User Login
[0983] The user launches the app, enters their user ID and password, and clicks the "Login" button. The device sends this information to the server, which then authenticates them by checking it against the user information in its database. If authentication is successful, the server sends the home screen data to the device, and the device displays the home screen.
[0984] Input: User ID, Password
[0985] Output: Home screen display
[0986] Specific operation: The device sends input to the server, the server performs authentication, generates a home screen and sends it to the device, which then displays the home screen.
[0987] Step 2:
[0988] Search and select singers and songs
[0989] The user enters "Singer A" in the search bar, and the device sends this query to the server. The server retrieves related singer information from the database and sends it to the device. The user selects "Singer A" from the displayed candidates, then enters "Song B" in the song search bar to search and select a song in the same way. This selection information is sent to the server.
[0990] Input: Artist name, song name
[0991] Output: Artist and song selection information
[0992] Specific operation: The device sends a search query to the server, the server retrieves candidates from the database and sends them to the device, the user enters selection information, which is sent to the server.
[0993] Step 3:
[0994] Data Acquisition and Speech Synthesis
[0995] The server retrieves voice data and music score data from a database based on the received singer name and song title. This data is passed to a voice synthesis engine, which synthesizes music according to the specified conditions. The resulting voice data is returned to the server and sent to the user's device.
[0996] Input: Artist name, song name
[0997] Output: Synthesized speech data
[0998] Specific operation: The server retrieves the necessary data from the database, sends a request to the speech synthesis engine to synthesize speech, and sends the result to the user's device.
[0999] Step 4:
[1000] Music playback and emotion recognition
[1001] The user clicks on the music playback link from their device. The device plays the music data and sends the user's voice and text input during playback to the emotion engine. The emotion engine analyzes this and sends emotional data to the server.
[1002] Input: Music link, voice or text input
[1003] Output: Emotion data
[1004] Specific operation: The device plays music, sends the data entered by the user to the emotion engine, and sends the analysis results to the server.
[1005] Step 5:
[1006] Providing recommended content
[1007] The server references the database based on the emotion data received from the emotion engine to obtain suitable singers and song candidates. It then generates a list of recommended content and sends it to the user's device. The device then displays the recommended content.
[1008] Input: Emotion data
[1009] Output: Recommended content list
[1010] Specific operation: The server analyzes the emotion data, generates recommended content, and sends it to the device. The device then displays the recommended content.
[1011] Step 6:
[1012] Create and share playlists
[1013] The user adds the songs they played to a playlist and accesses the playlist creation screen. The device sends the list of songs selected by the user to the server, which then saves the playlist information in a database and generates a playlist ID. The generated link is provided to the user's device, and the user can share it on social media, etc.
[1014] Input: Song list
[1015] Output: Playlist link
[1016] Specific operation: The device sends the song list to the server, which stores it in a database, generates a link, and sends it to the device. The user can then use the link to share the playlist.
[1017] (Application example 2)
[1018] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[1019] Current music distribution services limit users to simply selecting and playing songs, and do not adequately provide a personalized music experience based on their emotions. Furthermore, few services recognize users' input emotions and recommend content based on those emotions, making it difficult to provide a music experience that best suits the user's emotions. The objective of this invention is to provide a system that not only plays singers and songs freely selected by the user, but also recognizes the user's emotions in real time and provides recommended content based on those emotions.
[1020] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1021] In this invention, the server includes means for processing singer and song selection information received from a user terminal, means for retrieving singer voice data and song score data from a database based on the received selection information, means for synthesizing music using a voice synthesis engine using the retrieved voice data and song score data, means for transmitting the synthesized voice data to the user terminal, means for processing emotion data input by the user and recommending related information, means for recommending and providing content such as news and music based on the emotion data, and means for recognizing the emotion input by the user, thereby enabling a personalized music experience that matches the user's emotions.
[1022] "User terminal" means a device that a user accesses and uses to input and make selections.
[1023] "Selection information" is information about the singers and songs selected by the user.
[1024] "Audio data" refers to data that stores the voice of a particular singer in digital form.
[1025] "Music score data" refers to data that stores the melody and chords of a particular piece of music in digital format.
[1026] A "voice synthesis engine" is software or a system that generates new music using voice data and musical score data.
[1027] "Emotional data" refers to data about emotions derived from user input and actions.
[1028] "Recommended content" refers to music and information provided based on a user's emotional data.
[1029] A "search interface" is a screen or input field that allows a user to search for a particular artist or song.
[1030] A "playlist" is a function that organizes songs into a list, making it easy to play them consecutively and share them.
[1031] A "unique link" is a URL or hyperlink that provides access to a specific playlist or content.
[1032] This invention is a system that plays back singer-music combinations freely selected by the user in real time, and also recognizes the user's emotions to provide recommended content. The system is composed of a user terminal, a server, a database, a speech synthesis engine, and an emotion engine.
[1033] Overall system configuration
[1034] 1. User Device
[1035] The user terminal is a device such as a smartphone, tablet, or PC, and provides a user interface. From the user interface, users can search for and select singers and songs. It also has the function of sending the user's voice and text input to the emotion engine.
[1036] 2. Server
[1037] The server is responsible for processing the singer and song selection information received from the user device, as well as the emotion data from the emotion engine. The server analyzes the received data and retrieves the singer's voice data and song score data from the database. It also sends requests to the voice synthesis engine and transmits the synthesized voice data to the user device. It also has the function of generating recommended content based on the emotion data.
[1038] 3. Database
[1039] The database stores the singer's voice data, the music score data, and the user's emotion history data. After the server receives the request, it retrieves the necessary data from the database.
[1040] 4. Speech synthesis engine
[1041] The voice synthesis engine uses the acquired singer's voice data and the music score data to synthesize a song according to the specified conditions, and the synthesized voice data is sent to the user's device via the server.
[1042] 5. Emotion Engine
[1043] The emotion engine recognizes emotions from the user's voice or text and sends the emotion data to the server. The recognized emotion data is stored in a database along with the user's emotion history data and used to generate recommended content.
[1044] Program processing and specific examples
[1045] The program supports a series of operations, from user login to playing music, recognizing emotions, providing recommended content, and sharing. While the program's specific processing steps are not included, the following concrete examples are provided:
[1046] Specific examples
[1047] 1. Log in
[1048] The user logs in to the app. They enter their username and password, and the server authenticates them. If authentication is successful, the user is taken to the home screen.
[1049] 2. Selection of singers and songs
[1050] Users can enter a specific artist in the search bar and choose from a list of artists, then enter a song name in the song search bar and choose from a list of songs.
[1051] 3. Data Acquisition and Speech Synthesis
[1052] The server processes the selection information, retrieves the respective data from the database, and passes this data to the speech synthesis engine to synthesize the music.
[1053] 4. Music playback and emotion recognition
[1054] The server sends the synthesized voice data to the user's device. The user clicks the provided link and plays the music. During playback, the user's voice and text input are analyzed by the emotion engine.
[1055] 5. Providing recommended content
[1056] The emotion data analyzed by the emotion engine is sent to the server, which then recommends singers and songs that are suitable for the user based on the emotion data.
[1057] 6. Create and share playlists
[1058] Users can save the songs they play as playlists and share them on social media. Users can easily share music with friends and family using the playlist link.
[1059] Prompt Sentence Examples
[1060] "User text input: 'I'm feeling a bit down...'"
[1061] "The generative AI model's output: 'That's sad. Let's pick a song that's uplifting.'"
[1062] In this way, the present invention is a system that provides recommended content that corresponds to the user's emotions, thereby realizing a personalized music experience.
[1063] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1064] Step 1:
[1065] A user logs in to the app.
[1066] Input: User ID, Password
[1067] Output: Authentication result (success / failure), home screen
[1068] Specific operation: The user enters their user ID and password on the login screen and sends them to the server. The server retrieves the relevant information from the database and compares it with the entered information. If authentication is successful, the user is taken to the home screen; if it is unsuccessful, an error message is displayed.
[1069] Step 2:
[1070] The user searches for and selects an artist and song.
[1071] Input: Search keyword (singer name, song name)
[1072] Output: Search result list, selection information
[1073] Specific operation: The user enters the name of an artist or song in the search bar and presses the Enter key. The device sends the input information to the server, which retrieves the corresponding artist and song information from the database and sends a list of search results to the user's device. The user then selects a specific artist and song from the list.
[1074] Step 3:
[1075] The server retrieves the necessary data from the database based on the selection information.
[1076] Input: Selection information (singer name, song name)
[1077] Output: Singer's voice data, music score data
[1078] Specific operation: The server analyzes the selection information received from the user terminal and sends a request to the database, which then returns the corresponding singer's voice data and the music score data to the server.
[1079] Step 4:
[1080] The server synthesizes the music using a speech synthesis engine.
[1081] Input: Audio data, music score data
[1082] Output: Synthesized speech data
[1083] Specific operation: The server passes the acquired voice data and musical score data to the speech synthesis engine, which synthesizes the music. The speech synthesis engine processes the data and generates new synthetic voice data, which is then sent back to the server.
[1084] Step 5:
[1085] The synthesized speech data is transmitted to the user terminal.
[1086] Input: Synthetic speech data
[1087] Output: Music playback link
[1088] Specific operation: The server sends the synthesized voice data to the user's device and provides a music playback link. The user can click the link to play the synthesized music in real time.
[1089] Step 6:
[1090] The emotion engine analyzes the user's voice and text input.
[1091] Input: User voice or text input
[1092] Output: Emotion data
[1093] Specific operation: The user inputs voice or text into the emotion engine while playing music. The emotion engine receives the data, analyzes it, and generates emotion data as a result. The emotion data is then sent to the server.
[1094] Step 7:
[1095] Based on the emotion data, the server provides recommended content.
[1096] Input: Emotion data
[1097] Output: Recommended singer and song information
[1098] Specific operation: The server analyzes the emotion data and generates recommended content based on it. Specifically, it selects appropriate singers and songs from the database and sends that information to the user's device.
[1099] Step 8:
[1100] Users can save the songs they play as a playlist and share it on social media.
[1101] Input: Information about the song played
[1102] Output: Playlist link
[1103] Specific operation: The user selects the songs they played and saves them as a playlist. The server saves the playlist information in a database and generates a dedicated link. The generated link is sent to the user's device, and the user can share it via social media or with other users.
[1104] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1105] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1106] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[1107] [Third embodiment]
[1108] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[1109] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[1110] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1111] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[1112] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1113] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1114] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1115] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1116] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1117] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1118] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1119] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[1120] The present invention is a system that allows users to freely select singers and song combinations, play them in real time, and share them. This system is composed of a user terminal, a server, a database, and a voice synthesis engine.
[1121] Overall system configuration
[1122] 1. User Device
[1123] The user terminal is a device such as a smartphone, tablet, or PC, and provides a user interface through which the user can search for and select singers and songs.
[1124] 2. Server
[1125] The server is responsible for processing the singer and song selection information received from the user device. The server analyzes the received data, retrieves the singer's voice data and song score data from the database, sends a request to the voice synthesis engine, and transmits the synthesized voice data to the user device.
[1126] 3. Database
[1127] The database stores the singer's voice data and the music score data. After the server receives the request, it retrieves the necessary data from the database.
[1128] 4. Speech synthesis engine
[1129] The voice synthesis engine uses the acquired singer's voice data and the music score data to synthesize a song according to the specified conditions. The synthesized voice data is sent to the user's device via the server.
[1130] Program processing and specific examples
[1131] Program processing
[1132] The program supports users in a series of operations from logging in to playing and sharing songs. The specific processing steps of the program are as follows:
[1133] 1. User login
[1134] The user enters the information required to log in to the app (user ID, password, etc.) and is authenticated. If authentication is successful, the user is redirected to the home screen.
[1135] 2. Search and select an artist and song
[1136] The user searches for singers and songs using keywords in the search interface and selects from the corresponding candidates. The selected information is sent to the server.
[1137] 3. Data Acquisition and Speech Synthesis
[1138] The server retrieves the singer's voice data and the song's score data from a database based on the received selection information. The retrieved data is passed to a voice synthesis engine, which synthesizes the song according to the specified conditions. The resulting voice data is returned to the server and sent to the user's device.
[1139] 4. Playing Music and Creating Playlists
[1140] Users can play the synthesized music using the link sent to them, and can also save multiple songs as a playlist and share it on social media.
[1141] Specific examples
[1142] For example, suppose a user wants to listen to "Song B" in the voice of "Singer A."
[1143] 1. Log in
[1144] A user logs in to the app.
[1145] Enter your username and password and the server will authenticate you.
[1146] If the authentication is successful, the user will be taken to the home screen.
[1147] 2. Selection of singers and songs
[1148] The user enters "singer A" in the search bar and selects "singer A" from the list of singers.
[1149] Next, enter "Song B" in the song search bar and select "Song B" from the list of songs.
[1150] 3. Data Acquisition and Speech Synthesis
[1151] The server processes the combination of "singer A" and "song B" and retrieves the respective data from the database.
[1152] The server passes this data to a voice synthesis engine, which synthesizes music according to the specified conditions.
[1153] 4. Playing songs
[1154] The server transmits the synthesized voice data to the user terminal.
[1155] The user plays "Song B" with the voice of "Singer A" on the device and enjoys the music.
[1156] 5. Create and share playlists
[1157] Users can save the songs they play as a playlist and share it on social media.
[1158] Users can easily share music with friends and family using the links of their saved playlists.
[1159] The present invention is a system that can provide users with a new way of enjoying music.
[1160] The processing flow will be explained below.
[1161] Step 1:
[1162] A user logs in to the app
[1163] The user launches the app on the device and accesses the login screen.
[1164] The terminal receives the user ID and password entered by the user.
[1165] The terminal sends this login information to the server.
[1166] The server checks the received user ID and password in a database and returns the authentication result to the terminal.
[1167] If the authentication is successful, the terminal transitions the user to the home screen.
[1168] Step 2:
[1169] User searches and selects singer
[1170] The user enters the singer's name into the device's search bar.
[1171] The terminal sends the entered name to the server.
[1172] The server searches the database for a list of singers based on the received singer name and returns the list to the terminal.
[1173] The user selects a particular singer from a list of singers displayed on the terminal.
[1174] The terminal transmits information about the selected singer to the server.
[1175] Step 3:
[1176] User searches and selects a song
[1177] The user enters the title of the song into the device's search bar.
[1178] The terminal transmits the input song title to the server.
[1179] The server searches the database for a list of songs based on the received song title and returns it to the terminal.
[1180] The user selects a specific song from the song list displayed on the terminal.
[1181] The terminal transmits information about the selected song to the server.
[1182] Step 4:
[1183] The server retrieves the data
[1184] The server analyzes the received selection information (singer name and song title).
[1185] The server retrieves the singer's voice data and the music score data from the database.
[1186] The acquired data is temporarily stored on the server.
[1187] Step 5:
[1188] Perform speech synthesis
[1189] The server passes the acquired voice data and musical score data to the voice synthesis engine.
[1190] The speech synthesis engine generates synthetic music based on this data.
[1191] The synthesized music file is sent back to the server.
[1192] Step 6:
[1193] Sending synthesized voice data
[1194] The server transmits the synthesized voice data to the user terminal as streaming or a download link.
[1195] The terminal receives this and makes it available for playback by the user.
[1196] Step 7:
[1197] User plays music
[1198] The user clicks on the link provided by the device and plays the music.
[1199] The device streams the audio data and plays it back in real time.
[1200] Step 8:
[1201] Create and share playlists
[1202] After listening to a number of songs, the user accesses a playlist creation screen.
[1203] The terminal sends a request to the server to save the list of songs selected by the user as a playlist.
[1204] The server stores the playlist information in a database and generates a playlist ID.
[1205] A link is created based on the generated playlist ID and sent to the device.
[1206] Users can share this link on social media and with other users.
[1207] Example 1
[1208] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1209] With conventional systems, it was difficult for users to play back their desired singer-song combinations in real time and easily share them. Furthermore, the process of synthesizing audio data and sheet music data was complex and time-consuming, which sometimes compromised the user experience. This resulted in a lack of flexibility and convenience for proposing new ways to enjoy music.
[1210] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1211] In this invention, the server includes means for processing selection information of voice data and music data received from a user terminal, means for acquiring voice data and music score data from a storage device based on the received selection information, means for synthesizing music using the acquired voice data and music score data with a voice synthesizer, and means for transmitting the synthesized voice data to the user terminal, thereby enabling users to freely select combinations of singers and music to be played in real time and easily shared.
[1212] A "user terminal" is a device that allows a user to interact with the system through an interface and perform various operations, and specifically refers to a smartphone, tablet, PC, etc.
[1213] "Selection Information" refers to data about a particular artist and song that a user enters through the system.
[1214] A "server" refers to a computer system that receives requests from user terminals and processes data and provides services in response to those requests.
[1215] "Audio data" refers to acoustic information that is a digital recording of the voice of a particular singer.
[1216] "Musical score data" refers to digital data that records musical information such as the melody, rhythm, and chords of a song.
[1217] "Storage device" refers to hardware and software for storing digital information such as audio data, sheet music data, and song selection lists.
[1218] A "speech synthesizer" refers to hardware or software that uses analog or digital voice data to generate new voices.
[1219] "Synthesized voice data" refers to music data generated by a voice synthesizer based on the voice data of a singer selected by the user and music score data.
[1220] A "search screen" refers to a user interface that provides an interface for a user to search for a specific singer or song.
[1221] "Song selection list" refers to a data list for managing multiple songs selected and saved by the user.
[1222] "Information sharing service" means an online platform that enables users to share information with other users via the Internet.
[1223] "Dedicated link" refers to a URL or URI generated to directly access a specific song selection list or voice synthesis data.
[1224] The present invention provides a system that allows users to freely select and share combinations of voice data and music data in real time. This system is comprised of a user terminal, a server, a storage device, and a voice synthesizer.
[1225] Overall system configuration
[1226] 1. User Device
[1227] The user terminal is a device such as a smartphone, tablet, or PC, and provides a user interface. The user can search and select audio data and music data through a search screen, allowing the user to easily find the music they want.
[1228] 2. Server
[1229] The server is responsible for processing the voice data and music data selection information received from the user terminal. The server analyzes the received data and retrieves the voice data and music score data from the storage device. It also sends a request to the voice synthesizer and sends the synthesized voice data to the user terminal. This allows the user to generate the music they want in real time.
[1230] 3. Storage device
[1231] The storage device stores the voice data and the musical score data. After the server receives the request, it retrieves the necessary data from the storage device. This data is provided to the voice synthesizer and used to synthesize the music.
[1232] 4. Speech synthesizer
[1233] The speech synthesizer uses the acquired voice data and musical score data to synthesize music under specified conditions. The synthesized voice data is sent to the user's terminal via a server. The speech synthesizer can generate music in real time.
[1234] Specific examples
[1235] Below is a specific example where a user wants to listen to "Song B" in the voice of "Singer A."
[1236] 1. Log in
[1237] A user logs in to a smartphone app.
[1238] Enter your username and password and the server will authenticate you.
[1239] If the authentication is successful, the user will be taken to the home screen.
[1240] 2. Selecting audio and music data
[1241] The user enters "singer A" in the search bar and selects "singer A" from the list of audio data.
[1242] Next, enter "Song B" in the song search bar and select "Song B" from the list of song data.
[1243] 3. Data Acquisition and Speech Synthesis
[1244] The server processes the combination of "singer A" and "song B" and retrieves the respective data from the storage device.
[1245] The server passes this data to a voice synthesizer, which synthesizes music according to the specified conditions.
[1246] 4. Playing songs
[1247] The server transmits the synthesized voice data to the user terminal.
[1248] The user plays "Song B" with the voice of "Singer A" on the device and enjoys the music.
[1249] 5. Create and share song lists
[1250] The songs played by the user are saved as a song selection list and shared via an information sharing service.
[1251] Users can easily share music with friends and family using links to their saved selections.
[1252] Prompt Sentence Examples
[1253] This is a system that generates music in real time using "Singer A" and "Song B." Specifically, a user logs into the app and enters "Singer A" and "Song B" in the search bar. The data for the selected singer and song is sent to a server, which then retrieves the necessary voice data and musical score data from a storage device and passes them to a voice synthesizer to synthesize the song. The synthesized song is sent to the user's device, where the user can play it.
[1254] With the above configuration, this system can provide users with a new way of enjoying music.
[1255] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1256] Step 1: The user launches the app and enters their user ID and password on the login screen.
[1257] Input: User ID, Password
[1258] Output: Authentication request
[1259] Specific operation: When the user taps the login button, the entered information is sent to the server.
[1260] Step 2: The server processes the received authentication information and checks it against a database.
[1261] Input: User ID, Password
[1262] Output: Authentication result
[1263] Specific operation: The server checks the user ID and password against the database, and if authentication is successful, generates a response indicating successful authentication.
[1264] Step 3: If the authentication is successful, the server sends an instruction to the user terminal to transition to the home screen.
[1265] Input: Authentication result (success)
[1266] Output: Home screen display instruction
[1267] Specific operation: Information indicating access rights to the home screen is sent to the user terminal, and the home screen is displayed on the user terminal.
[1268] Step 4: The user enters keywords for audio data (e.g., "singer A") and song data (e.g., "song B") into the search bar.
[1269] Input: Keywords for audio data and music data
[1270] Output: Search request
[1271] Specific operation: When the user taps the search button, the entered keywords are sent to the server.
[1272] Step 5: The server searches the storage device for related audio data and music data based on the received search keyword.
[1273] Input: Search keywords for audio data and music data
[1274] Output: Search results
[1275] Specific operation: The server searches the storage device and obtains the corresponding audio data and music data.
[1276] Step 6: The server sends the search results to the user terminal, and the user selects the appropriate singer and song.
[1277] Input: Search results
[1278] Output: User's choice
[1279] Specific operation: The user taps to select the desired artist and song from the list of search results.
[1280] Step 7: The server receives the user's selection information and retrieves the corresponding audio data and music score data from the storage device.
[1281] Input: User selection information
[1282] Output: Audio data and sheet music data
[1283] Specific operation: The server accesses the storage device and acquires the selected audio data and music score data.
[1284] Step 8: The server passes the acquired data to a voice synthesizer, which synthesizes music according to the specified conditions.
[1285] Input: Audio data and sheet music data
[1286] Output: Synthesized voice data
[1287] Specific operation: The voice synthesizer synthesizes music according to the specified conditions, and the generated voice data is sent back to the server.
[1288] Step 9: The server transmits the synthesized voice data to the user terminal.
[1289] Input: Synthesized voice data
[1290] Output: Playback link for audio data
[1291] Specific operation: The synthesized voice data is sent to the user's terminal and a playback link is displayed.
[1292] Step 10: The user taps the play link to play the synthesized voice data.
[1293] Input: Audio data playback link
[1294] Output: Playing music
[1295] Specific operation: The user device opens the link and plays the synthesized voice data.
[1296] Step 11: The user selects multiple songs, creates a song selection list, and shares it via an information sharing service.
[1297] Input: Song list
[1298] Output: Link to song selection list
[1299] Specific operation: The song list is saved in a storage device, and a link is generated and displayed on the user's device. The user can then share the link on social media.
[1300] (Application example 1)
[1301] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1302] Conventional music distribution services have difficulty generating and playing combinations of singers and songs freely selected by users in real time, limiting the variety of music experiences they can offer. Furthermore, there is a lack of systems that allow users to easily share the music they create through social networking sites or messaging applications. A system that can solve these issues and provide users with a new music experience is needed.
[1303] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1304] In this invention, the server includes means for processing singer and song selection information received from a user terminal, means for retrieving singer voice data and song score data from a database based on the received selection information, means for synthesizing a song using a voice synthesis engine using the retrieved voice data and song score data, means for transmitting the synthesized voice data to the user terminal, and means for users to play and share the song using a link to the synthesized voice data. This allows users to freely create and play back combinations of singers and songs in real time, and to easily share the created songs via social networking sites and messaging applications.
[1305] A "user terminal" is a device used by a user to operate the system, and includes smartphones, tablets, PCs, etc.
[1306] "Singer and song selection information" is information input by the user to combine a specific singer with a song.
[1307] "Audio data" refers to data that digitally records the voice of a particular singer.
[1308] "Musical score data" refers to data in which the notes, chords, rhythms, etc. of a piece of music are recorded in digital format.
[1309] "Database" means a data storage system for storing and managing audio data and musical score data.
[1310] A "voice synthesis engine" is software that generates music by combining voice data and musical score data.
[1311] "Synthetic voice data" is data of a song sung in the voice of a specific singer, generated by a voice synthesis engine.
[1312] A "search interface" is a screen or input device that allows a user to search for and select an artist and song.
[1313] A "server" is a computer system that processes requests from user terminals and provides information in cooperation with a database and a speech synthesis engine.
[1314] A "link" is a URL or URI that a user uses to play back synthesized speech data.
[1315] "SNS" is an abbreviation for social networking service, a platform for users to share information online.
[1316] A "messaging application" is software that allows users to exchange messages and information in real time.
[1317] A "playlist" is a collection of songs saved by a user and organized for easy playback and sharing.
[1318] The present invention is a system that receives information on singer and song selection from a user terminal, synthesizes the song in real time using a specific voice synthesis engine, and provides the generated voice data to the user. Specifically, the present invention is implemented using a user terminal, a server, a database, and a voice synthesis engine.
[1319] User terminal
[1320] The user terminal consists of a smartphone, tablet, PC, etc., and is the device through which the user operates the system. The user terminal is provided with a search interface for searching and selecting singers and songs. It also has the functionality to receive a link to the synthesized voice data and play and share it.
[1321] server
[1322] The server processes the selection information received from the user's device and retrieves the necessary audio data and music score data from the database. Specifically, the server was built using the Python Flask framework and receives input from the search interface. The server then works with a speech synthesis engine to generate music using the retrieved data. It is then responsible for sending a link to the generated audio data to the user's device.
[1323] Database
[1324] The database is a data storage system for storing and managing singers' voice data and music score data. The server quickly retrieves the necessary data from the database based on the selection information received.
[1325] Text-to-speech engine
[1326] A speech synthesis engine is software that synthesizes music based on specified conditions using acquired voice data and musical score data. Specifically, it can generate music in real time using external services such as speech synthesis APIs.
[1327] Specific examples of user operations
[1328] The user launches the app and logs in. They enter their favorite singer and song in the search bar and select from the search results. For example, if they want to listen to "Song B" in the voice of "Singer A," they enter and select this. The server retrieves the necessary data from the database based on the selection and passes it to the speech synthesis engine. A link to the synthesized voice data is sent to the user's device, and the user can use this link to play and share the song.
[1329] Prompt Sentence Examples
[1330] User: Please create song B using singer A's voice.
[1331] System: Synthesized audio data URL: http: / / example.com / audio / 1234
[1332] This invention allows users to freely select singers and song combinations, generate and play them in real time, and easily share them via social networking sites and messaging applications, providing a new musical experience.
[1333] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1334] Step 1:
[1335] The user starts the application on their device, such as a smartphone or PC, and goes to the login screen. The user enters their user ID and password and presses the login button.
[1336] Input: User ID, Password
[1337] Action: User device sends authentication information to server
[1338] Output: Authentication result from the server (success or failure)
[1339] Step 2:
[1340] The server authenticates the user based on the received authentication information. If authentication is successful, the user is redirected to the home screen and the search interface is displayed.
[1341] Input: Authentication information (user ID, password)
[1342] How it works: The server checks the authentication information against a database and sends the result back to the user's device.
[1343] Output: Authentication result (home screen if successful, error message if unsuccessful)
[1344] Step 3:
[1345] The user searches for and selects an artist and song in the search interface, for example, by entering "artist A" and "song B" in the search bar and selecting from the matching results.
[1346] Input: Search keyword (singer name, song name)
[1347] Action: Sends selections from the search interface to the server
[1348] Output: The server receives the selection information and displays the results on the user's terminal.
[1349] Step 4:
[1350] The server acquires the singer's voice data and the musical score data of the song from the database based on the received selection information.
[1351] Input: Selection information (singer name, song name)
[1352] How it works: The server sends a request to the database and retrieves the required data.
[1353] Output: Audio data, music score data
[1354] Step 5:
[1355] The server passes the acquired voice data and musical score data to a voice synthesis engine, which synthesizes music according to the specified conditions.
[1356] Input: Audio data, music score data
[1357] How it works: The server calls the speech synthesis engine and sends a synthesis request.
[1358] Output: Synthesized audio data
[1359] Step 6:
[1360] The server transmits the synthesized voice data to the user terminal and provides the user with a link to the synthesized voice data.
[1361] Input: Synthetic speech data
[1362] Operation: The server stores the synthesized voice data, generates a link to it, and sends it to the user's device.
[1363] Output: Link to the synthesized speech data
[1364] Step 7:
[1365] Users can use the provided link to play the synthesized voice and share it via social media or messaging applications.
[1366] Input: Synthetic speech data link
[1367] What happens: The user receives the link, plays the audio, and offers sharing options.
[1368] Output: Play audio, share
[1369] By following these steps, users can freely create and play combinations of singers and songs in real time, and can easily share the created songs via social media or messaging applications.
[1370] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1371] The present invention is a system that plays back singer-music combinations freely selected by the user in real time, and also recognizes the user's emotions to provide recommended content. This system is composed of a user terminal, a server, a database, a speech synthesis engine, and an emotion engine.
[1372] Overall system configuration
[1373] 1. User Device
[1374] The user terminal is a device such as a smartphone, tablet, or PC, and provides a user interface. From the user interface, users can search for and select singers and songs. It also has the function of sending the user's voice and text input to the emotion engine.
[1375] 2. Server
[1376] The server is responsible for processing the singer and song selection information received from the user terminal, as well as the emotion data from the emotion engine. The server analyzes the received data and retrieves the singer's voice data and song score data from the database. It also sends requests to the voice synthesis engine and transmits the synthesized voice data to the user terminal. It also has the function of generating recommended content based on the emotion data.
[1377] 3. Database
[1378] The database stores the singer's voice data, the music score data, and the user's emotion history data. After the server receives the request, it retrieves the required data from the database.
[1379] 4. Speech synthesis engine
[1380] The voice synthesis engine uses the acquired singer's voice data and the music score data to synthesize a song according to the specified conditions. The synthesized voice data is sent to the user's device via the server.
[1381] 5. Emotion Engine
[1382] The emotion engine recognizes emotions from the user's voice or text and sends the emotion data to the server. The recognized emotion data is stored in a database along with the user's emotion history data and used to generate recommended content.
[1383] Program processing and specific examples
[1384] Program processing
[1385] The program supports users in a series of operations, from logging in to playing songs, recognizing emotions, providing recommended content, and sharing. The specific processing steps of the program are as follows:
[1386] 1. User login
[1387] The user enters the information required to log in to the app (user ID, password, etc.) and is authenticated. If authentication is successful, the user is redirected to the home screen.
[1388] 2. Search and select an artist and song
[1389] The user searches for singers and songs using keywords in the search interface and selects from the corresponding candidates. The selected information is sent to the server.
[1390] 3. Data Acquisition and Speech Synthesis
[1391] The server analyzes the received selection information (singer name and song title) and retrieves the singer's voice data and the song's score data from the database. The retrieved data is passed to a voice synthesis engine, which synthesizes the song according to the specified conditions. The resulting voice data is returned to the server and sent to the user's device.
[1392] 4. Music playback and emotion recognition
[1393] The user can use the sent link to play the synthesized music, while the user's voice and text input are analyzed by the emotion engine.
[1394] 5. Providing recommended content
[1395] Based on the emotional data analyzed by the emotion engine, the server presents recommended singers and songs to the user, allowing the user to have a music experience that best suits their current emotions.
[1396] 6. Create and share playlists
[1397] After listening to multiple songs, the user accesses the playlist creation screen. The device sends a request to the server to save the list of songs selected by the user as a playlist. The server saves the playlist information in a database and generates a playlist ID. A link is created based on the generated playlist ID and sent to the device. The user can share this link via social media or with other users.
[1398] Specific examples
[1399] For example, suppose a user wants to listen to "Song B" in the voice of "Singer A."
[1400] 1. Log in
[1401] A user logs in to the app.
[1402] Enter your username and password and the server will authenticate you.
[1403] If the authentication is successful, the user will be taken to the home screen.
[1404] 2. Selection of singers and songs
[1405] The user enters "singer A" in the search bar and selects "singer A" from the list of singers.
[1406] Next, enter "Song B" in the song search bar and select "Song B" from the list of songs.
[1407] 3. Data Acquisition and Speech Synthesis
[1408] The server processes the combination of "singer A" and "song B" and retrieves the respective data from the database.
[1409] The server passes this data to a voice synthesis engine, which synthesizes music according to the specified conditions.
[1410] 4. Music playback and emotion recognition
[1411] The server transmits the synthesized voice data to the user terminal.
[1412] The user clicks on the link provided by the device and plays the music.
[1413] During playback, the user's voice and text input are analyzed by the emotion engine.
[1414] 5. Providing recommended content
[1415] The emotion data analyzed by the emotion engine is sent to the server.
[1416] The server recommends singers and songs that suit the user based on the emotion data.
[1417] 6. Create and share playlists
[1418] Users can save the songs they play as a playlist and share it on social media.
[1419] Users can easily share music with friends and family using the links of their saved playlists.
[1420] This invention is a system that provides users with a new way of enjoying music and also realizes a personalized music experience that matches the user's emotions.
[1421] The processing flow will be explained below.
[1422] Step 1:
[1423] A user logs in to the app
[1424] The user launches the app on the device and accesses the login screen.
[1425] The terminal receives the user ID and password entered by the user.
[1426] The terminal sends this login information to the server.
[1427] The server checks the received user ID and password in a database and returns the authentication result to the terminal.
[1428] If the authentication is successful, the terminal transitions the user to the home screen.
[1429] Step 2:
[1430] User searches and selects singer
[1431] The user enters the singer's name into the device's search bar.
[1432] The terminal sends the entered name to the server.
[1433] The server searches the database for a list of singers based on the received singer name and returns the list to the terminal.
[1434] The user selects a particular singer from a list of singers displayed on the terminal.
[1435] The terminal transmits information about the selected singer to the server.
[1436] Step 3:
[1437] User searches and selects a song
[1438] The user enters the title of the song into the device's search bar.
[1439] The terminal transmits the input song title to the server.
[1440] The server searches the database for a list of songs based on the received song title and returns it to the terminal.
[1441] The user selects a specific song from the song list displayed on the terminal.
[1442] The terminal transmits information about the selected song to the server.
[1443] Step 4:
[1444] The server retrieves the data
[1445] The server analyzes the received selection information (singer name and song title).
[1446] The server retrieves the singer's voice data and the music score data from the database.
[1447] The acquired data is temporarily stored on the server.
[1448] Step 5:
[1449] Perform speech synthesis
[1450] The server passes the acquired voice data and musical score data to the voice synthesis engine.
[1451] The speech synthesis engine generates synthetic music based on this data.
[1452] The synthesized music file is sent back to the server.
[1453] Step 6:
[1454] Sending synthesized voice data
[1455] The server transmits the synthesized voice data to the user terminal as streaming or a download link.
[1456] The terminal receives this and makes it available for playback by the user.
[1457] Step 7:
[1458] User plays music
[1459] The user clicks on the link provided by the device and plays the music.
[1460] The device streams the audio data and plays it back in real time.
[1461] Step 8:
[1462] Emotion recognition processing
[1463] Users can enter text and voice messages while listening to music.
[1464] The device sends the user's voice and text input to the emotion engine.
[1465] The emotion engine analyzes the user's input, recognizes the emotion, and sends the result to the server.
[1466] Step 9:
[1467] Providing recommended content
[1468] The server searches the database for singers and songs that suit the user's emotions based on the emotion data received from the emotion engine.
[1469] The server generates a list of relevant singers and songs and transmits it to the user terminal.
[1470] The user can check the recommended content displayed on the terminal and select new songs.
[1471] Step 10:
[1472] Create and share playlists
[1473] After playing a number of songs, the user accesses the playlist creation screen.
[1474] The terminal sends a request to the server to save the list of songs selected by the user as a playlist.
[1475] The server stores the playlist information in a database and generates a playlist ID.
[1476] A link is created based on the generated playlist ID and sent to the device.
[1477] Users can share this link on social media and with other users.
[1478] Example 2
[1479] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1480] Conventional music playback systems have difficulty in playing a combination of singers and songs freely selected by the user in real time. Furthermore, they lack the functionality to recognize the user's emotions and provide appropriate recommended content, making it difficult to provide a music experience that best suits the user's current emotions. To solve this problem, a system is needed that can efficiently process user selection information and emotional data, synthesize music in real time, and make appropriate recommendations.
[1481] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1482] In this invention, the server includes means for processing singer and song selection information received from a user terminal, means for retrieving singer voice data and song score data from a database based on the received selection information, means for synthesizing a song using a voice synthesis engine using the retrieved voice data and song score data, means for transmitting the user's voice and text input to an emotion engine, means for recognizing the user's emotion data and transmitting it to the server to generate recommended content, and means for transmitting the recommended content to the user terminal. This allows the user to freely select singers and songs to be played in real time, and further allows the user to enjoy a personalized music experience based on their emotions.
[1483] A "user terminal" is a device that allows a user to access and operate the system through an interface, and includes smartphones, tablets, PCs, etc.
[1484] A "server" is a computer system that processes data received from a user terminal, retrieves necessary data from a database based on selection information, and performs appropriate processing.
[1485] "Singer and song selection information" is data relating to the singer name and song title selected by the user through the system.
[1486] The "database" is a data storage system that stores necessary information such as singer's voice data, music score data, and user's emotional history data.
[1487] A "voice synthesis engine" is software or hardware that synthesizes music under specified conditions based on acquired voice data and musical score data.
[1488] An "emotion engine" is software or hardware that recognizes emotions from a user's voice or text input and generates emotion data.
[1489] "Recommended content" is information about singers and songs that the system determines to be appropriate based on the user's emotional data.
[1490] "Playlist" is a function for saving and playing a list of multiple songs selected by the user.
[1491] The present invention is a system that plays back singer-music combinations freely selected by the user in real time, and also recognizes the user's emotions to provide recommended content. This system is composed of a user terminal, a server, a database, a speech synthesis engine, and an emotion engine.
[1492] Overall system configuration
[1493] 1. User Device
[1494] The user terminal is a device such as a smartphone, tablet, or PC, and provides a user interface through which users can search for and select singers and songs. It also has the function of sending the user's voice and text input to the emotion engine.
[1495] 2. Server
[1496] The server is responsible for processing the singer and song selection information received from the user device, as well as the emotion data from the emotion engine. The server analyzes the received data and retrieves the singer's voice data and song score data from a database. It also sends a request to a speech synthesis engine (e.g., Google Cloud Text-to-Speech API or Amazon Polly) and sends the synthesized voice data to the user device. It also has the function of generating recommended content based on the emotion data.
[1497] 3. Database
[1498] The database stores the singer's voice data, the music score data, and the user's emotion history data. After the server receives the request, it retrieves the necessary data from the database.
[1499] 4. Speech synthesis engine
[1500] The voice synthesis engine uses the acquired singer's voice data and the music score data to synthesize a song according to the specified conditions. The synthesized voice data is sent to the user's device via the server.
[1501] 5. Emotion Engine
[1502] The emotion engine recognizes emotions from the user's voice or text and sends the emotion data to the server. The recognized emotion data is stored in a database along with the user's emotion history data and is used to generate recommended content.
[1503] Program processing and specific examples
[1504] The program helps users with a series of operations, from logging in to playing songs, recognizing emotions, providing recommended content, and sharing.
[1505] Specific examples
[1506] If a user wants to listen to "Song B" with the voice of "Singer A":
[1507] 1. User login
[1508] The user enters the information required to log in to the app (user ID, password, etc.) and is authenticated. If authentication is successful, the user is redirected to the home screen.
[1509] 2. Search and select an artist and song
[1510] The user enters "Singer A" in the search bar and selects "Singer A" from the displayed candidates. Next, the user enters "Song B" in the song search bar and selects "Song B" from the displayed candidates.
[1511] 3. Data Acquisition and Speech Synthesis
[1512] The server processes the combination of "Singer A" and "Song B" and retrieves the respective data from the database. This data is passed to the voice synthesis engine, which synthesizes the song according to the specified conditions. The resulting voice data is returned to the server and sent to the user's device.
[1513] 4. Music playback and emotion recognition
[1514] The user clicks on the link provided by the device to play the music, and during playback, the user's voice and text input are analyzed by the emotion engine.
[1515] 5. Providing recommended content
[1516] The server recommends singers and songs suitable for the user based on the emotion data received from the emotion engine.
[1517] 6. Create and share playlists
[1518] Users can add the songs they play to a playlist, save the playlist, and share the link to the playlist via social media or email.
[1519] This system allows users to enjoy new musical experiences and provides a personalized musical experience that is tailored to their emotions.
[1520] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1521] Step 1:
[1522] User Login
[1523] The user launches the app, enters their user ID and password, and clicks the "Login" button. The device sends this information to the server, which then authenticates them by checking it against the user information in its database. If authentication is successful, the server sends the home screen data to the device, and the device displays the home screen.
[1524] Input: User ID, Password
[1525] Output: Home screen display
[1526] Specific operation: The device sends input to the server, the server performs authentication, generates a home screen and sends it to the device, which then displays the home screen.
[1527] Step 2:
[1528] Search and select singers and songs
[1529] The user enters "Singer A" in the search bar, and the device sends this query to the server. The server retrieves related singer information from the database and sends it to the device. The user selects "Singer A" from the displayed candidates, then enters "Song B" in the song search bar to search and select a song in the same way. This selection information is sent to the server.
[1530] Input: Artist name, song name
[1531] Output: Artist and song selection information
[1532] Specific operation: The device sends a search query to the server, the server retrieves candidates from the database and sends them to the device, the user enters selection information, which is sent to the server.
[1533] Step 3:
[1534] Data Acquisition and Speech Synthesis
[1535] The server retrieves voice data and music score data from a database based on the received singer name and song title. This data is passed to a voice synthesis engine, which synthesizes music according to the specified conditions. The resulting voice data is returned to the server and sent to the user's device.
[1536] Input: Artist name, song name
[1537] Output: Synthesized speech data
[1538] Specific operation: The server retrieves the necessary data from the database, sends a request to the speech synthesis engine to synthesize speech, and sends the result to the user's device.
[1539] Step 4:
[1540] Music playback and emotion recognition
[1541] The user clicks on the music playback link from their device. The device plays the music data and sends the user's voice and text input during playback to the emotion engine. The emotion engine analyzes this and sends emotional data to the server.
[1542] Input: Music link, voice or text input
[1543] Output: Emotion data
[1544] Specific operation: The device plays music, sends the data entered by the user to the emotion engine, and sends the analysis results to the server.
[1545] Step 5:
[1546] Providing recommended content
[1547] The server references the database based on the emotion data received from the emotion engine to obtain suitable singers and song candidates. It then generates a list of recommended content and sends it to the user's device. The device then displays the recommended content.
[1548] Input: Emotion data
[1549] Output: Recommended content list
[1550] Specific operation: The server analyzes the emotion data, generates recommended content, and sends it to the device. The device then displays the recommended content.
[1551] Step 6:
[1552] Create and share playlists
[1553] The user adds the songs they played to a playlist and accesses the playlist creation screen. The device sends the list of songs selected by the user to the server, which then saves the playlist information in a database and generates a playlist ID. The generated link is provided to the user's device, and the user can share it on social media, etc.
[1554] Input: Song list
[1555] Output: Playlist link
[1556] Specific operation: The device sends the song list to the server, which stores it in a database, generates a link, and sends it to the device. The user can then use the link to share the playlist.
[1557] (Application example 2)
[1558] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1559] Current music distribution services limit users to simply selecting and playing songs, and do not adequately provide a personalized music experience based on their emotions. Furthermore, few services recognize users' input emotions and recommend content based on those emotions, making it difficult to provide a music experience that best suits the user's emotions. The objective of this invention is to provide a system that not only plays singers and songs freely selected by the user, but also recognizes the user's emotions in real time and provides recommended content based on those emotions.
[1560] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1561] In this invention, the server includes means for processing singer and song selection information received from a user terminal, means for retrieving singer voice data and song score data from a database based on the received selection information, means for synthesizing music using a voice synthesis engine using the retrieved voice data and song score data, means for transmitting the synthesized voice data to the user terminal, means for processing emotion data input by the user and recommending related information, means for recommending and providing content such as news and music based on the emotion data, and means for recognizing the emotion input by the user, thereby enabling a personalized music experience that matches the user's emotions.
[1562] "User terminal" means a device that a user accesses and uses to input and make selections.
[1563] "Selection information" is information about the singers and songs selected by the user.
[1564] "Audio data" refers to data that stores the voice of a particular singer in digital form.
[1565] "Music score data" refers to data that stores the melody and chords of a particular piece of music in digital format.
[1566] A "voice synthesis engine" is software or a system that generates new music using voice data and musical score data.
[1567] "Emotional data" refers to data about emotions derived from user input and actions.
[1568] "Recommended content" refers to music and information provided based on a user's emotional data.
[1569] A "search interface" is a screen or input field that allows a user to search for a particular artist or song.
[1570] A "playlist" is a function that organizes songs into a list, making it easy to play them consecutively and share them.
[1571] A "unique link" is a URL or hyperlink that provides access to a specific playlist or content.
[1572] This invention is a system that plays back singer-music combinations freely selected by the user in real time, and also recognizes the user's emotions to provide recommended content. The system is composed of a user terminal, a server, a database, a speech synthesis engine, and an emotion engine.
[1573] Overall system configuration
[1574] 1. User Device
[1575] The user terminal is a device such as a smartphone, tablet, or PC, and provides a user interface. From the user interface, users can search for and select singers and songs. It also has the function of sending the user's voice and text input to the emotion engine.
[1576] 2. Server
[1577] The server is responsible for processing the singer and song selection information received from the user device, as well as the emotion data from the emotion engine. The server analyzes the received data and retrieves the singer's voice data and song score data from the database. It also sends requests to the voice synthesis engine and transmits the synthesized voice data to the user device. It also has the function of generating recommended content based on the emotion data.
[1578] 3. Database
[1579] The database stores the singer's voice data, the music score data, and the user's emotion history data. After the server receives the request, it retrieves the necessary data from the database.
[1580] 4. Speech synthesis engine
[1581] The voice synthesis engine uses the acquired singer's voice data and the music score data to synthesize a song according to the specified conditions, and the synthesized voice data is sent to the user's device via the server.
[1582] 5. Emotion Engine
[1583] The emotion engine recognizes emotions from the user's voice or text and sends the emotion data to the server. The recognized emotion data is stored in a database along with the user's emotion history data and used to generate recommended content.
[1584] Program processing and specific examples
[1585] The program supports a series of operations, from user login to playing music, recognizing emotions, providing recommended content, and sharing. While the program's specific processing steps are not included, the following concrete examples are provided:
[1586] Specific examples
[1587] 1. Log in
[1588] The user logs in to the app. They enter their username and password, and the server authenticates them. If authentication is successful, the user is taken to the home screen.
[1589] 2. Selection of singers and songs
[1590] Users can enter a specific artist in the search bar and choose from a list of artists, then enter a song name in the song search bar and choose from a list of songs.
[1591] 3. Data Acquisition and Speech Synthesis
[1592] The server processes the selection information, retrieves the respective data from the database, and passes this data to the speech synthesis engine to synthesize the music.
[1593] 4. Music playback and emotion recognition
[1594] The server sends the synthesized voice data to the user's device. The user clicks the provided link and plays the music. During playback, the user's voice and text input are analyzed by the emotion engine.
[1595] 5. Providing recommended content
[1596] The emotion data analyzed by the emotion engine is sent to the server, which then recommends singers and songs that are suitable for the user based on the emotion data.
[1597] 6. Create and share playlists
[1598] Users can save the songs they play as playlists and share them on social media. Users can easily share music with friends and family using the playlist link.
[1599] Prompt Sentence Examples
[1600] "User text input: 'I'm feeling a bit down...'"
[1601] "The generative AI model's output: 'That's sad. Let's pick a song that's uplifting.'"
[1602] In this way, the present invention is a system that provides recommended content that corresponds to the user's emotions, thereby realizing a personalized music experience.
[1603] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1604] Step 1:
[1605] A user logs in to the app.
[1606] Input: User ID, Password
[1607] Output: Authentication result (success / failure), home screen
[1608] Specific operation: The user enters their user ID and password on the login screen and sends them to the server. The server retrieves the relevant information from the database and compares it with the entered information. If authentication is successful, the user is taken to the home screen; if it is unsuccessful, an error message is displayed.
[1609] Step 2:
[1610] The user searches for and selects an artist and song.
[1611] Input: Search keyword (singer name, song name)
[1612] Output: Search result list, selection information
[1613] Specific operation: The user enters the name of an artist or song in the search bar and presses the Enter key. The device sends the input information to the server, which retrieves the corresponding artist and song information from the database and sends a list of search results to the user's device. The user then selects a specific artist and song from the list.
[1614] Step 3:
[1615] The server retrieves the necessary data from the database based on the selection information.
[1616] Input: Selection information (singer name, song name)
[1617] Output: Singer's voice data, music score data
[1618] Specific operation: The server analyzes the selection information received from the user terminal and sends a request to the database, which then returns the corresponding singer's voice data and the music score data to the server.
[1619] Step 4:
[1620] The server synthesizes the music using a speech synthesis engine.
[1621] Input: Audio data, music score data
[1622] Output: Synthesized speech data
[1623] Specific operation: The server passes the acquired voice data and musical score data to the speech synthesis engine, which synthesizes the music. The speech synthesis engine processes the data and generates new synthetic voice data, which is then sent back to the server.
[1624] Step 5:
[1625] The synthesized speech data is transmitted to the user terminal.
[1626] Input: Synthetic speech data
[1627] Output: Music playback link
[1628] Specific operation: The server sends the synthesized voice data to the user's device and provides a music playback link. The user can click the link to play the synthesized music in real time.
[1629] Step 6:
[1630] The emotion engine analyzes the user's voice and text input.
[1631] Input: User voice or text input
[1632] Output: Emotion data
[1633] Specific operation: The user inputs voice or text into the emotion engine while playing music. The emotion engine receives the data, analyzes it, and generates emotion data as a result. The emotion data is then sent to the server.
[1634] Step 7:
[1635] Based on the emotion data, the server provides recommended content.
[1636] Input: Emotion data
[1637] Output: Recommended singer and song information
[1638] Specific operation: The server analyzes the emotion data and generates recommended content based on it. Specifically, it selects appropriate singers and songs from the database and sends that information to the user's device.
[1639] Step 8:
[1640] Users can save the songs they play as a playlist and share it on social media.
[1641] Input: Information about the song played
[1642] Output: Playlist link
[1643] Specific operation: The user selects the songs they played and saves them as a playlist. The server saves the playlist information in a database and generates a dedicated link. The generated link is sent to the user's device, and the user can share it via social media or with other users.
[1644] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1645] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1646] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1647] [Fourth embodiment]
[1648] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1649] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1650] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1651] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1652] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1653] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1654] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1655] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1656] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1657] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1658] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1659] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1660] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1661] The present invention is a system that allows users to freely select singers and song combinations, play them in real time, and share them. This system is composed of a user terminal, a server, a database, and a voice synthesis engine.
[1662] Overall system configuration
[1663] 1. User Device
[1664] The user terminal is a device such as a smartphone, tablet, or PC, and provides a user interface through which the user can search for and select singers and songs.
[1665] 2. Server
[1666] The server is responsible for processing the singer and song selection information received from the user device. The server analyzes the received data, retrieves the singer's voice data and song score data from the database, sends a request to the voice synthesis engine, and transmits the synthesized voice data to the user device.
[1667] 3. Database
[1668] The database stores the singer's voice data and the music score data. After the server receives the request, it retrieves the necessary data from the database.
[1669] 4. Speech synthesis engine
[1670] The voice synthesis engine uses the acquired singer's voice data and the music score data to synthesize a song according to the specified conditions. The synthesized voice data is sent to the user's device via the server.
[1671] Program processing and specific examples
[1672] Program processing
[1673] The program supports users in a series of operations from logging in to playing and sharing songs. The specific processing steps of the program are as follows:
[1674] 1. User login
[1675] The user enters the information required to log in to the app (user ID, password, etc.) and is authenticated. If authentication is successful, the user is redirected to the home screen.
[1676] 2. Search and select an artist and song
[1677] The user searches for singers and songs using keywords in the search interface and selects from the corresponding candidates. The selected information is sent to the server.
[1678] 3. Data Acquisition and Speech Synthesis
[1679] The server retrieves the singer's voice data and the song's score data from a database based on the received selection information. The retrieved data is passed to a voice synthesis engine, which synthesizes the song according to the specified conditions. The resulting voice data is returned to the server and sent to the user's device.
[1680] 4. Playing Music and Creating Playlists
[1681] Users can play the synthesized music using the link sent to them, and can also save multiple songs as a playlist and share it on social media.
[1682] Specific examples
[1683] For example, suppose a user wants to listen to "Song B" in the voice of "Singer A."
[1684] 1. Log in
[1685] A user logs in to the app.
[1686] Enter your username and password and the server will authenticate you.
[1687] If the authentication is successful, the user will be taken to the home screen.
[1688] 2. Selection of singers and songs
[1689] The user enters "singer A" in the search bar and selects "singer A" from the list of singers.
[1690] Next, enter "Song B" in the song search bar and select "Song B" from the list of songs.
[1691] 3. Data Acquisition and Speech Synthesis
[1692] The server processes the combination of "singer A" and "song B" and retrieves the respective data from the database.
[1693] The server passes this data to a voice synthesis engine, which synthesizes music according to the specified conditions.
[1694] 4. Playing songs
[1695] The server transmits the synthesized voice data to the user terminal.
[1696] The user plays "Song B" with the voice of "Singer A" on the device and enjoys the music.
[1697] 5. Create and share playlists
[1698] Users can save the songs they play as a playlist and share it on social media.
[1699] Users can easily share music with friends and family using the links of their saved playlists.
[1700] The present invention is a system that can provide users with a new way of enjoying music.
[1701] The processing flow will be explained below.
[1702] Step 1:
[1703] A user logs in to the app
[1704] The user launches the app on the device and accesses the login screen.
[1705] The terminal receives the user ID and password entered by the user.
[1706] The terminal sends this login information to the server.
[1707] The server checks the received user ID and password in a database and returns the authentication result to the terminal.
[1708] If the authentication is successful, the terminal transitions the user to the home screen.
[1709] Step 2:
[1710] User searches and selects singer
[1711] The user enters the singer's name into the device's search bar.
[1712] The terminal sends the entered name to the server.
[1713] The server searches the database for a list of singers based on the received singer name and returns the list to the terminal.
[1714] The user selects a particular singer from a list of singers displayed on the terminal.
[1715] The terminal transmits information about the selected singer to the server.
[1716] Step 3:
[1717] User searches and selects a song
[1718] The user enters the title of the song into the device's search bar.
[1719] The terminal transmits the input song title to the server.
[1720] The server searches the database for a list of songs based on the received song title and returns it to the terminal.
[1721] The user selects a specific song from the song list displayed on the terminal.
[1722] The terminal transmits information about the selected song to the server.
[1723] Step 4:
[1724] The server retrieves the data
[1725] The server analyzes the received selection information (singer name and song title).
[1726] The server retrieves the singer's voice data and the music score data from the database.
[1727] The acquired data is temporarily stored on the server.
[1728] Step 5:
[1729] Perform speech synthesis
[1730] The server passes the acquired voice data and musical score data to the voice synthesis engine.
[1731] The speech synthesis engine generates synthetic music based on this data.
[1732] The synthesized music file is sent back to the server.
[1733] Step 6:
[1734] Sending synthesized voice data
[1735] The server transmits the synthesized voice data to the user terminal as streaming or a download link.
[1736] The terminal receives this and makes it available for playback by the user.
[1737] Step 7:
[1738] User plays music
[1739] The user clicks on the link provided by the device and plays the music.
[1740] The device streams the audio data and plays it back in real time.
[1741] Step 8:
[1742] Create and share playlists
[1743] After listening to a number of songs, the user accesses a playlist creation screen.
[1744] The terminal sends a request to the server to save the list of songs selected by the user as a playlist.
[1745] The server stores the playlist information in a database and generates a playlist ID.
[1746] A link is created based on the generated playlist ID and sent to the device.
[1747] Users can share this link on social media and with other users.
[1748] Example 1
[1749] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1750] With conventional systems, it was difficult for users to play back their desired singer-song combinations in real time and easily share them. Furthermore, the process of synthesizing audio data and sheet music data was complex and time-consuming, which sometimes compromised the user experience. This resulted in a lack of flexibility and convenience for proposing new ways to enjoy music.
[1751] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1752] In this invention, the server includes means for processing selection information of voice data and music data received from a user terminal, means for acquiring voice data and music score data from a storage device based on the received selection information, means for synthesizing music using the acquired voice data and music score data with a voice synthesizer, and means for transmitting the synthesized voice data to the user terminal, thereby enabling users to freely select combinations of singers and music to be played in real time and easily shared.
[1753] A "user terminal" is a device that allows a user to interact with the system through an interface and perform various operations, and specifically refers to a smartphone, tablet, PC, etc.
[1754] "Selection Information" refers to data about a particular artist and song that a user enters through the system.
[1755] A "server" refers to a computer system that receives requests from user terminals and processes data and provides services in response to those requests.
[1756] "Audio data" refers to acoustic information that is a digital recording of the voice of a particular singer.
[1757] "Musical score data" refers to digital data that records musical information such as the melody, rhythm, and chords of a song.
[1758] "Storage device" refers to hardware and software for storing digital information such as audio data, sheet music data, and song selection lists.
[1759] A "speech synthesizer" refers to hardware or software that uses analog or digital voice data to generate new voices.
[1760] "Synthesized voice data" refers to music data generated by a voice synthesizer based on the voice data of a singer selected by the user and music score data.
[1761] A "search screen" refers to a user interface that provides an interface for a user to search for a specific singer or song.
[1762] "Song selection list" refers to a data list for managing multiple songs selected and saved by the user.
[1763] "Information sharing service" means an online platform that enables users to share information with other users via the Internet.
[1764] "Dedicated link" refers to a URL or URI generated to directly access a specific song selection list or voice synthesis data.
[1765] The present invention provides a system that allows users to freely select and share combinations of voice data and music data in real time. This system is comprised of a user terminal, a server, a storage device, and a voice synthesizer.
[1766] Overall system configuration
[1767] 1. User Device
[1768] The user terminal is a device such as a smartphone, tablet, or PC, and provides a user interface. The user can search and select audio data and music data through a search screen, allowing the user to easily find the music they want.
[1769] 2. Server
[1770] The server is responsible for processing the voice data and music data selection information received from the user terminal. The server analyzes the received data and retrieves the voice data and music score data from the storage device. It also sends a request to the voice synthesizer and sends the synthesized voice data to the user terminal. This allows the user to generate the music they want in real time.
[1771] 3. Storage device
[1772] The storage device stores the voice data and the musical score data. After the server receives the request, it retrieves the necessary data from the storage device. This data is provided to the voice synthesizer and used to synthesize the music.
[1773] 4. Speech synthesizer
[1774] The speech synthesizer uses the acquired voice data and musical score data to synthesize music under specified conditions. The synthesized voice data is sent to the user's terminal via a server. The speech synthesizer can generate music in real time.
[1775] Specific examples
[1776] Below is a specific example where a user wants to listen to "Song B" in the voice of "Singer A."
[1777] 1. Log in
[1778] A user logs in to a smartphone app.
[1779] Enter your username and password and the server will authenticate you.
[1780] If the authentication is successful, the user will be taken to the home screen.
[1781] 2. Selecting audio and music data
[1782] The user enters "singer A" in the search bar and selects "singer A" from the list of audio data.
[1783] Next, enter "Song B" in the song search bar and select "Song B" from the list of song data.
[1784] 3. Data Acquisition and Speech Synthesis
[1785] The server processes the combination of "singer A" and "song B" and retrieves the respective data from the storage device.
[1786] The server passes this data to a voice synthesizer, which synthesizes music according to the specified conditions.
[1787] 4. Playing songs
[1788] The server transmits the synthesized voice data to the user terminal.
[1789] The user plays "Song B" with the voice of "Singer A" on the device and enjoys the music.
[1790] 5. Create and share song lists
[1791] The songs played by the user are saved as a song selection list and shared via an information sharing service.
[1792] Users can easily share music with friends and family using links to their saved selections.
[1793] Prompt Sentence Examples
[1794] This is a system that generates music in real time using "Singer A" and "Song B." Specifically, a user logs into the app and enters "Singer A" and "Song B" in the search bar. The data for the selected singer and song is sent to a server, which then retrieves the necessary voice data and musical score data from a storage device and passes them to a voice synthesizer to synthesize the song. The synthesized song is sent to the user's device, where the user can play it.
[1795] With the above configuration, this system can provide users with a new way of enjoying music.
[1796] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1797] Step 1: The user launches the app and enters their user ID and password on the login screen.
[1798] Input: User ID, Password
[1799] Output: Authentication request
[1800] Specific operation: When the user taps the login button, the entered information is sent to the server.
[1801] Step 2: The server processes the received authentication information and checks it against a database.
[1802] Input: User ID, Password
[1803] Output: Authentication result
[1804] Specific operation: The server checks the user ID and password against the database, and if authentication is successful, generates a response indicating successful authentication.
[1805] Step 3: If the authentication is successful, the server sends an instruction to the user terminal to transition to the home screen.
[1806] Input: Authentication result (success)
[1807] Output: Home screen display instruction
[1808] Specific operation: Information indicating access rights to the home screen is sent to the user terminal, and the home screen is displayed on the user terminal.
[1809] Step 4: The user enters keywords for audio data (e.g., "singer A") and song data (e.g., "song B") into the search bar.
[1810] Input: Keywords for audio data and music data
[1811] Output: Search request
[1812] Specific operation: When the user taps the search button, the entered keywords are sent to the server.
[1813] Step 5: The server searches the storage device for related audio data and music data based on the received search keyword.
[1814] Input: Search keywords for audio data and music data
[1815] Output: Search results
[1816] Specific operation: The server searches the storage device and obtains the corresponding audio data and music data.
[1817] Step 6: The server sends the search results to the user terminal, and the user selects the appropriate singer and song.
[1818] Input: Search results
[1819] Output: User's choice
[1820] Specific operation: The user taps to select the desired artist and song from the list of search results.
[1821] Step 7: The server receives the user's selection information and retrieves the corresponding audio data and music score data from the storage device.
[1822] Input: User selection information
[1823] Output: Audio data and sheet music data
[1824] Specific operation: The server accesses the storage device and acquires the selected audio data and music score data.
[1825] Step 8: The server passes the acquired data to a voice synthesizer, which synthesizes music according to the specified conditions.
[1826] Input: Audio data and sheet music data
[1827] Output: Synthesized voice data
[1828] Specific operation: The voice synthesizer synthesizes music according to the specified conditions, and the generated voice data is sent back to the server.
[1829] Step 9: The server transmits the synthesized voice data to the user terminal.
[1830] Input: Synthesized voice data
[1831] Output: Playback link for audio data
[1832] Specific operation: The synthesized voice data is sent to the user's terminal and a playback link is displayed.
[1833] Step 10: The user taps the play link to play the synthesized voice data.
[1834] Input: Audio data playback link
[1835] Output: Playing music
[1836] Specific operation: The user device opens the link and plays the synthesized voice data.
[1837] Step 11: The user selects multiple songs, creates a song selection list, and shares it via an information sharing service.
[1838] Input: Song list
[1839] Output: Link to song selection list
[1840] Specific operation: The song list is saved in a storage device, and a link is generated and displayed on the user's device. The user can then share the link on social media.
[1841] (Application example 1)
[1842] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1843] Conventional music distribution services have difficulty generating and playing combinations of singers and songs freely selected by users in real time, limiting the variety of music experiences they can offer. Furthermore, there is a lack of systems that allow users to easily share the music they create through social networking sites or messaging applications. A system that can solve these issues and provide users with a new music experience is needed.
[1844] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1845] In this invention, the server includes means for processing singer and song selection information received from a user terminal, means for retrieving singer voice data and song score data from a database based on the received selection information, means for synthesizing a song using a voice synthesis engine using the retrieved voice data and song score data, means for transmitting the synthesized voice data to the user terminal, and means for users to play and share the song using a link to the synthesized voice data. This allows users to freely create and play back combinations of singers and songs in real time, and to easily share the created songs via social networking sites and messaging applications.
[1846] A "user terminal" is a device used by a user to operate the system, and includes smartphones, tablets, PCs, etc.
[1847] "Singer and song selection information" is information input by the user to combine a specific singer with a song.
[1848] "Audio data" refers to data that digitally records the voice of a particular singer.
[1849] "Musical score data" refers to data in which the notes, chords, rhythms, etc. of a piece of music are recorded in digital format.
[1850] "Database" means a data storage system for storing and managing audio data and musical score data.
[1851] A "voice synthesis engine" is software that generates music by combining voice data and musical score data.
[1852] "Synthetic voice data" is data of a song sung in the voice of a specific singer, generated by a voice synthesis engine.
[1853] A "search interface" is a screen or input device that allows a user to search for and select an artist and song.
[1854] A "server" is a computer system that processes requests from user terminals and provides information in cooperation with a database and a speech synthesis engine.
[1855] A "link" is a URL or URI that a user uses to play back synthesized speech data.
[1856] "SNS" is an abbreviation for social networking service, a platform for users to share information online.
[1857] A "messaging application" is software that allows users to exchange messages and information in real time.
[1858] A "playlist" is a collection of songs saved by a user and organized for easy playback and sharing.
[1859] The present invention is a system that receives information on singer and song selection from a user terminal, synthesizes the song in real time using a specific voice synthesis engine, and provides the generated voice data to the user. Specifically, the present invention is implemented using a user terminal, a server, a database, and a voice synthesis engine.
[1860] User terminal
[1861] The user terminal consists of a smartphone, tablet, PC, etc., and is the device through which the user operates the system. The user terminal is provided with a search interface for searching and selecting singers and songs. It also has the functionality to receive a link to the synthesized voice data and play and share it.
[1862] server
[1863] The server processes the selection information received from the user's device and retrieves the necessary audio data and music score data from the database. Specifically, the server was built using the Python Flask framework and receives input from the search interface. The server then works with a speech synthesis engine to generate music using the retrieved data. It is then responsible for sending a link to the generated audio data to the user's device.
[1864] Database
[1865] The database is a data storage system for storing and managing singers' voice data and music score data. The server quickly retrieves the necessary data from the database based on the selection information received.
[1866] Text-to-speech engine
[1867] A speech synthesis engine is software that synthesizes music based on specified conditions using acquired voice data and musical score data. Specifically, it can generate music in real time using external services such as speech synthesis APIs.
[1868] Specific examples of user operations
[1869] The user launches the app and logs in. They enter their favorite singer and song in the search bar and select from the search results. For example, if they want to listen to "Song B" in the voice of "Singer A," they enter and select this. The server retrieves the necessary data from the database based on the selection and passes it to the speech synthesis engine. A link to the synthesized voice data is sent to the user's device, and the user can use this link to play and share the song.
[1870] Prompt Sentence Examples
[1871] User: Please create song B using singer A's voice.
[1872] System: Synthesized audio data URL: http: / / example.com / audio / 1234
[1873] This invention allows users to freely select singers and song combinations, generate and play them in real time, and easily share them via social networking sites and messaging applications, providing a new musical experience.
[1874] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1875] Step 1:
[1876] The user starts the application on their device, such as a smartphone or PC, and goes to the login screen. The user enters their user ID and password and presses the login button.
[1877] Input: User ID, Password
[1878] Action: User device sends authentication information to server
[1879] Output: Authentication result from the server (success or failure)
[1880] Step 2:
[1881] The server authenticates the user based on the received authentication information. If authentication is successful, the user is redirected to the home screen and the search interface is displayed.
[1882] Input: Authentication information (user ID, password)
[1883] How it works: The server checks the authentication information against a database and sends the result back to the user's device.
[1884] Output: Authentication result (home screen if successful, error message if unsuccessful)
[1885] Step 3:
[1886] The user searches for and selects an artist and song in the search interface, for example, by entering "artist A" and "song B" in the search bar and selecting from the matching results.
[1887] Input: Search keyword (singer name, song name)
[1888] Action: Sends selections from the search interface to the server
[1889] Output: The server receives the selection information and displays the results on the user's terminal.
[1890] Step 4:
[1891] The server acquires the singer's voice data and the musical score data of the song from the database based on the received selection information.
[1892] Input: Selection information (singer name, song name)
[1893] How it works: The server sends a request to the database and retrieves the required data.
[1894] Output: Audio data, music score data
[1895] Step 5:
[1896] The server passes the acquired voice data and musical score data to a voice synthesis engine, which synthesizes music according to the specified conditions.
[1897] Input: Audio data, music score data
[1898] How it works: The server calls the speech synthesis engine and sends a synthesis request.
[1899] Output: Synthesized audio data
[1900] Step 6:
[1901] The server transmits the synthesized voice data to the user terminal and provides the user with a link to the synthesized voice data.
[1902] Input: Synthetic speech data
[1903] Operation: The server stores the synthesized voice data, generates a link to it, and sends it to the user's device.
[1904] Output: Link to the synthesized speech data
[1905] Step 7:
[1906] Users can use the provided link to play the synthesized voice and share it via social media or messaging applications.
[1907] Input: Synthetic speech data link
[1908] What happens: The user receives the link, plays the audio, and offers sharing options.
[1909] Output: Play audio, share
[1910] By following these steps, users can freely create and play combinations of singers and songs in real time, and can easily share the created songs via social media or messaging applications.
[1911] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1912] The present invention is a system that plays back singer-music combinations freely selected by the user in real time, and also recognizes the user's emotions to provide recommended content. This system is composed of a user terminal, a server, a database, a speech synthesis engine, and an emotion engine.
[1913] Overall system configuration
[1914] 1. User Device
[1915] The user terminal is a device such as a smartphone, tablet, or PC, and provides a user interface. From the user interface, users can search for and select singers and songs. It also has the function of sending the user's voice and text input to the emotion engine.
[1916] 2. Server
[1917] The server is responsible for processing the singer and song selection information received from the user terminal, as well as the emotion data from the emotion engine. The server analyzes the received data and retrieves the singer's voice data and song score data from the database. It also sends requests to the voice synthesis engine and transmits the synthesized voice data to the user terminal. It also has the function of generating recommended content based on the emotion data.
[1918] 3. Database
[1919] The database stores the singer's voice data, the music score data, and the user's emotion history data. After the server receives the request, it retrieves the required data from the database.
[1920] 4. Speech synthesis engine
[1921] The voice synthesis engine uses the acquired singer's voice data and the music score data to synthesize a song according to the specified conditions. The synthesized voice data is sent to the user's device via the server.
[1922] 5. Emotion Engine
[1923] The emotion engine recognizes emotions from the user's voice or text and sends the emotion data to the server. The recognized emotion data is stored in a database along with the user's emotion history data and used to generate recommended content.
[1924] Program processing and specific examples
[1925] Program processing
[1926] The program supports users in a series of operations, from logging in to playing songs, recognizing emotions, providing recommended content, and sharing. The specific processing steps of the program are as follows:
[1927] 1. User login
[1928] The user enters the information required to log in to the app (user ID, password, etc.) and is authenticated. If authentication is successful, the user is redirected to the home screen.
[1929] 2. Search and select an artist and song
[1930] The user searches for singers and songs using keywords in the search interface and selects from the corresponding candidates. The selected information is sent to the server.
[1931] 3. Data Acquisition and Speech Synthesis
[1932] The server analyzes the received selection information (singer name and song title) and retrieves the singer's voice data and the song's score data from the database. The retrieved data is passed to a voice synthesis engine, which synthesizes the song according to the specified conditions. The resulting voice data is returned to the server and sent to the user's device.
[1933] 4. Music playback and emotion recognition
[1934] The user can use the sent link to play the synthesized music, while the user's voice and text input are analyzed by the emotion engine.
[1935] 5. Providing recommended content
[1936] Based on the emotional data analyzed by the emotion engine, the server presents recommended singers and songs to the user, allowing the user to have a music experience that best suits their current emotions.
[1937] 6. Create and share playlists
[1938] After listening to multiple songs, the user accesses the playlist creation screen. The device sends a request to the server to save the list of songs selected by the user as a playlist. The server saves the playlist information in a database and generates a playlist ID. A link is created based on the generated playlist ID and sent to the device. The user can share this link via social media or with other users.
[1939] Specific examples
[1940] For example, suppose a user wants to listen to "Song B" in the voice of "Singer A."
[1941] 1. Log in
[1942] A user logs in to the app.
[1943] Enter your username and password and the server will authenticate you.
[1944] If the authentication is successful, the user will be taken to the home screen.
[1945] 2. Selection of singers and songs
[1946] The user enters "singer A" in the search bar and selects "singer A" from the list of singers.
[1947] Next, enter "Song B" in the song search bar and select "Song B" from the list of songs.
[1948] 3. Data Acquisition and Speech Synthesis
[1949] The server processes the combination of "singer A" and "song B" and retrieves the respective data from the database.
[1950] The server passes this data to a voice synthesis engine, which synthesizes music according to the specified conditions.
[1951] 4. Music playback and emotion recognition
[1952] The server transmits the synthesized voice data to the user terminal.
[1953] The user clicks on the link provided by the device and plays the music.
[1954] During playback, the user's voice and text input are analyzed by the emotion engine.
[1955] 5. Providing recommended content
[1956] The emotion data analyzed by the emotion engine is sent to the server.
[1957] The server recommends singers and songs that suit the user based on the emotion data.
[1958] 6. Create and share playlists
[1959] Users can save the songs they play as a playlist and share it on social media.
[1960] Users can easily share music with friends and family using the links of their saved playlists.
[1961] This invention is a system that provides users with a new way of enjoying music and also realizes a personalized music experience that matches the user's emotions.
[1962] The processing flow will be explained below.
[1963] Step 1:
[1964] A user logs in to the app
[1965] The user launches the app on the device and accesses the login screen.
[1966] The terminal receives the user ID and password entered by the user.
[1967] The terminal sends this login information to the server.
[1968] The server checks the received user ID and password in a database and returns the authentication result to the terminal.
[1969] If the authentication is successful, the terminal transitions the user to the home screen.
[1970] Step 2:
[1971] User searches and selects singer
[1972] The user enters the singer's name into the device's search bar.
[1973] The terminal sends the entered name to the server.
[1974] The server searches the database for a list of singers based on the received singer name and returns the list to the terminal.
[1975] The user selects a particular singer from a list of singers displayed on the terminal.
[1976] The terminal transmits information about the selected singer to the server.
[1977] Step 3:
[1978] User searches and selects a song
[1979] The user enters the title of the song into the device's search bar.
[1980] The terminal transmits the input song title to the server.
[1981] The server searches the database for a list of songs based on the received song title and returns it to the terminal.
[1982] The user selects a specific song from the song list displayed on the terminal.
[1983] The terminal transmits information about the selected song to the server.
[1984] Step 4:
[1985] The server retrieves the data
[1986] The server analyzes the received selection information (singer name and song title).
[1987] The server retrieves the singer's voice data and the music score data from the database.
[1988] The acquired data is temporarily stored on the server.
[1989] Step 5:
[1990] Perform speech synthesis
[1991] The server passes the acquired voice data and musical score data to the voice synthesis engine.
[1992] The speech synthesis engine generates synthetic music based on this data.
[1993] The synthesized music file is sent back to the server.
[1994] Step 6:
[1995] Sending synthesized voice data
[1996] The server transmits the synthesized voice data to the user terminal as streaming or a download link.
[1997] The terminal receives this and makes it available for playback by the user.
[1998] Step 7:
[1999] User plays music
[2000] The user clicks on the link provided by the device and plays the music.
[2001] The device streams the audio data and plays it back in real time.
[2002] Step 8:
[2003] Emotion recognition processing
[2004] Users can enter text and voice messages while listening to music.
[2005] The device sends the user's voice and text input to the emotion engine.
[2006] The emotion engine analyzes the user's input, recognizes the emotion, and sends the result to the server.
[2007] Step 9:
[2008] Providing recommended content
[2009] The server searches the database for singers and songs that suit the user's emotions based on the emotion data received from the emotion engine.
[2010] The server generates a list of relevant singers and songs and transmits it to the user terminal.
[2011] The user can check the recommended content displayed on the terminal and select new songs.
[2012] Step 10:
[2013] Create and share playlists
[2014] After playing a number of songs, the user accesses the playlist creation screen.
[2015] The terminal sends a request to the server to save the list of songs selected by the user as a playlist.
[2016] The server stores the playlist information in a database and generates a playlist ID.
[2017] A link is created based on the generated playlist ID and sent to the device.
[2018] Users can share this link on social media and with other users.
[2019] Example 2
[2020] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[2021] Conventional music playback systems have difficulty in playing a combination of singers and songs freely selected by the user in real time. Furthermore, they lack the functionality to recognize the user's emotions and provide appropriate recommended content, making it difficult to provide a music experience that best suits the user's current emotions. To solve this problem, a system is needed that can efficiently process user selection information and emotional data, synthesize music in real time, and make appropriate recommendations.
[2022] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[2023] In this invention, the server includes means for processing singer and song selection information received from a user terminal, means for retrieving singer voice data and song score data from a database based on the received selection information, means for synthesizing a song using a voice synthesis engine using the retrieved voice data and song score data, means for transmitting the user's voice and text input to an emotion engine, means for recognizing the user's emotion data and transmitting it to the server to generate recommended content, and means for transmitting the recommended content to the user terminal. This allows the user to freely select singers and songs to be played in real time, and further allows the user to enjoy a personalized music experience based on their emotions.
[2024] A "user terminal" is a device that allows a user to access and operate the system through an interface, and includes smartphones, tablets, PCs, etc.
[2025] A "server" is a computer system that processes data received from a user terminal, retrieves necessary data from a database based on selection information, and performs appropriate processing.
[2026] "Singer and song selection information" is data relating to the singer name and song title selected by the user through the system.
[2027] The "database" is a data storage system that stores necessary information such as singer's voice data, music score data, and user's emotional history data.
[2028] A "voice synthesis engine" is software or hardware that synthesizes music under specified conditions based on acquired voice data and musical score data.
[2029] An "emotion engine" is software or hardware that recognizes emotions from a user's voice or text input and generates emotion data.
[2030] "Recommended content" is information about singers and songs that the system determines to be appropriate based on the user's emotional data.
[2031] "Playlist" is a function for saving and playing a list of multiple songs selected by the user.
[2032] The present invention is a system that plays back singer-music combinations freely selected by the user in real time, and also recognizes the user's emotions to provide recommended content. This system is composed of a user terminal, a server, a database, a speech synthesis engine, and an emotion engine.
[2033] Overall system configuration
[2034] 1. User Device
[2035] The user terminal is a device such as a smartphone, tablet, or PC, and provides a user interface through which users can search for and select singers and songs. It also has the function of sending the user's voice and text input to the emotion engine.
[2036] 2. Server
[2037] The server is responsible for processing the singer and song selection information received from the user device, as well as the emotion data from the emotion engine. The server analyzes the received data and retrieves the singer's voice data and song score data from a database. It also sends a request to a speech synthesis engine (e.g., Google Cloud Text-to-Speech API or Amazon Polly) and sends the synthesized voice data to the user device. It also has the function of generating recommended content based on the emotion data.
[2038] 3. Database
[2039] The database stores the singer's voice data, the music score data, and the user's emotion history data. After the server receives the request, it retrieves the necessary data from the database.
[2040] 4. Speech synthesis engine
[2041] The voice synthesis engine uses the acquired singer's voice data and the music score data to synthesize a song according to the specified conditions. The synthesized voice data is sent to the user's device via the server.
[2042] 5. Emotion Engine
[2043] The emotion engine recognizes emotions from the user's voice or text and sends the emotion data to the server. The recognized emotion data is stored in a database along with the user's emotion history data and is used to generate recommended content.
[2044] Program processing and specific examples
[2045] The program helps users with a series of operations, from logging in to playing songs, recognizing emotions, providing recommended content, and sharing.
[2046] Specific examples
[2047] If a user wants to listen to "Song B" with the voice of "Singer A":
[2048] 1. User login
[2049] The user enters the information required to log in to the app (user ID, password, etc.) and is authenticated. If authentication is successful, the user is redirected to the home screen.
[2050] 2. Search and select an artist and song
[2051] The user enters "Singer A" in the search bar and selects "Singer A" from the displayed candidates. Next, the user enters "Song B" in the song search bar and selects "Song B" from the displayed candidates.
[2052] 3. Data Acquisition and Speech Synthesis
[2053] The server processes the combination of "Singer A" and "Song B" and retrieves the respective data from the database. This data is passed to the voice synthesis engine, which synthesizes the song according to the specified conditions. The resulting voice data is returned to the server and sent to the user's device.
[2054] 4. Music playback and emotion recognition
[2055] The user clicks on the link provided by the device to play the music, and during playback, the user's voice and text input are analyzed by the emotion engine.
[2056] 5. Providing recommended content
[2057] The server recommends singers and songs suitable for the user based on the emotion data received from the emotion engine.
[2058] 6. Create and share playlists
[2059] Users can add the songs they play to a playlist, save the playlist, and share the link to the playlist via social media or email.
[2060] This system allows users to enjoy new musical experiences and provides a personalized musical experience that is tailored to their emotions.
[2061] The flow of the identification process in the second embodiment will be described with reference to FIG.
[2062] Step 1:
[2063] User Login
[2064] The user launches the app, enters their user ID and password, and clicks the "Login" button. The device sends this information to the server, which then authenticates them by checking it against the user information in its database. If authentication is successful, the server sends the home screen data to the device, and the device displays the home screen.
[2065] Input: User ID, Password
[2066] Output: Home screen display
[2067] Specific operation: The device sends input to the server, the server performs authentication, generates a home screen and sends it to the device, which then displays the home screen.
[2068] Step 2:
[2069] Search and select singers and songs
[2070] The user enters "Singer A" in the search bar, and the device sends this query to the server. The server retrieves related singer information from the database and sends it to the device. The user selects "Singer A" from the displayed candidates, then enters "Song B" in the song search bar to search and select a song in the same way. This selection information is sent to the server.
[2071] Input: Artist name, song name
[2072] Output: Artist and song selection information
[2073] Specific operation: The device sends a search query to the server, the server retrieves candidates from the database and sends them to the device, the user enters selection information, which is sent to the server.
[2074] Step 3:
[2075] Data Acquisition and Speech Synthesis
[2076] The server retrieves voice data and music score data from a database based on the received singer name and song title. This data is passed to a voice synthesis engine, which synthesizes music according to the specified conditions. The resulting voice data is returned to the server and sent to the user's device.
[2077] Input: Artist name, song name
[2078] Output: Synthesized speech data
[2079] Specific operation: The server retrieves the necessary data from the database, sends a request to the speech synthesis engine to synthesize speech, and sends the result to the user's device.
[2080] Step 4:
[2081] Music playback and emotion recognition
[2082] The user clicks on the music playback link from their device. The device plays the music data and sends the user's voice and text input during playback to the emotion engine. The emotion engine analyzes this and sends emotional data to the server.
[2083] Input: Music link, voice or text input
[2084] Output: Emotion data
[2085] Specific operation: The device plays music, sends the data entered by the user to the emotion engine, and sends the analysis results to the server.
[2086] Step 5:
[2087] Providing recommended content
[2088] The server references the database based on the emotion data received from the emotion engine to obtain suitable singers and song candidates. It then generates a list of recommended content and sends it to the user's device. The device then displays the recommended content.
[2089] Input: Emotion data
[2090] Output: Recommended content list
[2091] Specific operation: The server analyzes the emotion data, generates recommended content, and sends it to the device. The device then displays the recommended content.
[2092] Step 6:
[2093] Create and share playlists
[2094] The user adds the songs they played to a playlist and accesses the playlist creation screen. The device sends the list of songs selected by the user to the server, which then saves the playlist information in a database and generates a playlist ID. The generated link is provided to the user's device, and the user can share it on social media, etc.
[2095] Input: Song list
[2096] Output: Playlist link
[2097] Specific operation: The device sends the song list to the server, which stores it in a database, generates a link, and sends it to the device. The user can then use the link to share the playlist.
[2098] (Application example 2)
[2099] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[2100] Current music distribution services limit users to simply selecting and playing songs, and do not adequately provide a personalized music experience based on their emotions. Furthermore, few services recognize users' input emotions and recommend content based on those emotions, making it difficult to provide a music experience that best suits the user's emotions. The objective of this invention is to provide a system that not only plays singers and songs freely selected by the user, but also recognizes the user's emotions in real time and provides recommended content based on those emotions.
[2101] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[2102] In this invention, the server includes means for processing singer and song selection information received from a user terminal, means for retrieving singer voice data and song score data from a database based on the received selection information, means for synthesizing music using a voice synthesis engine using the retrieved voice data and song score data, means for transmitting the synthesized voice data to the user terminal, means for processing emotion data input by the user and recommending related information, means for recommending and providing content such as news and music based on the emotion data, and means for recognizing the emotion input by the user, thereby enabling a personalized music experience that matches the user's emotions.
[2103] "User terminal" means a device that a user accesses and uses to input and make selections.
[2104] "Selection information" is information about the singers and songs selected by the user.
[2105] "Audio data" refers to data that stores the voice of a particular singer in digital form.
[2106] "Music score data" refers to data that stores the melody and chords of a particular piece of music in digital format.
[2107] A "voice synthesis engine" is software or a system that generates new music using voice data and musical score data.
[2108] "Emotional data" refers to data about emotions derived from user input and actions.
[2109] "Recommended content" refers to music and information provided based on a user's emotional data.
[2110] A "search interface" is a screen or input field that allows a user to search for a particular artist or song.
[2111] A "playlist" is a function that organizes songs into a list, making it easy to play them consecutively and share them.
[2112] A "unique link" is a URL or hyperlink that provides access to a specific playlist or content.
[2113] This invention is a system that plays back singer-music combinations freely selected by the user in real time, and also recognizes the user's emotions to provide recommended content. The system is composed of a user terminal, a server, a database, a speech synthesis engine, and an emotion engine.
[2114] Overall system configuration
[2115] 1. User Device
[2116] The user terminal is a device such as a smartphone, tablet, or PC, and provides a user interface. From the user interface, users can search for and select singers and songs. It also has the function of sending the user's voice and text input to the emotion engine.
[2117] 2. Server
[2118] The server is responsible for processing the singer and song selection information received from the user device, as well as the emotion data from the emotion engine. The server analyzes the received data and retrieves the singer's voice data and song score data from the database. It also sends requests to the voice synthesis engine and transmits the synthesized voice data to the user device. It also has the function of generating recommended content based on the emotion data.
[2119] 3. Database
[2120] The database stores the singer's voice data, the music score data, and the user's emotion history data. After the server receives the request, it retrieves the necessary data from the database.
[2121] 4. Speech synthesis engine
[2122] The voice synthesis engine uses the acquired singer's voice data and the music score data to synthesize a song according to the specified conditions, and the synthesized voice data is sent to the user's device via the server.
[2123] 5. Emotion Engine
[2124] The emotion engine recognizes emotions from the user's voice or text and sends the emotion data to the server. The recognized emotion data is stored in a database along with the user's emotion history data and used to generate recommended content.
[2125] Program processing and specific examples
[2126] The program supports a series of operations, from user login to playing music, recognizing emotions, providing recommended content, and sharing. While the program's specific processing steps are not included, the following concrete examples are provided:
[2127] Specific examples
[2128] 1. Log in
[2129] The user logs in to the app. They enter their username and password, and the server authenticates them. If authentication is successful, the user is taken to the home screen.
[2130] 2. Selection of singers and songs
[2131] Users can enter a specific artist in the search bar and choose from a list of artists, then enter a song name in the song search bar and choose from a list of songs.
[2132] 3. Data Acquisition and Speech Synthesis
[2133] The server processes the selection information, retrieves the respective data from the database, and passes this data to the speech synthesis engine to synthesize the music.
[2134] 4. Music playback and emotion recognition
[2135] The server sends the synthesized voice data to the user's device. The user clicks the provided link and plays the music. During playback, the user's voice and text input are analyzed by the emotion engine.
[2136] 5. Providing recommended content
[2137] The emotion data analyzed by the emotion engine is sent to the server, which then recommends singers and songs that are suitable for the user based on the emotion data.
[2138] 6. Create and share playlists
[2139] Users can save the songs they play as playlists and share them on social media. Users can easily share music with friends and family using the playlist link.
[2140] Prompt Sentence Examples
[2141] "User text input: 'I'm feeling a bit down...'"
[2142] "The generative AI model's output: 'That's sad. Let's pick a song that's uplifting.'"
[2143] In this way, the present invention is a system that provides recommended content that corresponds to the user's emotions, thereby realizing a personalized music experience.
[2144] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[2145] Step 1:
[2146] A user logs in to the app.
[2147] Input: User ID, Password
[2148] Output: Authentication result (success / failure), home screen
[2149] Specific operation: The user enters their user ID and password on the login screen and sends them to the server. The server retrieves the relevant information from the database and compares it with the entered information. If authentication is successful, the user is taken to the home screen; if it is unsuccessful, an error message is displayed.
[2150] Step 2:
[2151] The user searches for and selects an artist and song.
[2152] Input: Search keyword (singer name, song name)
[2153] Output: Search result list, selection information
[2154] Specific operation: The user enters the name of an artist or song in the search bar and presses the Enter key. The device sends the input information to the server, which retrieves the corresponding artist and song information from the database and sends a list of search results to the user's device. The user then selects a specific artist and song from the list.
[2155] Step 3:
[2156] The server retrieves the necessary data from the database based on the selection information.
[2157] Input: Selection information (singer name, song name)
[2158] Output: Singer's voice data, music score data
[2159] Specific operation: The server analyzes the selection information received from the user terminal and sends a request to the database, which then returns the corresponding singer's voice data and the music score data to the server.
[2160] Step 4:
[2161] The server synthesizes the music using a speech synthesis engine.
[2162] Input: Audio data, music score data
[2163] Output: Synthesized speech data
[2164] Specific operation: The server passes the acquired voice data and musical score data to the speech synthesis engine, which synthesizes the music. The speech synthesis engine processes the data and generates new synthetic voice data, which is then sent back to the server.
[2165] Step 5:
[2166] The synthesized speech data is transmitted to the user terminal.
[2167] Input: Synthetic speech data
[2168] Output: Music playback link
[2169] Specific operation: The server sends the synthesized voice data to the user's device and provides a music playback link. The user can click the link to play the synthesized music in real time.
[2170] Step 6:
[2171] The emotion engine analyzes the user's voice and text input.
[2172] Input: User voice or text input
[2173] Output: Emotion data
[2174] Specific operation: The user inputs voice or text into the emotion engine while playing music. The emotion engine receives the data, analyzes it, and generates emotion data as a result. The emotion data is then sent to the server.
[2175] Step 7:
[2176] Based on the emotion data, the server provides recommended content.
[2177] Input: Emotion data
[2178] Output: Recommended singer and song information
[2179] Specific operation: The server analyzes the emotion data and generates recommended content based on it. Specifically, it selects appropriate singers and songs from the database and sends that information to the user's device.
[2180] Step 8:
[2181] Users can save the songs they play as a playlist and share it on social media.
[2182] Input: Information about the song played
[2183] Output: Playlist link
[2184] Specific operation: The user selects the songs they played and saves them as a playlist. The server saves the playlist information in a database and generates a dedicated link. The generated link is sent to the user's device, and the user can share it via social media or with other users.
[2185] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[2186] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[2187] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[2188] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[2189] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[2190] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[2191] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[2192] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[2193] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[2194] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[2195] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[2196] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[2197] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[2198] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[2199] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[2200] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[2201] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[2202] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[2203] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[2204] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[2205] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[2206] The following is further disclosed regarding the above embodiment.
[2207] (Claim 1)
[2208] means for processing singer and song selection information received from a user terminal;
[2209] means for acquiring singer voice data and musical score data of a song from a database based on the received selection information;
[2210] means for synthesizing music using a voice synthesis engine using the acquired voice data and musical score data;
[2211] means for transmitting the synthesized voice data to a user terminal;
[2212] A system including:
[2213] (Claim 2)
[2214] means for providing a search interface for a user to select an artist and song;
[2215] means for transmitting input from said search interface to a server;
[2216] means for providing a synthesized speech link to the user based on the received selection;
[2217] 10. The system of claim 1, comprising:
[2218] (Claim 3)
[2219] A way for users to save songs and create playlists to share on social media,
[2220] means for storing said playlist in a database and generating a dedicated link;
[2221] means for providing the generated link to a user terminal;
[2222] 10. The system of claim 1, comprising:
[2223] "Example 1"
[2224] (Claim 1)
[2225] means for processing selection information of audio data and music data received from a user terminal;
[2226] means for acquiring the audio data and the musical score data from the storage device based on the received selection information;
[2227] means for synthesizing music using a voice synthesizer using the acquired voice data and musical score data;
[2228] means for transmitting the synthesized voice data to a user terminal;
[2229] A system including:
[2230] (Claim 2)
[2231] means for providing a search screen for a user to select audio data and music data;
[2232] means for transmitting input from the search screen to a server;
[2233] means for providing a synthesized speech link to the user based on the received selection;
[2234] 10. The system of claim 1, comprising:
[2235] (Claim 3)
[2236] A means for a user to save songs and create a song selection list for sharing on an information sharing service;
[2237] means for storing the music selection list in a storage device and generating a dedicated link;
[2238] means for providing the generated link to a user terminal;
[2239] 10. The system of claim 1, comprising:
[2240] "Application Example 1"
[2241] (Claim 1)
[2242] means for processing singer and song selection information received from a user terminal;
[2243] means for acquiring singer voice data and musical score data of a song from a database based on the received selection information;
[2244] means for synthesizing music using a voice synthesis engine using the acquired voice data and musical score data;
[2245] means for transmitting the synthesized voice data to a user terminal;
[2246] a means for a user to play and share the song using a link of the synthesized speech data;
[2247] A system including:
[2248] (Claim 2)
[2249] means for providing a search interface for a user to select an artist and song;
[2250] means for transmitting input from said search interface to a server;
[2251] means for providing a synthesized speech link to the user based on the received selection;
[2252] A way to share the synthesized voice link on social media or messaging applications;
[2253] 10. The system of claim 1, comprising:
[2254] (Claim 3)
[2255] A way for users to save songs and create playlists to share on social media,
[2256] means for storing said playlist in a database and generating a dedicated link;
[2257] means for providing the generated link to a user terminal;
[2258] means for playing songs in the playlist in real time;
[2259] 10. The system of claim 1, comprising:
[2260] "Example 2: Combining Emotion Engines"
[2261] (Claim 1)
[2262] means for processing singer and song selection information received from a user terminal;
[2263] means for acquiring singer voice data and musical score data of a song from a database based on the received selection information;
[2264] means for synthesizing music using a voice synthesis engine using the acquired voice data and musical score data;
[2265] means for transmitting the synthesized voice data to a user terminal;
[2266] A means of transmitting user voice and text input to the emotion engine;
[2267] means for recognizing user emotion data and sending it to a server for generating recommended content;
[2268] means for transmitting recommended content to a user terminal;
[2269] A system including:
[2270] (Claim 2)
[2271] means for providing a search interface for a user to select an artist and song;
[2272] means for transmitting input from said search interface to a server;
[2273] means for providing a synthesized speech link to the user based on the received selection;
[2274] means for recommending singers and songs based on the user's emotional data;
[2275] 10. The system of claim 1, comprising:
[2276] (Claim 3)
[2277] A way for users to save songs and create playlists to share on social media,
[2278] means for storing said playlist in a database and generating a dedicated link;
[2279] means for providing the generated link to a user terminal;
[2280] 10. The system of claim 1, comprising:
[2281] "Application example 2 when combining emotion engines"
[2282] (Claim 1)
[2283] means for processing singer and song selection information received from a user terminal;
[2284] means for acquiring singer voice data and musical score data of a song from a database based on the received selection information;
[2285] means for synthesizing music using a voice synthesis engine using the acquired voice data and musical score data;
[2286] means for transmitting the synthesized voice data to a user terminal;
[2287] means for processing the user's input emotion data and recommending relevant information;
[2288] A means to recommend and provide content such as news and music based on emotional data;
[2289] means for recognizing an emotion input by a user;
[2290] A system including:
[2291] (Claim 2)
[2292] means for providing a search interface for a user to select an artist and song;
[2293] means for transmitting input from said search interface to a server;
[2294] means for providing a synthesized speech link to the user based on the received selection;
[2295] means for providing recommended content links based on a user's emotions via emotion recognition;
[2296] 10. The system of claim 1, comprising:
[2297] (Claim 3)
[2298] A way for users to save songs and create playlists to share on social media,
[2299] means for storing said playlist in a database and generating a dedicated link;
[2300] means for providing the generated link to a user terminal;
[2301] A method for automatically generating playlists based on emotional data,
[2302] 10. The system of claim 1, comprising: [Explanation of symbols]
[2303] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. means for processing singer and song selection information received from a user terminal; means for acquiring singer voice data and musical score data of a song from a database based on the received selection information; means for synthesizing music using a voice synthesis engine using the acquired voice data and musical score data; means for transmitting the synthesized voice data to a user terminal; A system including:
2. means for providing a search interface for a user to select an artist and song; means for transmitting input from said search interface to a server; means for providing a synthesized speech link to the user based on the received selection; The system of claim 1 , comprising:
3. A way for users to save songs and create playlists to share on social media, means for storing said playlist in a database and generating a dedicated link; means for providing the generated link to a user terminal; The system of claim 1 , comprising:
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A