System
The system addresses the inefficiencies in conventional audiobooks by allowing voice-controlled operation and instant term explanations, improving user convenience and understanding of technical content.
Patent Information
- Application Number
- JP2024116542
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-19
- Publication Date
- 2026-01-29
AI Technical Summary
Conventional audiobooks require users to rewind or skip playback to find specific information, which is time-consuming, and systems relying on voice control lack accurate speech recognition and subsequent processing, making it difficult to understand technical or new terminology.
A system with a voice input unit, voice recognition unit, control unit, search unit, and voice synthesis unit that allows users to play, stop, rewind, and skip audiobooks by voice, and provides instant explanations of technical terms and new terms, with functions to resume playback from the previous position and adjust reading speed and tone.
Improves user convenience by enabling quick and accurate operation of audiobooks and providing immediate explanations of technical terms, enhancing the listening experience.
Smart Images

Figure 2026015068000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] Create a "problem the invention aims to solve" and a "means for solving the problem."
[0005] In conventional audiobooks, users have to rewind or skip playback to check specific information, which is time-consuming and makes it difficult to understand technical or new terminology. Furthermore, systems that rely on voice control require highly accurate speech recognition and subsequent processing to properly recognize user instructions. The present invention aims to solve these problems and improve the audiobook listening experience. [Means for solving the problem]
[0006] The present invention provides a system that includes a voice input unit, a voice recognition unit that converts voice data into text, a control unit that analyzes the text data and controls audio content, a search unit that provides explanations of technical terms and new terms, and a voice synthesis unit that outputs search results by voice. This system allows users to play, stop, rewind, and skip audiobooks by voice, and instantly obtain explanations of technical terms and new terms. The system also provides functions to resume playback from the previous playback position and to adjust the reading speed and tone by voice, significantly improving user convenience.
[0007] Understood. Below are definitions for important terms included in the claims.
[0008] The "voice input means" is a device or module for capturing the user's speech as voice data.
[0009] "Speech recognition means" refers to the technology or algorithms used to convert voice data into text data.
[0010] The "control means" refers to a module or program that analyzes text data and performs processing based on the operation of audio content or user instructions.
[0011] "Search methods" refer to technologies and databases for obtaining the meanings and explanations of technical terms and new terms.
[0012] "Speech synthesis means" refers to the technology or algorithm for converting text data into speech data and outputting it.
[0013] "Audio content" means content such as books, papers, articles, etc. that is provided in audio format.
[0014] "User" refers to a person who uses the system.
[0015] The "previous playback position" is the point where the user last stopped playback of the audio content.
[0016] "Reading speed" refers to the playback speed of the voice generated by speech synthesis.
[0017] "Tone" refers to the tone or pitch of the voice generated by speech synthesis. [Brief explanation of the drawings]
[0018] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0019] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0020] First, the terms used in the following description will be explained.
[0021] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0022] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0023] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0024] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0025] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0026] [First embodiment]
[0027] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0028] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0029] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0030] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0031] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0032] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0033] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0034] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0035] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0036] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0037] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0038] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0039] Understood. Below is the description of the "Mode for Carrying Out the Invention" of the patent specification.
[0040] The present invention provides a system including a voice input unit, a voice recognition unit, a control unit, a search unit, and a voice synthesis unit. This system allows a user to easily operate audio content by voice and quickly obtain explanations of technical terms and new terms. Specific embodiments of the present invention are described in detail below.
[0041] System configuration:
[0042] 1. Voice input method:
[0043] The device is equipped with a microphone that captures the user's speech as voice data.
[0044] This voice input means provides a basis for responding to operational commands and questions uttered by the user.
[0045] 2. Voice recognition means:
[0046] The captured voice data is sent to a voice recognition means and converted into text data.
[0047] This speech recognizer uses advanced algorithms to extract accurate text from speech.
[0048] 3. Control measures:
[0049] The converted text data is analyzed by the control means and the corresponding operation is performed.
[0050] For example, if the user says "play," the control will begin playing audio content.
[0051] 4. Search methods:
[0052] When a user requests an explanation of a particular term, the control means uses the search means to quickly obtain the meaning of that term.
[0053] The search tool accesses an extensive database and retrieves detailed explanations of technical terms.
[0054] 5. Speech synthesis means:
[0055] The acquired text information is passed to a voice synthesis means and converted into voice data.
[0056] The speech synthesis means generates natural speech in real time and provides audible feedback to the user.
[0057] Explanation of program operation:
[0058] 1. Speech Recognition:
[0059] The device captures the user's speech and converts it into text using a speech recognition service.
[0060] The converted text data is passed to the next processing step.
[0061] 2. Command analysis:
[0062] The control means analyzes the converted text data and determines the user's intention.
[0063] Specific actions are determined in response to control commands (e.g., play, stop, rewind, etc.) and questions (e.g., "What does AI mean?").
[0064] 3. Perform a search:
[0065] When an explanation of a technical term is required, the control means calls the search means and retrieves the meaning of the relevant term from the database.
[0066] The information obtained is stored as text.
[0067] 4. Audio output:
[0068] The final information is passed to a speech synthesis means and converted into speech data.
[0069] The terminal plays the generated voice data to the user through a speaker.
[0070] Examples:
[0071] 1. Playback instructions:
[0072] When the user utters "play", the terminal captures this speech and converts it into text "play" through speech recognition means.
[0073] The control means recognizes the command "play" and starts playing the audio content.
[0074] 2. Terminology:
[0075] If a user says, "Please explain artificial intelligence," the device converts the speech into text and recognizes it as "Please explain artificial intelligence."
[0076] The control means determines that an explanation of "artificial intelligence" is desired, and uses the search means to obtain the meaning of "artificial intelligence."
[0077] The acquired commentary is converted into voice data by a voice synthesis means and provided to the user.
[0078] In this way, the system of the present invention can meet the diverse needs of users, improving operability and convenience.
[0079] The processing flow will be explained below.
[0080] Understood. Below, I will explain the process flow in concrete steps.
[0081] Step 1:
[0082] The user speaks into the device's microphone to give voice commands such as play, stop, rewind, skip, or explain technical terms.
[0083] Step 2:
[0084] The device captures the user's speech as voice data through a microphone.
[0085] Step 3:
[0086] The voice data captured by the device is sent to the voice recognition means and converted into text data using the Google voice recognition service.
[0087] Step 4:
[0088] The terminal receives the converted text data and passes it to the control means.
[0089] Step 5:
[0090] The control means of the terminal analyzes the text data to determine the user's intended command (eg, play, stop, rewind, skip, or commentary).
[0091] Step 6:
[0092] If the command is "play," "stop," "rewind," or "skip," the device's control invokes the corresponding audio player function (e.g., the play, pause, rewind, or skip method) to perform the audio content operation.
[0093] Step 7:
[0094] If the command requests an explanation of a technical term, the control means of the terminal utilizes the search means to obtain an explanation of the specified term from the database.
[0095] Step 8:
[0096] A search means on the server searches a database for an explanation of the requested technical term and returns it to the terminal in text format.
[0097] Step 9:
[0098] The explanatory text acquired by the terminal is passed to a speech synthesis means, which generates voice data using TTS (Text-To-Speech).
[0099] Step 10:
[0100] The device outputs the generated voice data to the user through a speaker, providing feedback in response to explanations and commands.
[0101] The above are the steps for processing voice instructions in the system of the present invention. This detailed processing procedure allows the user to intuitively operate audio content and receive explanations of technical terms.
[0102] Example 1
[0103] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0104] In conventional voice-operated systems, when users operate audio content by voice, feedback on operations and explanations of technical terms can be delayed, resulting in a lack of convenience. Additionally, there is a lack of a mechanism for rapid and accurate voice recognition and subsequent processing.
[0105] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0106] In this invention, the server includes a voice input means, a voice recognition means for converting voice data into text, a control means for analyzing the text data and operating the media content, a search means for providing explanations of technical terms and undefined terms, and a voice synthesis means for outputting the obtained explanations by voice, thereby enabling quick and accurate provision of information and operation of the media content by voice operation.
[0107] 1. "Voice input means" means any device or technology used to capture a user's speech as voice data.
[0108] 2. "Speech recognition means" means a technology or device for converting voice data captured by a voice input means into text data.
[0109] 3. "Control means" refers to the technology or device that analyzes the text data converted by the speech recognition means and controls the media content.
[0110] 4. "Search means" refers to the technology or device that allows users to obtain explanations of technical terms or undefined terms they are looking for.
[0111] 5. "Speech synthesis means" refers to technology or equipment for outputting the acquired commentary in audio form.
[0112] 6. "Media Content" refers to all content, including audio and video, that can be used through voice control.
[0113] 7. "Prompt" means a voice command that allows a user to give verbal instructions to a system.
[0114] The present invention provides a system including a voice input unit, a voice recognition unit, a control unit, a search unit, and a voice synthesis unit. This system allows users to easily operate media content by voice and quickly obtain explanations of technical terms and new terms. Specific embodiments of the present invention are described in detail below.
[0115] System configuration
[0116] 1. Voice input method:
[0117] The device is equipped with a microphone to capture the user's speech as voice data. This voice input means provides a basis for responding to user commands and questions.
[0118] 2. Voice recognition means:
[0119] The voice data captured by the device is sent to a server, which converts the voice data into text using a speech recognition service such as the Google Cloud Speech-to-Text API, which uses advanced algorithms to extract accurate text from the speech.
[0120] 3. Control measures:
[0121] The server analyzes the converted text data to determine the user's intent and determines the specific actions to take in response to commands (e.g., play, stop, rewind, etc.) or questions (e.g., "What does AI mean?").
[0122] 4. Search methods:
[0123] When a user requests an explanation of a particular term, the server uses a search engine to quickly retrieve the meaning of the term, accessing the Wikipedia API and other extensive databases to retrieve detailed explanations of technical terms.
[0124] 5. Speech synthesis means:
[0125] The text information acquired by the server is converted into voice data using a speech synthesis service such as Amazon Polly, which generates natural-sounding voices in real time and provides voice feedback to the user.
[0126] Specific examples
[0127] 1. Playback instructions:
[0128] When the user says "play," the terminal captures this speech and converts it into text through speech recognition means as "play." The server recognizes the command "play" and starts playing the media content.
[0129] 2. Terminology:
[0130] When a user says, "Please explain artificial intelligence," the device converts the speech into text and recognizes it as "Please explain artificial intelligence." The server determines that an explanation of "artificial intelligence" is being sought, and uses a search means to obtain the meaning of "artificial intelligence." The obtained explanation is converted into voice data by a voice synthesis means and provided to the user.
[0131] Prompt Sentence Examples
[0132] 1. Playback request prompt:
[0133] When someone says "play," use speech recognition to convert it to text and provide instructions for starting to play media content.
[0134] 2. Glossary prompt:
[0135] If someone says "Please explain artificial intelligence," please provide instructions on how to use the Wikipedia API to obtain an explanation of artificial intelligence, synthesize it into voice, and provide it to the user.
[0136] The system of the present invention is thus able to meet the diverse needs of users and greatly improves operability and convenience.
[0137] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0138] Step 1: Voice Input
[0139] Users speak voice commands to control media content and to explain technical terms.
[0140] The terminal uses a microphone to capture the user's speech as voice data.
[0141] Input: User speech
[0142] Output: Audio data
[0143] Step 2: Voice Recognition
[0144] The terminal transmits the captured audio data to the server.
[0145] The server converts the voice data into text data using the Google Cloud Speech-to-Text API.
[0146] Input: Audio data
[0147] Data processing: Converting voice data into text data
[0148] Output: Text data
[0149] Step 3: Command analysis
[0150] The server analyzes the text data obtained by the voice recognition means and determines whether the spoken content is an operation command or a question.
[0151] For example, text data such as "play" is determined to be an operation command, and text data such as "please explain artificial intelligence" is determined to be a question.
[0152] Input: Text data
[0153] Data operation: Content analysis and category determination of text data (operation commands or questions)
[0154] Output: The result of the operation command or question
[0155] Step 4: Execute the operation command
[0156] If the server recognizes the operation command, it performs the corresponding operation on the media content.
[0157] For example, the command "play" will call the play function of the media player.
[0158] Input: Operation command
[0159] Data operation: Executes media operations corresponding to commands.
[0160] Output: Playing, stopping, etc. media content
[0161] Step 5: Question Processing
[0162] When the server recognizes the question, it uses a search means to obtain explanations of technical terms and undefined terms.
[0163] For example, in response to the question "Please explain artificial intelligence," an explanation of "artificial intelligence" is retrieved from the Wikipedia API.
[0164] Input: Question text data
[0165] Data processing: Performing searches that correspond to the query
[0166] Output: Search results text data
[0167] Step 6: Text-to-Speech
[0168] The server converts the acquired text data of the explanation into audio data using Amazon Polly.
[0169] Input: Text data explaining the question
[0170] Data processing: Convert text data into audio data
[0171] Output: Audio data
[0172] Step 7: Audio Output
[0173] The terminal outputs the audio data transmitted from the server to the user through a speaker.
[0174] Input: Audio data
[0175] Output: Play audio
[0176] This completes the processing flow. The system of the present invention has the advantage that users can quickly and accurately use information and media content through voice operations.
[0177] (Application example 1)
[0178] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0179] Conventional content distribution services require users to physically operate the service and check search results, resulting in problems with usability and convenience. In particular, when requesting explanations of technical terms, users are often forced to search and check text information, which often interrupts the viewing experience. While systems that allow voice control exist, they lack the functionality to simultaneously operate viewing content and provide terminology explanations in real time. This degrades the user experience and makes it difficult to provide a comfortable viewing environment.
[0180] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0181] In this invention, the server includes a voice input means, a voice recognition means for converting voice data into text, a control means for analyzing the text data and operating the voice and audiovisual content, a search means for providing explanations of technical terms and new terms, a voice synthesis means for outputting the search results by voice, and a means for providing audio instructions for operating the audiovisual content and explanations of terms in real time. This allows the user to intuitively operate the audiovisual content by voice instructions while instantly obtaining audio explanations of technical terms, thereby creating an optimal usage environment.
[0182] "Voice input means" refers to a device or system, such as a microphone, for capturing a user's speech as voice data.
[0183] A "voice recognition means" is a technique or system for converting captured voice data into text data.
[0184] The "control means" is a system or method that analyzes text data and operates the audiovisual content based on the analysis.
[0185] A "search tool" is a system or method for quickly retrieving information about a particular term or terminology.
[0186] "Speech synthesis means" refers to a system or technology that converts text information into voice data and provides it to the user in voice form.
[0187] "Viewing content" refers to media content such as audio and video that a user views.
[0188] A "real-time audio feedback mechanism" is a system or method for providing immediate audio feedback to a user upon request.
[0189] System configuration
[0190] This invention provides a system that allows a user to operate viewing content through voice input and obtain audio explanations of technical terms in real time. The system includes the following main means.
[0191] 1. Voice input method:
[0192] - A microphone is used as a voice input means to capture the user's speech as voice data. A microphone built into a smartphone, smart glasses, a head-mounted display, or a robot is used.
[0193] 2. Voice recognition means:
[0194] - Use speech recognition software such as Google Cloud Speech-to-Text API to convert the captured voice data into text data, thereby capturing the user's voice instructions as text.
[0195] 3. Control measures:
[0196] - Analyze text data and manipulate viewing content. Use Python and natural language processing libraries (NLTK and spaCy) to analyze text data from speech recognition tools and perform operations such as play, stop, and rewind.
[0197] 4. Search methods:
[0198] - Use Elasticsearch or other database search technologies to provide explanations of technical terms and new terms. Search for explanations of terms captured as text and return them in text format.
[0199] 5. Speech synthesis means:
[0200] - Use speech synthesis technology such as Google Cloud Text-to-Speech API to convert text information obtained from search tools into audio data and provide feedback to users.
[0201] 6. Real-time feedback methods:
[0202] - Provide audio instructions and glossary of terms for the content in real time, ensuring seamless operation and commentary of the content.
[0203] Usage and Examples
[0204] In a real-world usage scenario, a user will interface with the system using a smartphone or other compatible device. The following example illustrates how to operate the system.
[0205] 1. Playback instructions:
[0206] - When the user says "play," the microphone captures the voice and converts it into text "play" using the Google Cloud Speech-to-Text API. The control means recognizes this and starts playing the content.
[0207] 2. Terminology:
[0208] - When a user says, "Please explain artificial intelligence," the speech is converted into text in the same way. This text is analyzed by the control means, and an explanation of "artificial intelligence" is retrieved from the Elasticsearch database through the search means. The retrieved text information is converted into audio data using the Google Cloud Text-to-Speech API and provided to the user in real time.
[0209] Example prompt sentence:
[0210] "Explain the following statement: 'Artificial intelligence refers to systems designed to have human intelligence.'"
[0211] These methods allow users to intuitively control content through voice commands and instantly receive explanations of technical terms, improving the user experience and ease of use.
[0212] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0213] Step 1:
[0214] The user speaks a voice command to the device (e.g., "play," "please explain with AI," etc.). The input is the user's voice, and the output is the audio data captured by the device's microphone.
[0215] Step 2:
[0216] The terminal captures the user's voice using the voice input means and obtains it as voice data. The input is the user's voice and the output is voice data.
[0217] Step 3:
[0218] The device sends voice data to the Google Cloud Speech-to-Text API, which converts the voice data into text. The input is voice data, and the output is text data. The server calls a Google Cloud service to convert the voice data.
[0219] Step 4:
[0220] The server uses a control means to analyze the received text data and determine the operation and explanation acquisition instructions based on that. The input is text data, and the output is the operation instructions or search query as the analysis result. The specific operation is to determine the user's intention using a text data analysis algorithm.
[0221] Step 5:
[0222] When the server operates the viewing content based on the analysis results, it executes operations such as play, stop, and rewind. The input is the operation command, and the output is the actual content operation. The specific operation is to call the media player's control API and execute the required operation.
[0223] Step 6:
[0224] When a term explanation is requested, the server uses a search mechanism to search for relevant information in a database. The input is a search query, and the output is text information as search results. The specific operation is to retrieve the appropriate information using Elasticsearch or other database search technologies.
[0225] Step 7:
[0226] The server sends the text information acquired through the search tool to the Google Cloud Text-to-Speech API and converts it into voice data. The input is text information and the output is voice data. The specific operation is to generate voice from text using the voice synthesis API.
[0227] Step 8:
[0228] The terminal plays back the voice data obtained from the voice synthesis means and provides real-time feedback to the user. The input is voice data, and the output is voice feedback to the user. The specific operation is to play back the voice data using the terminal's speaker.
[0229] Step 9:
[0230] The flow from step 1 to step 8 is repeated until the user utters a new command or performs the next operation. The input is the next voice command, and the output is the capture of the next voice data.
[0231] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0232] Understood. Below is a description of the "Mode for carrying out the invention" based on the invention that combines an emotion engine.
[0233] The present invention provides a system including a voice input unit, a voice recognition unit, a control unit, a search unit, a voice synthesis unit, and an emotion recognition unit. This system allows users to easily operate audio content by voice, quickly obtain explanations of technical terms and new terms, and also realizes more appropriate feedback and content suggestions based on the user's emotions. Specific embodiments of the present invention are described in detail below.
[0234] System configuration:
[0235] 1. Voice input method:
[0236] The device is equipped with a microphone that captures the user's speech as voice data.
[0237] The voice input means provides a basis for responding to operational commands and questions uttered by the user.
[0238] 2. Voice recognition means:
[0239] The captured voice data is sent to a voice recognition means and converted into text data.
[0240] The speech recognizer uses advanced algorithms to extract accurate text from speech.
[0241] 3. Emotion recognition means:
[0242] Analyze user emotions from captured voice data.
[0243] The emotion recognition means identifies the emotional state of the user based on information such as the tone, speed, and strength of the user's voice.
[0244] 4. Control measures:
[0245] The converted text data and the emotion data obtained by the emotion recognition means are analyzed, and a corresponding operation is performed.
[0246] For example, when the user says "play," the control means starts playing audio content and adjusts the speaking speed and tone based on the user's emotions.
[0247] 5. Search methods:
[0248] When a user requests an explanation of a particular term, the control means uses the search means to quickly obtain the meaning of that term.
[0249] The search tool accesses an extensive database and retrieves detailed explanations of technical terms.
[0250] 6. Speech synthesis means:
[0251] The acquired text information is passed to a voice synthesis means and converted into voice data.
[0252] The speech synthesis means generates natural speech in real time and provides audible feedback to the user.
[0253] Explanation of program operation:
[0254] 1. Speech Recognition:
[0255] The device captures the user's speech and converts it into text using a speech recognition service.
[0256] The converted text data is passed to the next processing step.
[0257] 2. Emotion analysis:
[0258] The emotion recognition means of the terminal analyzes the user's emotions from the captured voice data.
[0259] The results of the sentiment analysis are passed to the control means and used in the next processing step.
[0260] 3. Command Analysis:
[0261] The control means analyzes the converted text data and the emotion analysis data from the emotion recognition means to determine the user's intention.
[0262] Specific actions corresponding to operation commands (e.g., play, stop, rewind) or questions (e.g., "What does XX mean?") are determined.
[0263] 4. Perform a search:
[0264] When an explanation of a technical term is required, the control means calls the search means and retrieves the meaning of the relevant term from the database.
[0265] The information obtained is stored as text.
[0266] 5. Adjustment and Audio Output:
[0267] Based on the emotion recognition results, the reading speed and tone of the audio content are adjusted.
[0268] The final information is passed to a speech synthesis means and converted into speech data.
[0269] The device outputs the generated voice data to the user through a speaker, providing feedback in response to explanations and commands.
[0270] Examples:
[0271] 1. Playback instructions:
[0272] When the user utters "play", the terminal captures this speech and converts it into text "play" through speech recognition means.
[0273] The control means recognizes the command "play" and starts playing the audio content. Furthermore, the emotion recognition means analyzes the user's emotion, and if the user is nervous, for example, starts playing the audio content in a calm tone.
[0274] 2. Terminology:
[0275] If a user says, "Please explain artificial intelligence," the device converts the speech into text and recognizes it as "Please explain artificial intelligence."
[0276] The control means determines that an explanation of "artificial intelligence" is desired, and uses the search means to obtain the meaning of "artificial intelligence."
[0277] The acquired commentary is converted into voice data by a voice synthesis means and provided to the user. If the user shows interest, the emotion recognition means can detect this and continue with a more detailed commentary.
[0278] In this way, the system of the present invention can meet the diverse needs of users, not only improving operability and convenience but also providing a more personalized experience according to the user's emotions.
[0279] The processing flow will be explained below.
[0280] Understood. Below I will explain the process in concrete steps.
[0281] Step 1:
[0282] The user speaks into the device's microphone to give voice commands such as play, stop, rewind, skip, or explain technical terms.
[0283] Step 2:
[0284] The device captures the user's speech as voice data through a microphone.
[0285] Step 3:
[0286] The voice data captured by the terminal is sent to a voice recognition means, which converts the voice data into text data.
[0287] Step 4:
[0288] The terminal receives the converted text data and passes it to the control means.
[0289] Step 5:
[0290] An emotion recognition means of the terminal analyzes the user's emotions from the captured voice data.
[0291] Step 6:
[0292] The control means analyzes the text data and the emotion recognition results to determine the user's intention.
[0293] Step 7:
[0294] If the command is "play," "stop," "rewind," or "skip," the control invokes the corresponding audio player function (e.g., play, pause, rewind, skip method) to perform the operation on the audio content.
[0295] Step 8:
[0296] If the command requests an explanation of technical terms, the control means uses the search means to obtain an explanation of the specified term from the database.
[0297] Step 9:
[0298] A search means on the server searches a database for an explanation of the requested technical term and returns it to the terminal in text format.
[0299] Step 10:
[0300] The explanatory text acquired by the terminal is passed to a speech synthesis means, which generates voice data using TTS (Text-To-Speech).
[0301] Step 11:
[0302] Based on the emotion recognition results, the reading speed and tone of the audio content are adjusted.
[0303] Step 12:
[0304] The device outputs the generated voice data to the user through a speaker, providing feedback in response to explanations and commands.
[0305] Examples:
[0306] Playback instructions example
[0307] 1. Step 1: The user says "play."
[0308] 2. Step 2: The device captures audio data through the microphone.
[0309] 3. Step 3: The voice data is converted into text using a speech recognition tool.
[0310] 4. Step 4: The converted text "play" is sent to the control means.
[0311] 5. Step 5: The emotion recognition means analyzes the user's emotions and determines, for example, that the user is nervous.
[0312] 6. Step 6: The control means analyzes the "play" command and also takes into account the user's emotions.
[0313] 7. Step 7: The control means calls the play method of the audio player to begin playback.
[0314] 8. Step 11: Based on the emotion recognition results, start playing at a gentle speed and tone.
[0315] 9. Step 12: The terminal outputs the playback of the audio content to the user through the speaker.
[0316] When requesting an explanation of technical terms
[0317] 1. Step 1: The user says, "Please explain to me, AI."
[0318] 2. Step 2: The device captures audio data through the microphone.
[0319] 3. Step 3: The voice data is converted into text using a speech recognition tool.
[0320] 4. Step 4: The converted text "Please explain the AI" is sent to the control means.
[0321] 5. Step 5: The emotion recognition means analyzes the user's emotion and determines, for example, that the user is interested.
[0322] 6. Step 6: The control means analyzes the text "Please explain the AI" and determines the user's intention.
[0323] 7. Step 8: The control means instructs the search means to search for an explanation of "AI."
[0324] 8. Step 9: The server's search function retrieves explanations about "AI" from the database and returns them to the terminal in text format.
[0325] 9. Step 10: The terminal passes the explanatory text to the speech synthesis means to generate speech data.
[0326] 10. Step 11: Based on the emotion recognition results, a detailed commentary is generated in an appropriate tone of voice.
[0327] 11. Step 12: The terminal outputs the generated voice data to the user through the speaker to provide an explanation.
[0328] The above are the processing steps for voice instruction and emotion recognition in the system of the present invention. This detailed processing procedure allows users to operate audio content and receive technical terminology explanations in an intuitive and personalized manner.
[0329] Example 2
[0330] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0331] Conventional speech recognition systems have the ability to operate audio content and provide information based on user voice commands. However, they have the problem of being unable to provide feedback or suggest content based on the user's emotions, which limits the user experience. Another issue is that it is difficult to quickly provide explanations of technical terms and new terms, making it difficult to meet diverse user needs.
[0332] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes a voice input means, a voice recognition means for converting voice data into text data, an emotion recognition means for analyzing emotions from the voice data, a control means for analyzing the text data and the emotion data to operate the audio content, a search means for providing explanations of technical terms and new terms, and a voice synthesis means for converting search results into voice data and outputting the voice. This enables personalized feedback and content suggestions based on the user's emotions, and also makes it possible to quickly provide explanations of technical terms and new terms.
[0333] "Voice input means" refers to a device for capturing a user's speech, and primarily refers to a microphone or similar hardware.
[0334] A "voice recognition means" is a technique or device for converting captured voice data into text data, utilizing a voice recognition algorithm.
[0335] "Emotion recognition means" refers to a means for analyzing and identifying a user's emotions from voice data, and is a technology for determining an emotional state based on information such as voice tone, speed, and strength.
[0336] "Control means" refers to a device or system for analyzing the converted text data and emotion data and for performing manipulations on the audio content.
[0337] "Search tools" refers to a mechanism for retrieving technical terms or explanations of new terms from online or offline databases.
[0338] A "voice synthesis means" is a technique or device for converting text data into voice data and providing voice feedback to the user.
[0339] The present invention provides a system including a voice input unit, a voice recognition unit, a control unit, a search unit, a voice synthesis unit, and an emotion recognition unit. This system allows users to easily operate audio content by voice, and not only quickly obtain explanations of technical terms and new terms, but also realizes more appropriate feedback and content suggestions based on the user's emotions. Specific embodiments are described in detail below.
[0340] System configuration
[0341] 1. Voice input method:
[0342] The device is equipped with a microphone that captures the user's speech as voice data.
[0343] Example: Use a commercially available microphone.
[0344] 2. Voice recognition means:
[0345] The captured voice data is sent to a server and converted into text data by a voice recognition means.
[0346] Examples of software used: General speech recognition engines (e.g., Google Speech-to-Text, Amazon Transcribe)
[0347] 3. Emotion recognition means:
[0348] The user's emotions are analyzed based on the same voice data.
[0349] Example of software used: Emotion recognition engine (e.g. IBM Watson Tone Analyzer)
[0350] 4. Control measures:
[0351] The server analyzes the converted text data and the emotion recognition results and generates instructions to perform appropriate operations.
[0352] Example: When the user says "play", start playing audio content.
[0353] 5. Search methods:
[0354] When an explanation of technical terms or new terms is required, the server uses search means to retrieve relevant information from the database.
[0355] Example: General search engines (e.g. Google Search API)
[0356] 6. Speech synthesis means:
[0357] The obtained information and processing results are converted from text data to audio data and provided to the user through the device's speaker.
[0358] Examples of software used: Speech synthesis engines (e.g., Amazon Polly, Google Text-to-Speech)
[0359] Specific examples
[0360] 1. Playback instructions:
[0361] When a user says "play," the device's microphone captures this voice. The captured voice data is sent to the server, where it is converted into text data ("play") using Google Speech-to-Text. The IBM Watson Tone Analyzer then analyzes the user's emotion, rating it as "excited." Based on this, the server instructs playback of audio content, and uses Amazon Polly to generate voice data such as "Start playback," which is played over the device's speaker.
[0362] 2. Terminology:
[0363] If a user says, "Please explain artificial intelligence," the audio is captured and converted to text using Google Speech-to-Text. Next, the IBM Watson Tone Analyzer evaluates the user's emotion as "interested." The server uses the Google Search API to obtain an explanation of "artificial intelligence" and provides it to the user using Amazon Polly.
[0364] Prompt Sentence Examples
[0365] "How can I start playback with a tone of voice that users find relaxing?"
[0366] "Explain how combining speech recognition and emotion recognition can respond to the user."
[0367] This system allows users to not only control audio content naturally through voice, but also receive personalized feedback based on their emotions. It also provides quick explanations of technical terms and new terms, greatly improving user convenience.
[0368] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0369] Step 1: Capture audio
[0370] Description: The device uses a microphone to capture what the user says.
[0371] Input: User's spoken utterance
[0372] Output: Audio data file
[0373] What it does: When the microphone detects the user speaking, audio data is recorded in real time and temporarily stored for later processing steps.
[0374] Step 2: Voice Recognition
[0375] Description: The device sends the captured voice data to the server, where it is converted into text data using speech recognition means.
[0376] Input: Audio data file
[0377] Output: Text data
[0378] Specific operation: The captured voice data is sent to the server and converted into text data using the Google Speech-to-Text API, a speech recognition engine. For example, if a user says "play," the speech recognition engine generates the text "play."
[0379] Step 3: Sentiment Analysis
[0380] Description: The server performs emotion recognition based on text and voice data.
[0381] Input: Audio data file, text data
[0382] Output: Emotion data
[0383] How it works: The voice data file is passed to an emotion recognition engine such as IBM Watson Tone Analyzer, which analyzes the user's voice for emotions (e.g., "excited" or "relaxed") and saves the results as emotion data.
[0384] Step 4: Command Parsing
[0385] Description: The server analyzes text and emotion data to determine the user's intent.
[0386] Input: Text data, emotion data
[0387] Output: Operation command
[0388] Specific operation: The server combines the converted text data with the analyzed emotional data to determine the user's intention. For example, if the user's text data is "play" and the emotional data is recognized as "excited," the server will determine that "the user wants to play audio content" and generate the corresponding operation command.
[0389] Step 5: Run a search
[0390] Description: When a term explanation is needed, the server uses a search tool to retrieve the relevant information.
[0391] Input: Text data of technical terms
[0392] Output: explanatory text data
[0393] Specific operation: The server uses the Google Search API to search online for the technical term the user is looking for (for example, "artificial intelligence") and obtains the results as explanatory text data.
[0394] Step 6: Audio adjustment and output
[0395] Description: The server generates and outputs voice based on the final text data and emotion data.
[0396] Input: explanatory text data (or text data related to operations), emotion data
[0397] Output: Audio data
[0398] How it works: The server uses Amazon Polly to convert text data into speech, adjusting the speech (e.g., tone and speech rate) based on emotion data, and then sending the resulting speech to the device and playing it back to the user through the speaker.
[0399] (Application example 2)
[0400] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0401] Conventional audio content manipulation systems operate based solely on the user's voice commands. This makes it difficult to provide feedback and suggest content that fully reflects the user's emotions and intentions. Furthermore, explanations of technical terms and new terms rely solely on voice commands, making it difficult to deepen understanding according to the situation. This makes it difficult for users to obtain an intuitive and personalized experience.
[0402] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a voice recognition means, a control means, and a search means. This enables intuitive and personalized operation and feedback of audio content based on the user's voice and emotions.
[0403] Understood. Below are the definitions of important words included in the rewritten patent claims.
[0404] "Audio input means" refers to a device or function that captures the user's voice.
[0405] "Speech recognition means" refers to a device or function for converting captured voice data into text.
[0406] "Controller" refers to a device or function that analyzes text data and emotional state to manipulate and adjust audio content.
[0407] "Search means" refers to a device or function for retrieving explanations of technical terms or new terms from a database.
[0408] "Speech synthesis means" refers to a device or function that converts acquired text information into voice data.
[0409] "Emotion recognition means" refers to a device or function that analyzes the user's emotional state from voice data.
[0410] The "playback position" refers to the current progress position during playback of audio content.
[0411] "Voice operation" refers to a method in which a user inputs operation commands by voice.
[0412] "Speech rate" refers to the speed at which audio content is played aloud.
[0413] "Tone" refers to the tone and intonation of a voice.
[0414] A system for implementing the present invention includes a voice input means, a voice recognition means, an emotion recognition means, a control means, a search means, and a voice synthesis means.
[0415] Program processing explanation
[0416] The server coordinates these means to manipulate and adjust audio content based on the user's voice input and emotional state.
[0417] 1. Voice Recognition Method
[0418] It uses the smartphone's microphone to capture the user's voice.
[0419] The captured voice data is converted into text data using a voice recognition service (Google Cloud Speech-to-Text API).
[0420] 2. Emotion recognition means
[0421] The voice data is sent to an emotion recognition service (IBM Watson Tone Analyzer) to analyze the user's emotional state, which is classified as "happiness," "sadness," "surprise," "anger," etc.
[0422] 3. Control Measures
[0423] The control means analyzes the text data and emotion recognition results to determine the user's intention. For example, if the user says, "Play the next song," the control means understands the intention and issues a command to start playing the next song.
[0424] Adjust reading speed and tone depending on emotional state.
[0425] 4. Search Methods
[0426] When an explanation of a technical term or a new term is needed, the search tool accesses a database (such as the Wikipedia API) to retrieve the relevant information, which is stored as text data.
[0427] 5. Speech synthesis means
[0428] The acquired text information is converted into voice data by a speech synthesis service (Amazon Polly) and provided to the user, with the tone and speed adjusted based on the user's emotional state.
[0429] Specific example explanation
[0430] Example 1: Relaxation scenario
[0431] User: "Play some relaxing music"
[0432] Server Action:
[0433] 1. The voice recognition means generates text data such as "Play some relaxing music."
[0434] 2. The emotion recognition means determines that the user's emotional state is one of seeking "relaxation."
[0435] 3. The control means instructs playback of a relaxation playlist.
[0436] 4. The speech synthesis means audibly conveys the relevant announcement to the user.
[0437] Prompt Sentence Examples
[0438] If the user says they want to relax, this prompt should suggest music to play.
[0439] [Relax, Music, Play]
[0440] Example 2: Scenario where terminology needs to be explained
[0441] User: "What is artificial intelligence?"
[0442] Server Action:
[0443] 1. The text data "What is artificial intelligence?" is generated by the speech recognition means.
[0444] 2. Emotion recognition means identifies the user's interests.
[0445] 3. The control means obtains the necessary information through the search means and analyzes its meaning.
[0446] 4. The search tool retrieves "artificial intelligence" information from the database.
[0447] 5. The speech synthesis means provides the acquired information to the user by speech.
[0448] Prompt Sentence Examples
[0449] When a user asks "What is artificial intelligence?", generate an explanation using the Wikipedia API.
[0450] The above is a specific embodiment for carrying out the invention. This system allows users to enjoy intuitive and personalized audio content manipulation and feedback through voice and emotion.
[0451] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0452] Step 1:
[0453] The user inputs voice data using the smartphone's microphone. For example, the user might say, "Play some relaxing music." This voice data becomes the input.
[0454] Step 2:
[0455] The device captures the voice input data and sends it to a speech recognition service (Google Cloud Speech-to-Text API). The speech recognition service converts the voice data into text data. The input is voice data, and the output is the converted text data.
[0456] Step 3:
[0457] The device sends the converted text data to an emotion recognition service (IBM Watson Tone Analyzer), which analyzes the user's emotional state. For example, it detects that the user wants to relax. The input is text data, and the output is emotion data.
[0458] Step 4:
[0459] The server sends the text data and emotional data to the control means for analysis. The control means determines the user's intention based on the text data. For example, based on the text "Play relaxing music" and the emotional data indicating a desire to relax, the control means selects an appropriate music playlist. The input is text data and emotional data, and the output is an operation command.
[0460] Step 5:
[0461] The server retrieves information from a database (such as Wikipedia API) using search tools to provide explanations of technical terms and new terms as needed. For example, this step is executed when a user asks "What is artificial intelligence?" The input is text data, and the output is the retrieved information.
[0462] Step 6:
[0463] The server passes the acquired information or the selected music playlist to a speech synthesis means (Amazon Polly) and converts it into voice data. For example, the title and summary of the playlist are announced to the user by voice. This voice data is the output. The input is text data, and the output is voice data.
[0464] Step 7:
[0465] The device provides the generated voice data to the user through a speaker. The output from the voice synthesis means is played back as is. For example, a voice notification such as "Relaxing music will be played" is given. The input is the voice data, and the output is the voice feedback provided to the user.
[0466] These processing steps enable users to have intuitive and personalized audio content manipulation and feedback based on their voice input and emotional state.
[0467] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0468] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0469] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0470] [Second embodiment]
[0471] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0472] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0473] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0474] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0475] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0476] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0477] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0478] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0479] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0480] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0481] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0482] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0483] Understood. Below is the description of the "Mode for Carrying Out the Invention" of the patent specification.
[0484] The present invention provides a system including a voice input unit, a voice recognition unit, a control unit, a search unit, and a voice synthesis unit. This system allows a user to easily operate audio content by voice and quickly obtain explanations of technical terms and new terms. Specific embodiments of the present invention are described in detail below.
[0485] System configuration:
[0486] 1. Voice input method:
[0487] The device is equipped with a microphone that captures the user's speech as voice data.
[0488] This voice input means provides a basis for responding to operational commands and questions uttered by the user.
[0489] 2. Voice recognition means:
[0490] The captured voice data is sent to a voice recognition means and converted into text data.
[0491] This speech recognizer uses advanced algorithms to extract accurate text from speech.
[0492] 3. Control measures:
[0493] The converted text data is analyzed by the control means and the corresponding operation is performed.
[0494] For example, if the user says "play," the control will begin playing audio content.
[0495] 4. Search methods:
[0496] When a user requests an explanation of a particular term, the control means uses the search means to quickly obtain the meaning of that term.
[0497] The search tool accesses an extensive database and retrieves detailed explanations of technical terms.
[0498] 5. Speech synthesis means:
[0499] The acquired text information is passed to a voice synthesis means and converted into voice data.
[0500] The speech synthesis means generates natural speech in real time and provides audible feedback to the user.
[0501] Explanation of program operation:
[0502] 1. Speech Recognition:
[0503] The device captures the user's speech and converts it into text using a speech recognition service.
[0504] The converted text data is passed to the next processing step.
[0505] 2. Command analysis:
[0506] The control means analyzes the converted text data and determines the user's intention.
[0507] Specific actions are determined in response to control commands (e.g., play, stop, rewind, etc.) and questions (e.g., "What does AI mean?").
[0508] 3. Perform a search:
[0509] When an explanation of a technical term is required, the control means calls the search means and retrieves the meaning of the relevant term from the database.
[0510] The information obtained is stored as text.
[0511] 4. Audio output:
[0512] The final information is passed to a speech synthesis means and converted into speech data.
[0513] The terminal plays the generated voice data to the user through a speaker.
[0514] Examples:
[0515] 1. Playback instructions:
[0516] When the user utters "play", the terminal captures this speech and converts it into text "play" through speech recognition means.
[0517] The control means recognizes the command "play" and starts playing the audio content.
[0518] 2. Terminology:
[0519] If a user says, "Please explain artificial intelligence," the device converts the speech into text and recognizes it as "Please explain artificial intelligence."
[0520] The control means determines that an explanation of "artificial intelligence" is desired, and uses the search means to obtain the meaning of "artificial intelligence."
[0521] The acquired commentary is converted into voice data by a voice synthesis means and provided to the user.
[0522] In this way, the system of the present invention can meet the diverse needs of users, improving operability and convenience.
[0523] The processing flow will be explained below.
[0524] Understood. Below, I will explain the process flow in concrete steps.
[0525] Step 1:
[0526] The user speaks into the device's microphone to give voice commands such as play, stop, rewind, skip, or explain technical terms.
[0527] Step 2:
[0528] The device captures the user's speech as voice data through a microphone.
[0529] Step 3:
[0530] The voice data captured by the device is sent to the voice recognition means and converted into text data using the Google voice recognition service.
[0531] Step 4:
[0532] The terminal receives the converted text data and passes it to the control means.
[0533] Step 5:
[0534] The control means of the terminal analyzes the text data to determine the user's intended command (eg, play, stop, rewind, skip, or commentary).
[0535] Step 6:
[0536] If the command is "play," "stop," "rewind," or "skip," the device's control invokes the corresponding audio player function (e.g., the play, pause, rewind, or skip method) to perform the audio content operation.
[0537] Step 7:
[0538] If the command requests an explanation of a technical term, the control means of the terminal utilizes the search means to obtain an explanation of the specified term from the database.
[0539] Step 8:
[0540] A search means on the server searches a database for an explanation of the requested technical term and returns it to the terminal in text format.
[0541] Step 9:
[0542] The explanatory text acquired by the terminal is passed to a speech synthesis means, which generates voice data using TTS (Text-To-Speech).
[0543] Step 10:
[0544] The device outputs the generated voice data to the user through a speaker, providing feedback in response to explanations and commands.
[0545] The above are the steps for processing voice instructions in the system of the present invention. This detailed processing procedure allows the user to intuitively operate audio content and receive explanations of technical terms.
[0546] Example 1
[0547] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0548] In conventional voice-operated systems, when users operate audio content by voice, feedback on operations and explanations of technical terms can be delayed, resulting in a lack of convenience. Additionally, there is a lack of a mechanism for rapid and accurate voice recognition and subsequent processing.
[0549] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0550] In this invention, the server includes a voice input means, a voice recognition means for converting voice data into text, a control means for analyzing the text data and operating the media content, a search means for providing explanations of technical terms and undefined terms, and a voice synthesis means for outputting the obtained explanations by voice, thereby enabling quick and accurate provision of information and operation of the media content by voice operation.
[0551] 1. "Voice input means" means any device or technology used to capture a user's speech as voice data.
[0552] 2. "Speech recognition means" means a technology or device for converting voice data captured by a voice input means into text data.
[0553] 3. "Control means" refers to the technology or device that analyzes the text data converted by the speech recognition means and controls the media content.
[0554] 4. "Search means" refers to the technology or device that allows users to obtain explanations of technical terms or undefined terms they are looking for.
[0555] 5. "Speech synthesis means" refers to technology or equipment for outputting the acquired commentary in audio form.
[0556] 6. "Media Content" refers to all content, including audio and video, that can be used through voice control.
[0557] 7. "Prompt" means a voice command that allows a user to give verbal instructions to a system.
[0558] The present invention provides a system including a voice input unit, a voice recognition unit, a control unit, a search unit, and a voice synthesis unit. This system allows users to easily operate media content by voice and quickly obtain explanations of technical terms and new terms. Specific embodiments of the present invention are described in detail below.
[0559] System configuration
[0560] 1. Voice input method:
[0561] The device is equipped with a microphone to capture the user's speech as voice data. This voice input means provides a basis for responding to user commands and questions.
[0562] 2. Voice recognition means:
[0563] The voice data captured by the device is sent to a server, which converts the voice data into text using a speech recognition service such as the Google Cloud Speech-to-Text API, which uses advanced algorithms to extract accurate text from the speech.
[0564] 3. Control measures:
[0565] The server analyzes the converted text data to determine the user's intent and determines the specific actions to take in response to commands (e.g., play, stop, rewind, etc.) or questions (e.g., "What does AI mean?").
[0566] 4. Search methods:
[0567] When a user requests an explanation of a particular term, the server uses a search engine to quickly retrieve the meaning of the term, accessing the Wikipedia API and other extensive databases to retrieve detailed explanations of technical terms.
[0568] 5. Speech synthesis means:
[0569] The text information acquired by the server is converted into voice data using a speech synthesis service such as Amazon Polly, which generates natural-sounding voices in real time and provides voice feedback to the user.
[0570] Specific examples
[0571] 1. Playback instructions:
[0572] When the user says "play," the terminal captures this speech and converts it into text through speech recognition means as "play." The server recognizes the command "play" and starts playing the media content.
[0573] 2. Terminology:
[0574] When a user says, "Please explain artificial intelligence," the device converts the speech into text and recognizes it as "Please explain artificial intelligence." The server determines that an explanation of "artificial intelligence" is being sought, and uses a search means to obtain the meaning of "artificial intelligence." The obtained explanation is converted into voice data by a voice synthesis means and provided to the user.
[0575] Prompt Sentence Examples
[0576] 1. Playback request prompt:
[0577] When someone says "play," use speech recognition to convert it to text and provide instructions for starting to play media content.
[0578] 2. Glossary prompt:
[0579] If someone says "Please explain artificial intelligence," please provide instructions on how to use the Wikipedia API to obtain an explanation of artificial intelligence, synthesize it into voice, and provide it to the user.
[0580] The system of the present invention is thus able to meet the diverse needs of users and greatly improves operability and convenience.
[0581] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0582] Step 1: Voice Input
[0583] Users speak voice commands to control media content and to explain technical terms.
[0584] The terminal uses a microphone to capture the user's speech as voice data.
[0585] Input: User speech
[0586] Output: Audio data
[0587] Step 2: Voice Recognition
[0588] The terminal transmits the captured audio data to the server.
[0589] The server converts the voice data into text data using the Google Cloud Speech-to-Text API.
[0590] Input: Audio data
[0591] Data processing: Converting voice data into text data
[0592] Output: Text data
[0593] Step 3: Command analysis
[0594] The server analyzes the text data obtained by the voice recognition means and determines whether the spoken content is an operation command or a question.
[0595] For example, text data such as "play" is determined to be an operation command, and text data such as "please explain artificial intelligence" is determined to be a question.
[0596] Input: Text data
[0597] Data operation: Content analysis and category determination of text data (operation commands or questions)
[0598] Output: The result of the operation command or question
[0599] Step 4: Execute the operation command
[0600] If the server recognizes the operation command, it performs the corresponding operation on the media content.
[0601] For example, the command "play" will call the play function of the media player.
[0602] Input: Operation command
[0603] Data operation: Executes media operations corresponding to commands.
[0604] Output: Playing, stopping, etc. media content
[0605] Step 5: Question Processing
[0606] When the server recognizes the question, it uses a search means to obtain explanations of technical terms and undefined terms.
[0607] For example, in response to the question "Please explain artificial intelligence," an explanation of "artificial intelligence" is retrieved from the Wikipedia API.
[0608] Input: Question text data
[0609] Data processing: Performing searches that correspond to the query
[0610] Output: Search results text data
[0611] Step 6: Text-to-Speech
[0612] The server converts the acquired text data of the explanation into audio data using Amazon Polly.
[0613] Input: Text data explaining the question
[0614] Data processing: Convert text data into audio data
[0615] Output: Audio data
[0616] Step 7: Audio Output
[0617] The terminal outputs the audio data transmitted from the server to the user through a speaker.
[0618] Input: Audio data
[0619] Output: Play audio
[0620] This completes the processing flow. The system of the present invention has the advantage that users can quickly and accurately use information and media content through voice operations.
[0621] (Application example 1)
[0622] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0623] Conventional content distribution services require users to physically operate the service and check search results, resulting in problems with usability and convenience. In particular, when requesting explanations of technical terms, users are often forced to search and check text information, which often interrupts the viewing experience. While systems that allow voice control exist, they lack the functionality to simultaneously operate viewing content and provide terminology explanations in real time. This degrades the user experience and makes it difficult to provide a comfortable viewing environment.
[0624] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0625] In this invention, the server includes a voice input means, a voice recognition means for converting voice data into text, a control means for analyzing the text data and operating the voice and audiovisual content, a search means for providing explanations of technical terms and new terms, a voice synthesis means for outputting the search results by voice, and a means for providing audio instructions for operating the audiovisual content and explanations of terms in real time. This allows the user to intuitively operate the audiovisual content by voice instructions while instantly obtaining audio explanations of technical terms, thereby creating an optimal usage environment.
[0626] "Voice input means" refers to a device or system, such as a microphone, for capturing a user's speech as voice data.
[0627] A "voice recognition means" is a technique or system for converting captured voice data into text data.
[0628] The "control means" is a system or method that analyzes text data and operates the audiovisual content based on the analysis.
[0629] A "search tool" is a system or method for quickly retrieving information about a particular term or terminology.
[0630] "Speech synthesis means" refers to a system or technology that converts text information into voice data and provides it to the user in voice form.
[0631] "Viewing content" refers to media content such as audio and video that a user views.
[0632] A "real-time audio feedback mechanism" is a system or method for providing immediate audio feedback to a user upon request.
[0633] System configuration
[0634] This invention provides a system that allows a user to operate viewing content through voice input and obtain audio explanations of technical terms in real time. The system includes the following main means.
[0635] 1. Voice input method:
[0636] - A microphone is used as a voice input means to capture the user's speech as voice data. A microphone built into a smartphone, smart glasses, a head-mounted display, or a robot is used.
[0637] 2. Voice recognition means:
[0638] - Use speech recognition software such as Google Cloud Speech-to-Text API to convert the captured voice data into text data, thereby capturing the user's voice instructions as text.
[0639] 3. Control measures:
[0640] - Analyze text data and manipulate viewing content. Use Python and natural language processing libraries (NLTK and spaCy) to analyze text data from speech recognition tools and perform operations such as play, stop, and rewind.
[0641] 4. Search methods:
[0642] - Use Elasticsearch or other database search technologies to provide explanations of technical terms and new terms. Search for explanations of terms captured as text and return them in text format.
[0643] 5. Speech synthesis means:
[0644] - Use speech synthesis technology such as Google Cloud Text-to-Speech API to convert text information obtained from search tools into audio data and provide feedback to users.
[0645] 6. Real-time feedback methods:
[0646] - Provide audio instructions and glossary of terms for the content in real time, ensuring seamless operation and commentary of the content.
[0647] Usage and Examples
[0648] In a real-world usage scenario, a user will interface with the system using a smartphone or other compatible device. The following example illustrates how to operate the system.
[0649] 1. Playback instructions:
[0650] - When the user says "play," the microphone captures the voice and converts it into text "play" using the Google Cloud Speech-to-Text API. The control means recognizes this and starts playing the content.
[0651] 2. Terminology:
[0652] - When a user says, "Please explain artificial intelligence," the speech is converted into text in the same way. This text is analyzed by the control means, and an explanation of "artificial intelligence" is retrieved from the Elasticsearch database through the search means. The retrieved text information is converted into audio data using the Google Cloud Text-to-Speech API and provided to the user in real time.
[0653] Example prompt sentence:
[0654] "Explain the following statement: 'Artificial intelligence refers to systems designed to have human intelligence.'"
[0655] These methods allow users to intuitively control content through voice commands and instantly receive explanations of technical terms, improving the user experience and ease of use.
[0656] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0657] Step 1:
[0658] The user speaks a voice command to the device (e.g., "play," "please explain with AI," etc.). The input is the user's voice, and the output is the audio data captured by the device's microphone.
[0659] Step 2:
[0660] The terminal captures the user's voice using the voice input means and obtains it as voice data. The input is the user's voice and the output is voice data.
[0661] Step 3:
[0662] The device sends voice data to the Google Cloud Speech-to-Text API, which converts the voice data into text. The input is voice data, and the output is text data. The server calls a Google Cloud service to convert the voice data.
[0663] Step 4:
[0664] The server uses a control means to analyze the received text data and determine the operation and explanation acquisition instructions based on that. The input is text data, and the output is the operation instructions or search query as the analysis result. The specific operation is to determine the user's intention using a text data analysis algorithm.
[0665] Step 5:
[0666] When the server operates the viewing content based on the analysis results, it executes operations such as play, stop, and rewind. The input is the operation command, and the output is the actual content operation. The specific operation is to call the media player's control API and execute the required operation.
[0667] Step 6:
[0668] When a term explanation is requested, the server uses a search mechanism to search for relevant information in a database. The input is a search query, and the output is text information as search results. The specific operation is to retrieve the appropriate information using Elasticsearch or other database search technologies.
[0669] Step 7:
[0670] The server sends the text information acquired through the search tool to the Google Cloud Text-to-Speech API and converts it into voice data. The input is text information and the output is voice data. The specific operation is to generate voice from text using the voice synthesis API.
[0671] Step 8:
[0672] The terminal plays back the voice data obtained from the voice synthesis means and provides real-time feedback to the user. The input is voice data, and the output is voice feedback to the user. The specific operation is to play back the voice data using the terminal's speaker.
[0673] Step 9:
[0674] The flow from step 1 to step 8 is repeated until the user utters a new command or performs the next operation. The input is the next voice command, and the output is the capture of the next voice data.
[0675] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0676] Understood. Below is a description of the "Mode for carrying out the invention" based on the invention that combines an emotion engine.
[0677] The present invention provides a system including a voice input unit, a voice recognition unit, a control unit, a search unit, a voice synthesis unit, and an emotion recognition unit. This system allows users to easily operate audio content by voice, quickly obtain explanations of technical terms and new terms, and also realizes more appropriate feedback and content suggestions based on the user's emotions. Specific embodiments of the present invention are described in detail below.
[0678] System configuration:
[0679] 1. Voice input method:
[0680] The device is equipped with a microphone that captures the user's speech as voice data.
[0681] The voice input means provides a basis for responding to operational commands and questions uttered by the user.
[0682] 2. Voice recognition means:
[0683] The captured voice data is sent to a voice recognition means and converted into text data.
[0684] The speech recognizer uses advanced algorithms to extract accurate text from speech.
[0685] 3. Emotion recognition means:
[0686] Analyze user emotions from captured voice data.
[0687] The emotion recognition means identifies the emotional state of the user based on information such as the tone, speed, and strength of the user's voice.
[0688] 4. Control measures:
[0689] The converted text data and the emotion data obtained by the emotion recognition means are analyzed, and a corresponding operation is performed.
[0690] For example, when the user says "play," the control means starts playing audio content and adjusts the speaking speed and tone based on the user's emotions.
[0691] 5. Search methods:
[0692] When a user requests an explanation of a particular term, the control means uses the search means to quickly obtain the meaning of that term.
[0693] The search tool accesses an extensive database and retrieves detailed explanations of technical terms.
[0694] 6. Speech synthesis means:
[0695] The acquired text information is passed to a voice synthesis means and converted into voice data.
[0696] The speech synthesis means generates natural speech in real time and provides audible feedback to the user.
[0697] Explanation of program operation:
[0698] 1. Speech Recognition:
[0699] The device captures the user's speech and converts it into text using a speech recognition service.
[0700] The converted text data is passed to the next processing step.
[0701] 2. Emotion analysis:
[0702] The emotion recognition means of the terminal analyzes the user's emotions from the captured voice data.
[0703] The results of the sentiment analysis are passed to the control means and used in the next processing step.
[0704] 3. Command Analysis:
[0705] The control means analyzes the converted text data and the emotion analysis data from the emotion recognition means to determine the user's intention.
[0706] Specific actions corresponding to operation commands (e.g., play, stop, rewind) or questions (e.g., "What does XX mean?") are determined.
[0707] 4. Perform a search:
[0708] When an explanation of a technical term is required, the control means calls the search means and retrieves the meaning of the relevant term from the database.
[0709] The information obtained is stored as text.
[0710] 5. Adjustment and Audio Output:
[0711] Based on the emotion recognition results, the reading speed and tone of the audio content are adjusted.
[0712] The final information is passed to a speech synthesis means and converted into speech data.
[0713] The device outputs the generated voice data to the user through a speaker, providing feedback in response to explanations and commands.
[0714] Examples:
[0715] 1. Playback instructions:
[0716] When the user utters "play", the terminal captures this speech and converts it into text "play" through speech recognition means.
[0717] The control means recognizes the command "play" and starts playing the audio content. Furthermore, the emotion recognition means analyzes the user's emotion, and if the user is nervous, for example, starts playing the audio content in a calm tone.
[0718] 2. Terminology:
[0719] If a user says, "Please explain artificial intelligence," the device converts the speech into text and recognizes it as "Please explain artificial intelligence."
[0720] The control means determines that an explanation of "artificial intelligence" is desired, and uses the search means to obtain the meaning of "artificial intelligence."
[0721] The acquired commentary is converted into voice data by a voice synthesis means and provided to the user. If the user shows interest, the emotion recognition means can detect this and continue with a more detailed commentary.
[0722] In this way, the system of the present invention can meet the diverse needs of users, not only improving operability and convenience but also providing a more personalized experience according to the user's emotions.
[0723] The processing flow will be explained below.
[0724] Understood. Below I will explain the process in concrete steps.
[0725] Step 1:
[0726] The user speaks into the device's microphone to give voice commands such as play, stop, rewind, skip, or explain technical terms.
[0727] Step 2:
[0728] The device captures the user's speech as voice data through a microphone.
[0729] Step 3:
[0730] The voice data captured by the terminal is sent to a voice recognition means, which converts the voice data into text data.
[0731] Step 4:
[0732] The terminal receives the converted text data and passes it to the control means.
[0733] Step 5:
[0734] An emotion recognition means of the terminal analyzes the user's emotions from the captured voice data.
[0735] Step 6:
[0736] The control means analyzes the text data and the emotion recognition results to determine the user's intention.
[0737] Step 7:
[0738] If the command is "play," "stop," "rewind," or "skip," the control invokes the corresponding audio player function (e.g., play, pause, rewind, skip method) to perform the operation on the audio content.
[0739] Step 8:
[0740] If the command requests an explanation of technical terms, the control means uses the search means to obtain an explanation of the specified term from the database.
[0741] Step 9:
[0742] A search means on the server searches a database for an explanation of the requested technical term and returns it to the terminal in text format.
[0743] Step 10:
[0744] The explanatory text acquired by the terminal is passed to a speech synthesis means, which generates voice data using TTS (Text-To-Speech).
[0745] Step 11:
[0746] Based on the emotion recognition results, the reading speed and tone of the audio content are adjusted.
[0747] Step 12:
[0748] The device outputs the generated voice data to the user through a speaker, providing feedback in response to explanations and commands.
[0749] Examples:
[0750] Playback instructions example
[0751] 1. Step 1: The user says "play."
[0752] 2. Step 2: The device captures audio data through the microphone.
[0753] 3. Step 3: The voice data is converted into text using a speech recognition tool.
[0754] 4. Step 4: The converted text "play" is sent to the control means.
[0755] 5. Step 5: The emotion recognition means analyzes the user's emotions and determines, for example, that the user is nervous.
[0756] 6. Step 6: The control means analyzes the "play" command and also takes into account the user's emotions.
[0757] 7. Step 7: The control means calls the play method of the audio player to begin playback.
[0758] 8. Step 11: Based on the emotion recognition results, start playing at a gentle speed and tone.
[0759] 9. Step 12: The terminal outputs the playback of the audio content to the user through the speaker.
[0760] When requesting an explanation of technical terms
[0761] 1. Step 1: The user says, "Please explain to me, AI."
[0762] 2. Step 2: The device captures audio data through the microphone.
[0763] 3. Step 3: The voice data is converted into text using a speech recognition tool.
[0764] 4. Step 4: The converted text "Please explain the AI" is sent to the control means.
[0765] 5. Step 5: The emotion recognition means analyzes the user's emotion and determines, for example, that the user is interested.
[0766] 6. Step 6: The control means analyzes the text "Please explain the AI" and determines the user's intention.
[0767] 7. Step 8: The control means instructs the search means to search for an explanation of "AI."
[0768] 8. Step 9: The server's search function retrieves explanations about "AI" from the database and returns them to the terminal in text format.
[0769] 9. Step 10: The terminal passes the explanatory text to the speech synthesis means to generate speech data.
[0770] 10. Step 11: Based on the emotion recognition results, a detailed commentary is generated in an appropriate tone of voice.
[0771] 11. Step 12: The terminal outputs the generated voice data to the user through the speaker to provide an explanation.
[0772] The above are the processing steps for voice instruction and emotion recognition in the system of the present invention. This detailed processing procedure allows users to operate audio content and receive technical terminology explanations in an intuitive and personalized manner.
[0773] Example 2
[0774] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0775] Conventional speech recognition systems have the ability to operate audio content and provide information based on user voice commands. However, they have the problem of being unable to provide feedback or suggest content based on the user's emotions, which limits the user experience. Another issue is that it is difficult to quickly provide explanations of technical terms and new terms, making it difficult to meet diverse user needs.
[0776] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes a voice input means, a voice recognition means for converting voice data into text data, an emotion recognition means for analyzing emotions from the voice data, a control means for analyzing the text data and the emotion data to operate the audio content, a search means for providing explanations of technical terms and new terms, and a voice synthesis means for converting search results into voice data and outputting the voice. This enables personalized feedback and content suggestions based on the user's emotions, and also makes it possible to quickly provide explanations of technical terms and new terms.
[0777] "Voice input means" refers to a device for capturing a user's speech, and primarily refers to a microphone or similar hardware.
[0778] A "voice recognition means" is a technique or device for converting captured voice data into text data, utilizing a voice recognition algorithm.
[0779] "Emotion recognition means" refers to a means for analyzing and identifying a user's emotions from voice data, and is a technology for determining an emotional state based on information such as voice tone, speed, and strength.
[0780] "Control means" refers to a device or system for analyzing the converted text data and emotion data and for performing manipulations on the audio content.
[0781] "Search tools" refers to a mechanism for retrieving technical terms or explanations of new terms from online or offline databases.
[0782] A "voice synthesis means" is a technique or device for converting text data into voice data and providing voice feedback to the user.
[0783] The present invention provides a system including a voice input unit, a voice recognition unit, a control unit, a search unit, a voice synthesis unit, and an emotion recognition unit. This system allows users to easily operate audio content by voice, and not only quickly obtain explanations of technical terms and new terms, but also realizes more appropriate feedback and content suggestions based on the user's emotions. Specific embodiments are described in detail below.
[0784] System configuration
[0785] 1. Voice input method:
[0786] The device is equipped with a microphone that captures the user's speech as voice data.
[0787] Example: Use a commercially available microphone.
[0788] 2. Voice recognition means:
[0789] The captured voice data is sent to a server and converted into text data by a voice recognition means.
[0790] Examples of software used: General speech recognition engines (e.g., Google Speech-to-Text, Amazon Transcribe)
[0791] 3. Emotion recognition means:
[0792] The user's emotions are analyzed based on the same voice data.
[0793] Example of software used: Emotion recognition engine (e.g. IBM Watson Tone Analyzer)
[0794] 4. Control measures:
[0795] The server analyzes the converted text data and the emotion recognition results and generates instructions to perform appropriate operations.
[0796] Example: When the user says "play", start playing audio content.
[0797] 5. Search methods:
[0798] When an explanation of technical terms or new terms is required, the server uses search means to retrieve relevant information from the database.
[0799] Example: General search engines (e.g. Google Search API)
[0800] 6. Speech synthesis means:
[0801] The obtained information and processing results are converted from text data to audio data and provided to the user through the device's speaker.
[0802] Examples of software used: Speech synthesis engines (e.g., Amazon Polly, Google Text-to-Speech)
[0803] Specific examples
[0804] 1. Playback instructions:
[0805] When a user says "play," the device's microphone captures this voice. The captured voice data is sent to the server, where it is converted into text data ("play") using Google Speech-to-Text. The IBM Watson Tone Analyzer then analyzes the user's emotion, rating it as "excited." Based on this, the server instructs playback of audio content, and uses Amazon Polly to generate voice data such as "Start playback," which is played over the device's speaker.
[0806] 2. Terminology:
[0807] If a user says, "Please explain artificial intelligence," the audio is captured and converted to text using Google Speech-to-Text. Next, the IBM Watson Tone Analyzer evaluates the user's emotion as "interested." The server uses the Google Search API to obtain an explanation of "artificial intelligence" and provides it to the user using Amazon Polly.
[0808] Prompt Sentence Examples
[0809] "How can I start playback with a tone of voice that users find relaxing?"
[0810] "Explain how combining speech recognition and emotion recognition can respond to the user."
[0811] This system allows users to not only control audio content naturally through voice, but also receive personalized feedback based on their emotions. It also provides quick explanations of technical terms and new terms, greatly improving user convenience.
[0812] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0813] Step 1: Capture audio
[0814] Description: The device uses a microphone to capture what the user says.
[0815] Input: User's spoken utterance
[0816] Output: Audio data file
[0817] What it does: When the microphone detects the user speaking, audio data is recorded in real time and temporarily stored for later processing steps.
[0818] Step 2: Voice Recognition
[0819] Description: The device sends the captured voice data to the server, where it is converted into text data using speech recognition means.
[0820] Input: Audio data file
[0821] Output: Text data
[0822] Specific operation: The captured voice data is sent to the server and converted into text data using the Google Speech-to-Text API, a speech recognition engine. For example, if a user says "play," the speech recognition engine generates the text "play."
[0823] Step 3: Sentiment Analysis
[0824] Description: The server performs emotion recognition based on text and voice data.
[0825] Input: Audio data file, text data
[0826] Output: Emotion data
[0827] How it works: The voice data file is passed to an emotion recognition engine such as IBM Watson Tone Analyzer, which analyzes the user's voice for emotions (e.g., "excited" or "relaxed") and saves the results as emotion data.
[0828] Step 4: Command Parsing
[0829] Description: The server analyzes text and emotion data to determine the user's intent.
[0830] Input: Text data, emotion data
[0831] Output: Operation command
[0832] Specific operation: The server combines the converted text data with the analyzed emotional data to determine the user's intention. For example, if the user's text data is "play" and the emotional data is recognized as "excited," the server will determine that "the user wants to play audio content" and generate the corresponding operation command.
[0833] Step 5: Run a search
[0834] Description: When a term explanation is needed, the server uses a search tool to retrieve the relevant information.
[0835] Input: Text data of technical terms
[0836] Output: explanatory text data
[0837] Specific operation: The server uses the Google Search API to search online for the technical term the user is looking for (for example, "artificial intelligence") and obtains the results as explanatory text data.
[0838] Step 6: Audio adjustment and output
[0839] Description: The server generates and outputs voice based on the final text data and emotion data.
[0840] Input: explanatory text data (or text data related to operations), emotion data
[0841] Output: Audio data
[0842] How it works: The server uses Amazon Polly to convert text data into speech, adjusting the speech (e.g., tone and speech rate) based on emotion data, and then sending the resulting speech to the device and playing it back to the user through the speaker.
[0843] (Application example 2)
[0844] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0845] Conventional audio content manipulation systems operate based solely on the user's voice commands. This makes it difficult to provide feedback and suggest content that fully reflects the user's emotions and intentions. Furthermore, explanations of technical terms and new terms rely solely on voice commands, making it difficult to deepen understanding according to the situation. This makes it difficult for users to obtain an intuitive and personalized experience.
[0846] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a voice recognition means, a control means, and a search means. This enables intuitive and personalized operation and feedback of audio content based on the user's voice and emotions.
[0847] Understood. Below are the definitions of important words included in the rewritten patent claims.
[0848] "Audio input means" refers to a device or function that captures the user's voice.
[0849] "Speech recognition means" refers to a device or function for converting captured voice data into text.
[0850] "Controller" refers to a device or function that analyzes text data and emotional state to manipulate and adjust audio content.
[0851] "Search means" refers to a device or function for retrieving explanations of technical terms or new terms from a database.
[0852] "Speech synthesis means" refers to a device or function that converts acquired text information into voice data.
[0853] "Emotion recognition means" refers to a device or function that analyzes the user's emotional state from voice data.
[0854] The "playback position" refers to the current progress position during playback of audio content.
[0855] "Voice operation" refers to a method in which a user inputs operation commands by voice.
[0856] "Speech rate" refers to the speed at which audio content is played aloud.
[0857] "Tone" refers to the tone and intonation of a voice.
[0858] A system for implementing the present invention includes a voice input means, a voice recognition means, an emotion recognition means, a control means, a search means, and a voice synthesis means.
[0859] Program processing explanation
[0860] The server coordinates these means to manipulate and adjust audio content based on the user's voice input and emotional state.
[0861] 1. Voice Recognition Method
[0862] It uses the smartphone's microphone to capture the user's voice.
[0863] The captured voice data is converted into text data using a voice recognition service (Google Cloud Speech-to-Text API).
[0864] 2. Emotion recognition means
[0865] The voice data is sent to an emotion recognition service (IBM Watson Tone Analyzer) to analyze the user's emotional state, which is classified as "happiness," "sadness," "surprise," "anger," etc.
[0866] 3. Control Measures
[0867] The control means analyzes the text data and emotion recognition results to determine the user's intention. For example, if the user says, "Play the next song," the control means understands the intention and issues a command to start playing the next song.
[0868] Adjust reading speed and tone depending on emotional state.
[0869] 4. Search Methods
[0870] When an explanation of a technical term or a new term is needed, the search tool accesses a database (such as the Wikipedia API) to retrieve the relevant information, which is stored as text data.
[0871] 5. Speech synthesis means
[0872] The acquired text information is converted into voice data by a speech synthesis service (Amazon Polly) and provided to the user, with the tone and speed adjusted based on the user's emotional state.
[0873] Specific example explanation
[0874] Example 1: Relaxation scenario
[0875] User: "Play some relaxing music"
[0876] Server Action:
[0877] 1. The voice recognition means generates text data such as "Play some relaxing music."
[0878] 2. The emotion recognition means determines that the user's emotional state is one of seeking "relaxation."
[0879] 3. The control means instructs playback of a relaxation playlist.
[0880] 4. The speech synthesis means audibly conveys the relevant announcement to the user.
[0881] Prompt Sentence Examples
[0882] If the user says they want to relax, this prompt should suggest music to play.
[0883] [Relax, Music, Play]
[0884] Example 2: Scenario where terminology needs to be explained
[0885] User: "What is artificial intelligence?"
[0886] Server Action:
[0887] 1. The text data "What is artificial intelligence?" is generated by the speech recognition means.
[0888] 2. Emotion recognition means identifies the user's interests.
[0889] 3. The control means obtains the necessary information through the search means and analyzes its meaning.
[0890] 4. The search tool retrieves "artificial intelligence" information from the database.
[0891] 5. The speech synthesis means provides the acquired information to the user by speech.
[0892] Prompt Sentence Examples
[0893] When a user asks "What is artificial intelligence?", generate an explanation using the Wikipedia API.
[0894] The above is a specific embodiment for carrying out the invention. This system allows users to enjoy intuitive and personalized audio content manipulation and feedback through voice and emotion.
[0895] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0896] Step 1:
[0897] The user inputs voice data using the smartphone's microphone. For example, the user might say, "Play some relaxing music." This voice data becomes the input.
[0898] Step 2:
[0899] The device captures the voice input data and sends it to a speech recognition service (Google Cloud Speech-to-Text API). The speech recognition service converts the voice data into text data. The input is voice data, and the output is the converted text data.
[0900] Step 3:
[0901] The device sends the converted text data to an emotion recognition service (IBM Watson Tone Analyzer), which analyzes the user's emotional state. For example, it detects that the user wants to relax. The input is text data, and the output is emotion data.
[0902] Step 4:
[0903] The server sends the text data and emotional data to the control means for analysis. The control means determines the user's intention based on the text data. For example, based on the text "Play relaxing music" and the emotional data indicating a desire to relax, the control means selects an appropriate music playlist. The input is text data and emotional data, and the output is an operation command.
[0904] Step 5:
[0905] The server retrieves information from a database (such as Wikipedia API) using search tools to provide explanations of technical terms and new terms as needed. For example, this step is executed when a user asks "What is artificial intelligence?" The input is text data, and the output is the retrieved information.
[0906] Step 6:
[0907] The server passes the acquired information or the selected music playlist to a speech synthesis means (Amazon Polly) and converts it into voice data. For example, the title and summary of the playlist are announced to the user by voice. This voice data is the output. The input is text data, and the output is voice data.
[0908] Step 7:
[0909] The device provides the generated voice data to the user through a speaker. The output from the voice synthesis means is played back as is. For example, a voice notification such as "Relaxing music will be played" is given. The input is the voice data, and the output is the voice feedback provided to the user.
[0910] These processing steps enable users to have intuitive and personalized audio content manipulation and feedback based on their voice input and emotional state.
[0911] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0912] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0913] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0914] [Third embodiment]
[0915] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0916] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[0917] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0918] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0919] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0920] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0921] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0922] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0923] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0924] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0925] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0926] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0927] Understood. Below is the description of the "Mode for Carrying Out the Invention" of the patent specification.
[0928] The present invention provides a system including a voice input unit, a voice recognition unit, a control unit, a search unit, and a voice synthesis unit. This system allows a user to easily operate audio content by voice and quickly obtain explanations of technical terms and new terms. Specific embodiments of the present invention are described in detail below.
[0929] System configuration:
[0930] 1. Voice input method:
[0931] The device is equipped with a microphone that captures the user's speech as voice data.
[0932] This voice input means provides a basis for responding to operational commands and questions uttered by the user.
[0933] 2. Voice recognition means:
[0934] The captured voice data is sent to a voice recognition means and converted into text data.
[0935] This speech recognizer uses advanced algorithms to extract accurate text from speech.
[0936] 3. Control measures:
[0937] The converted text data is analyzed by the control means and the corresponding operation is performed.
[0938] For example, if the user says "play," the control will begin playing audio content.
[0939] 4. Search methods:
[0940] When a user requests an explanation of a particular term, the control means uses the search means to quickly obtain the meaning of that term.
[0941] The search tool accesses an extensive database and retrieves detailed explanations of technical terms.
[0942] 5. Speech synthesis means:
[0943] The acquired text information is passed to a voice synthesis means and converted into voice data.
[0944] The speech synthesis means generates natural speech in real time and provides audible feedback to the user.
[0945] Explanation of program operation:
[0946] 1. Speech Recognition:
[0947] The device captures the user's speech and converts it into text using a speech recognition service.
[0948] The converted text data is passed to the next processing step.
[0949] 2. Command analysis:
[0950] The control means analyzes the converted text data and determines the user's intention.
[0951] Specific actions are determined in response to control commands (e.g., play, stop, rewind, etc.) and questions (e.g., "What does AI mean?").
[0952] 3. Perform a search:
[0953] When an explanation of a technical term is required, the control means calls the search means and retrieves the meaning of the relevant term from the database.
[0954] The information obtained is stored as text.
[0955] 4. Audio output:
[0956] The final information is passed to a speech synthesis means and converted into speech data.
[0957] The terminal plays the generated voice data to the user through a speaker.
[0958] Examples:
[0959] 1. Playback instructions:
[0960] When the user utters "play", the terminal captures this speech and converts it into text "play" through speech recognition means.
[0961] The control means recognizes the command "play" and starts playing the audio content.
[0962] 2. Terminology:
[0963] If a user says, "Please explain artificial intelligence," the device converts the speech into text and recognizes it as "Please explain artificial intelligence."
[0964] The control means determines that an explanation of "artificial intelligence" is desired, and uses the search means to obtain the meaning of "artificial intelligence."
[0965] The acquired commentary is converted into voice data by a voice synthesis means and provided to the user.
[0966] In this way, the system of the present invention can meet the diverse needs of users, improving operability and convenience.
[0967] The processing flow will be explained below.
[0968] Understood. Below, I will explain the process flow in concrete steps.
[0969] Step 1:
[0970] The user speaks into the device's microphone to give voice commands such as play, stop, rewind, skip, or explain technical terms.
[0971] Step 2:
[0972] The device captures the user's speech as voice data through a microphone.
[0973] Step 3:
[0974] The voice data captured by the device is sent to the voice recognition means and converted into text data using the Google voice recognition service.
[0975] Step 4:
[0976] The terminal receives the converted text data and passes it to the control means.
[0977] Step 5:
[0978] The control means of the terminal analyzes the text data to determine the user's intended command (eg, play, stop, rewind, skip, or commentary).
[0979] Step 6:
[0980] If the command is "play," "stop," "rewind," or "skip," the device's control invokes the corresponding audio player function (e.g., the play, pause, rewind, or skip method) to perform the audio content operation.
[0981] Step 7:
[0982] If the command requests an explanation of a technical term, the control means of the terminal utilizes the search means to obtain an explanation of the specified term from the database.
[0983] Step 8:
[0984] A search means on the server searches a database for an explanation of the requested technical term and returns it to the terminal in text format.
[0985] Step 9:
[0986] The explanatory text acquired by the terminal is passed to a speech synthesis means, which generates voice data using TTS (Text-To-Speech).
[0987] Step 10:
[0988] The device outputs the generated voice data to the user through a speaker, providing feedback in response to explanations and commands.
[0989] The above are the steps for processing voice instructions in the system of the present invention. This detailed processing procedure allows the user to intuitively operate audio content and receive explanations of technical terms.
[0990] Example 1
[0991] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0992] In conventional voice-operated systems, when users operate audio content by voice, feedback on operations and explanations of technical terms can be delayed, resulting in a lack of convenience. Additionally, there is a lack of a mechanism for rapid and accurate voice recognition and subsequent processing.
[0993] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0994] In this invention, the server includes a voice input means, a voice recognition means for converting voice data into text, a control means for analyzing the text data and operating the media content, a search means for providing explanations of technical terms and undefined terms, and a voice synthesis means for outputting the obtained explanations by voice, thereby enabling quick and accurate provision of information and operation of the media content by voice operation.
[0995] 1. "Voice input means" means any device or technology used to capture a user's speech as voice data.
[0996] 2. "Speech recognition means" means a technology or device for converting voice data captured by a voice input means into text data.
[0997] 3. "Control means" refers to the technology or device that analyzes the text data converted by the speech recognition means and controls the media content.
[0998] 4. "Search means" refers to the technology or device that allows users to obtain explanations of technical terms or undefined terms they are looking for.
[0999] 5. "Speech synthesis means" refers to technology or equipment for outputting the acquired commentary in audio form.
[1000] 6. "Media Content" refers to all content, including audio and video, that can be used through voice control.
[1001] 7. "Prompt" means a voice command that allows a user to give verbal instructions to a system.
[1002] The present invention provides a system including a voice input unit, a voice recognition unit, a control unit, a search unit, and a voice synthesis unit. This system allows users to easily operate media content by voice and quickly obtain explanations of technical terms and new terms. Specific embodiments of the present invention are described in detail below.
[1003] System configuration
[1004] 1. Voice input method:
[1005] The device is equipped with a microphone to capture the user's speech as voice data. This voice input means provides a basis for responding to user commands and questions.
[1006] 2. Voice recognition means:
[1007] The voice data captured by the device is sent to a server, which converts the voice data into text using a speech recognition service such as the Google Cloud Speech-to-Text API, which uses advanced algorithms to extract accurate text from the speech.
[1008] 3. Control measures:
[1009] The server analyzes the converted text data to determine the user's intent and determines the specific actions to take in response to commands (e.g., play, stop, rewind, etc.) or questions (e.g., "What does AI mean?").
[1010] 4. Search methods:
[1011] When a user requests an explanation of a particular term, the server uses a search engine to quickly retrieve the meaning of the term, accessing the Wikipedia API and other extensive databases to retrieve detailed explanations of technical terms.
[1012] 5. Speech synthesis means:
[1013] The text information acquired by the server is converted into voice data using a speech synthesis service such as Amazon Polly, which generates natural-sounding voices in real time and provides voice feedback to the user.
[1014] Specific examples
[1015] 1. Playback instructions:
[1016] When the user says "play," the terminal captures this speech and converts it into text through speech recognition means as "play." The server recognizes the command "play" and starts playing the media content.
[1017] 2. Terminology:
[1018] When a user says, "Please explain artificial intelligence," the device converts the speech into text and recognizes it as "Please explain artificial intelligence." The server determines that an explanation of "artificial intelligence" is being sought, and uses a search means to obtain the meaning of "artificial intelligence." The obtained explanation is converted into voice data by a voice synthesis means and provided to the user.
[1019] Prompt Sentence Examples
[1020] 1. Playback request prompt:
[1021] When someone says "play," use speech recognition to convert it to text and provide instructions for starting to play media content.
[1022] 2. Glossary prompt:
[1023] If someone says "Please explain artificial intelligence," please provide instructions on how to use the Wikipedia API to obtain an explanation of artificial intelligence, synthesize it into voice, and provide it to the user.
[1024] The system of the present invention is thus able to meet the diverse needs of users and greatly improves operability and convenience.
[1025] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1026] Step 1: Voice Input
[1027] Users speak voice commands to control media content and to explain technical terms.
[1028] The terminal uses a microphone to capture the user's speech as voice data.
[1029] Input: User speech
[1030] Output: Audio data
[1031] Step 2: Voice Recognition
[1032] The terminal transmits the captured audio data to the server.
[1033] The server converts the voice data into text data using the Google Cloud Speech-to-Text API.
[1034] Input: Audio data
[1035] Data processing: Converting voice data into text data
[1036] Output: Text data
[1037] Step 3: Command analysis
[1038] The server analyzes the text data obtained by the voice recognition means and determines whether the spoken content is an operation command or a question.
[1039] For example, text data such as "play" is determined to be an operation command, and text data such as "please explain artificial intelligence" is determined to be a question.
[1040] Input: Text data
[1041] Data operation: Content analysis and category determination of text data (operation commands or questions)
[1042] Output: The result of the operation command or question
[1043] Step 4: Execute the operation command
[1044] If the server recognizes the operation command, it performs the corresponding operation on the media content.
[1045] For example, the command "play" will call the play function of the media player.
[1046] Input: Operation command
[1047] Data operation: Executes media operations corresponding to commands.
[1048] Output: Playing, stopping, etc. media content
[1049] Step 5: Question Processing
[1050] When the server recognizes the question, it uses a search means to obtain explanations of technical terms and undefined terms.
[1051] For example, in response to the question "Please explain artificial intelligence," an explanation of "artificial intelligence" is retrieved from the Wikipedia API.
[1052] Input: Question text data
[1053] Data processing: Performing searches that correspond to the query
[1054] Output: Search results text data
[1055] Step 6: Text-to-Speech
[1056] The server converts the acquired text data of the explanation into audio data using Amazon Polly.
[1057] Input: Text data explaining the question
[1058] Data processing: Convert text data into audio data
[1059] Output: Audio data
[1060] Step 7: Audio Output
[1061] The terminal outputs the audio data transmitted from the server to the user through a speaker.
[1062] Input: Audio data
[1063] Output: Play audio
[1064] This completes the processing flow. The system of the present invention has the advantage that users can quickly and accurately use information and media content through voice operations.
[1065] (Application example 1)
[1066] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1067] Conventional content distribution services require users to physically operate the service and check search results, resulting in problems with usability and convenience. In particular, when requesting explanations of technical terms, users are often forced to search and check text information, which often interrupts the viewing experience. While systems that allow voice control exist, they lack the functionality to simultaneously operate viewing content and provide terminology explanations in real time. This degrades the user experience and makes it difficult to provide a comfortable viewing environment.
[1068] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1069] In this invention, the server includes a voice input means, a voice recognition means for converting voice data into text, a control means for analyzing the text data and operating the voice and audiovisual content, a search means for providing explanations of technical terms and new terms, a voice synthesis means for outputting the search results by voice, and a means for providing audio instructions for operating the audiovisual content and explanations of terms in real time. This allows the user to intuitively operate the audiovisual content by voice instructions while instantly obtaining audio explanations of technical terms, thereby creating an optimal usage environment.
[1070] "Voice input means" refers to a device or system, such as a microphone, for capturing a user's speech as voice data.
[1071] A "voice recognition means" is a technique or system for converting captured voice data into text data.
[1072] The "control means" is a system or method that analyzes text data and operates the audiovisual content based on the analysis.
[1073] A "search tool" is a system or method for quickly retrieving information about a particular term or terminology.
[1074] "Speech synthesis means" refers to a system or technology that converts text information into voice data and provides it to the user in voice form.
[1075] "Viewing content" refers to media content such as audio and video that a user views.
[1076] A "real-time audio feedback mechanism" is a system or method for providing immediate audio feedback to a user upon request.
[1077] System configuration
[1078] This invention provides a system that allows a user to operate viewing content through voice input and obtain audio explanations of technical terms in real time. The system includes the following main means.
[1079] 1. Voice input method:
[1080] - A microphone is used as a voice input means to capture the user's speech as voice data. A microphone built into a smartphone, smart glasses, a head-mounted display, or a robot is used.
[1081] 2. Voice recognition means:
[1082] - Use speech recognition software such as Google Cloud Speech-to-Text API to convert the captured voice data into text data, thereby capturing the user's voice instructions as text.
[1083] 3. Control measures:
[1084] - Analyze text data and manipulate viewing content. Use Python and natural language processing libraries (NLTK and spaCy) to analyze text data from speech recognition tools and perform operations such as play, stop, and rewind.
[1085] 4. Search methods:
[1086] - Use Elasticsearch or other database search technologies to provide explanations of technical terms and new terms. Search for explanations of terms captured as text and return them in text format.
[1087] 5. Speech synthesis means:
[1088] - Use speech synthesis technology such as Google Cloud Text-to-Speech API to convert text information obtained from search tools into audio data and provide feedback to users.
[1089] 6. Real-time feedback methods:
[1090] - Provide audio instructions and glossary of terms for the content in real time, ensuring seamless operation and commentary of the content.
[1091] Usage and Examples
[1092] In a real-world usage scenario, a user will interface with the system using a smartphone or other compatible device. The following example illustrates how to operate the system.
[1093] 1. Playback instructions:
[1094] - When the user says "play," the microphone captures the voice and converts it into text "play" using the Google Cloud Speech-to-Text API. The control means recognizes this and starts playing the content.
[1095] 2. Terminology:
[1096] - When a user says, "Please explain artificial intelligence," the speech is converted into text in the same way. This text is analyzed by the control means, and an explanation of "artificial intelligence" is retrieved from the Elasticsearch database through the search means. The retrieved text information is converted into audio data using the Google Cloud Text-to-Speech API and provided to the user in real time.
[1097] Example prompt sentence:
[1098] "Explain the following statement: 'Artificial intelligence refers to systems designed to have human intelligence.'"
[1099] These methods allow users to intuitively control content through voice commands and instantly receive explanations of technical terms, improving the user experience and ease of use.
[1100] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1101] Step 1:
[1102] The user speaks a voice command to the device (e.g., "play," "please explain with AI," etc.). The input is the user's voice, and the output is the audio data captured by the device's microphone.
[1103] Step 2:
[1104] The terminal captures the user's voice using the voice input means and obtains it as voice data. The input is the user's voice and the output is voice data.
[1105] Step 3:
[1106] The device sends voice data to the Google Cloud Speech-to-Text API, which converts the voice data into text. The input is voice data, and the output is text data. The server calls a Google Cloud service to convert the voice data.
[1107] Step 4:
[1108] The server uses a control means to analyze the received text data and determine the operation and explanation acquisition instructions based on that. The input is text data, and the output is the operation instructions or search query as the analysis result. The specific operation is to determine the user's intention using a text data analysis algorithm.
[1109] Step 5:
[1110] When the server operates the viewing content based on the analysis results, it executes operations such as play, stop, and rewind. The input is the operation command, and the output is the actual content operation. The specific operation is to call the media player's control API and execute the required operation.
[1111] Step 6:
[1112] When a term explanation is requested, the server uses a search mechanism to search for relevant information in a database. The input is a search query, and the output is text information as search results. The specific operation is to retrieve the appropriate information using Elasticsearch or other database search technologies.
[1113] Step 7:
[1114] The server sends the text information acquired through the search tool to the Google Cloud Text-to-Speech API and converts it into voice data. The input is text information and the output is voice data. The specific operation is to generate voice from text using the voice synthesis API.
[1115] Step 8:
[1116] The terminal plays back the voice data obtained from the voice synthesis means and provides real-time feedback to the user. The input is voice data, and the output is voice feedback to the user. The specific operation is to play back the voice data using the terminal's speaker.
[1117] Step 9:
[1118] The flow from step 1 to step 8 is repeated until the user utters a new command or performs the next operation. The input is the next voice command, and the output is the capture of the next voice data.
[1119] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1120] Understood. Below is a description of the "Mode for carrying out the invention" based on the invention that combines an emotion engine.
[1121] The present invention provides a system including a voice input unit, a voice recognition unit, a control unit, a search unit, a voice synthesis unit, and an emotion recognition unit. This system allows users to easily operate audio content by voice, quickly obtain explanations of technical terms and new terms, and also realizes more appropriate feedback and content suggestions based on the user's emotions. Specific embodiments of the present invention are described in detail below.
[1122] System configuration:
[1123] 1. Voice input method:
[1124] The device is equipped with a microphone that captures the user's speech as voice data.
[1125] The voice input means provides a basis for responding to operational commands and questions uttered by the user.
[1126] 2. Voice recognition means:
[1127] The captured voice data is sent to a voice recognition means and converted into text data.
[1128] The speech recognizer uses advanced algorithms to extract accurate text from speech.
[1129] 3. Emotion recognition means:
[1130] Analyze user emotions from captured voice data.
[1131] The emotion recognition means identifies the emotional state of the user based on information such as the tone, speed, and strength of the user's voice.
[1132] 4. Control measures:
[1133] The converted text data and the emotion data obtained by the emotion recognition means are analyzed, and a corresponding operation is performed.
[1134] For example, when the user says "play," the control means starts playing audio content and adjusts the speaking speed and tone based on the user's emotions.
[1135] 5. Search methods:
[1136] When a user requests an explanation of a particular term, the control means uses the search means to quickly obtain the meaning of that term.
[1137] The search tool accesses an extensive database and retrieves detailed explanations of technical terms.
[1138] 6. Speech synthesis means:
[1139] The acquired text information is passed to a voice synthesis means and converted into voice data.
[1140] The speech synthesis means generates natural speech in real time and provides audible feedback to the user.
[1141] Explanation of program operation:
[1142] 1. Speech Recognition:
[1143] The device captures the user's speech and converts it into text using a speech recognition service.
[1144] The converted text data is passed to the next processing step.
[1145] 2. Emotion analysis:
[1146] The emotion recognition means of the terminal analyzes the user's emotions from the captured voice data.
[1147] The results of the sentiment analysis are passed to the control means and used in the next processing step.
[1148] 3. Command Analysis:
[1149] The control means analyzes the converted text data and the emotion analysis data from the emotion recognition means to determine the user's intention.
[1150] Specific actions corresponding to operation commands (e.g., play, stop, rewind) or questions (e.g., "What does XX mean?") are determined.
[1151] 4. Perform a search:
[1152] When an explanation of a technical term is required, the control means calls the search means and retrieves the meaning of the relevant term from the database.
[1153] The information obtained is stored as text.
[1154] 5. Adjustment and Audio Output:
[1155] Based on the emotion recognition results, the reading speed and tone of the audio content are adjusted.
[1156] The final information is passed to a speech synthesis means and converted into speech data.
[1157] The device outputs the generated voice data to the user through a speaker, providing feedback in response to explanations and commands.
[1158] Examples:
[1159] 1. Playback instructions:
[1160] When the user utters "play", the terminal captures this speech and converts it into text "play" through speech recognition means.
[1161] The control means recognizes the command "play" and starts playing the audio content. Furthermore, the emotion recognition means analyzes the user's emotion, and if the user is nervous, for example, starts playing the audio content in a calm tone.
[1162] 2. Terminology:
[1163] If a user says, "Please explain artificial intelligence," the device converts the speech into text and recognizes it as "Please explain artificial intelligence."
[1164] The control means determines that an explanation of "artificial intelligence" is desired, and uses the search means to obtain the meaning of "artificial intelligence."
[1165] The acquired commentary is converted into voice data by a voice synthesis means and provided to the user. If the user shows interest, the emotion recognition means can detect this and continue with a more detailed commentary.
[1166] In this way, the system of the present invention can meet the diverse needs of users, not only improving operability and convenience but also providing a more personalized experience according to the user's emotions.
[1167] The processing flow will be explained below.
[1168] Understood. Below I will explain the process in concrete steps.
[1169] Step 1:
[1170] The user speaks into the device's microphone to give voice commands such as play, stop, rewind, skip, or explain technical terms.
[1171] Step 2:
[1172] The device captures the user's speech as voice data through a microphone.
[1173] Step 3:
[1174] The voice data captured by the terminal is sent to a voice recognition means, which converts the voice data into text data.
[1175] Step 4:
[1176] The terminal receives the converted text data and passes it to the control means.
[1177] Step 5:
[1178] An emotion recognition means of the terminal analyzes the user's emotions from the captured voice data.
[1179] Step 6:
[1180] The control means analyzes the text data and the emotion recognition results to determine the user's intention.
[1181] Step 7:
[1182] If the command is "play," "stop," "rewind," or "skip," the control invokes the corresponding audio player function (e.g., play, pause, rewind, skip method) to perform the operation on the audio content.
[1183] Step 8:
[1184] If the command requests an explanation of technical terms, the control means uses the search means to obtain an explanation of the specified term from the database.
[1185] Step 9:
[1186] A search means on the server searches a database for an explanation of the requested technical term and returns it to the terminal in text format.
[1187] Step 10:
[1188] The explanatory text acquired by the terminal is passed to a speech synthesis means, which generates voice data using TTS (Text-To-Speech).
[1189] Step 11:
[1190] Based on the emotion recognition results, the reading speed and tone of the audio content are adjusted.
[1191] Step 12:
[1192] The device outputs the generated voice data to the user through a speaker, providing feedback in response to explanations and commands.
[1193] Examples:
[1194] Playback instructions example
[1195] 1. Step 1: The user says "play."
[1196] 2. Step 2: The device captures audio data through the microphone.
[1197] 3. Step 3: The voice data is converted into text using a speech recognition tool.
[1198] 4. Step 4: The converted text "play" is sent to the control means.
[1199] 5. Step 5: The emotion recognition means analyzes the user's emotions and determines, for example, that the user is nervous.
[1200] 6. Step 6: The control means analyzes the "play" command and also takes into account the user's emotions.
[1201] 7. Step 7: The control means calls the play method of the audio player to begin playback.
[1202] 8. Step 11: Based on the emotion recognition results, start playing at a gentle speed and tone.
[1203] 9. Step 12: The terminal outputs the playback of the audio content to the user through the speaker.
[1204] When requesting an explanation of technical terms
[1205] 1. Step 1: The user says, "Please explain to me, AI."
[1206] 2. Step 2: The device captures audio data through the microphone.
[1207] 3. Step 3: The voice data is converted into text using a speech recognition tool.
[1208] 4. Step 4: The converted text "Please explain the AI" is sent to the control means.
[1209] 5. Step 5: The emotion recognition means analyzes the user's emotion and determines, for example, that the user is interested.
[1210] 6. Step 6: The control means analyzes the text "Please explain the AI" and determines the user's intention.
[1211] 7. Step 8: The control means instructs the search means to search for an explanation of "AI."
[1212] 8. Step 9: The server's search function retrieves explanations about "AI" from the database and returns them to the terminal in text format.
[1213] 9. Step 10: The terminal passes the explanatory text to the speech synthesis means to generate speech data.
[1214] 10. Step 11: Based on the emotion recognition results, a detailed commentary is generated in an appropriate tone of voice.
[1215] 11. Step 12: The terminal outputs the generated voice data to the user through the speaker to provide an explanation.
[1216] The above are the processing steps for voice instruction and emotion recognition in the system of the present invention. This detailed processing procedure allows users to operate audio content and receive technical terminology explanations in an intuitive and personalized manner.
[1217] Example 2
[1218] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1219] Conventional speech recognition systems have the ability to operate audio content and provide information based on user voice commands. However, they have the problem of being unable to provide feedback or suggest content based on the user's emotions, which limits the user experience. Another issue is that it is difficult to quickly provide explanations of technical terms and new terms, making it difficult to meet diverse user needs.
[1220] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes a voice input means, a voice recognition means for converting voice data into text data, an emotion recognition means for analyzing emotions from the voice data, a control means for analyzing the text data and the emotion data to operate the audio content, a search means for providing explanations of technical terms and new terms, and a voice synthesis means for converting search results into voice data and outputting the voice. This enables personalized feedback and content suggestions based on the user's emotions, and also makes it possible to quickly provide explanations of technical terms and new terms.
[1221] "Voice input means" refers to a device for capturing a user's speech, and primarily refers to a microphone or similar hardware.
[1222] A "voice recognition means" is a technique or device for converting captured voice data into text data, utilizing a voice recognition algorithm.
[1223] "Emotion recognition means" refers to a means for analyzing and identifying a user's emotions from voice data, and is a technology for determining an emotional state based on information such as voice tone, speed, and strength.
[1224] "Control means" refers to a device or system for analyzing the converted text data and emotion data and for performing manipulations on the audio content.
[1225] "Search tools" refers to a mechanism for retrieving technical terms or explanations of new terms from online or offline databases.
[1226] A "voice synthesis means" is a technique or device for converting text data into voice data and providing voice feedback to the user.
[1227] The present invention provides a system including a voice input unit, a voice recognition unit, a control unit, a search unit, a voice synthesis unit, and an emotion recognition unit. This system allows users to easily operate audio content by voice, and not only quickly obtain explanations of technical terms and new terms, but also realizes more appropriate feedback and content suggestions based on the user's emotions. Specific embodiments are described in detail below.
[1228] System configuration
[1229] 1. Voice input method:
[1230] The device is equipped with a microphone that captures the user's speech as voice data.
[1231] Example: Use a commercially available microphone.
[1232] 2. Voice recognition means:
[1233] The captured voice data is sent to a server and converted into text data by a voice recognition means.
[1234] Examples of software used: General speech recognition engines (e.g., Google Speech-to-Text, Amazon Transcribe)
[1235] 3. Emotion recognition means:
[1236] The user's emotions are analyzed based on the same voice data.
[1237] Example of software used: Emotion recognition engine (e.g. IBM Watson Tone Analyzer)
[1238] 4. Control measures:
[1239] The server analyzes the converted text data and the emotion recognition results and generates instructions to perform appropriate operations.
[1240] Example: When the user says "play", start playing audio content.
[1241] 5. Search methods:
[1242] When an explanation of technical terms or new terms is required, the server uses search means to retrieve relevant information from the database.
[1243] Example: General search engines (e.g. Google Search API)
[1244] 6. Speech synthesis means:
[1245] The obtained information and processing results are converted from text data to audio data and provided to the user through the device's speaker.
[1246] Examples of software used: Speech synthesis engines (e.g., Amazon Polly, Google Text-to-Speech)
[1247] Specific examples
[1248] 1. Playback instructions:
[1249] When a user says "play," the device's microphone captures this voice. The captured voice data is sent to the server, where it is converted into text data ("play") using Google Speech-to-Text. The IBM Watson Tone Analyzer then analyzes the user's emotion, rating it as "excited." Based on this, the server instructs playback of audio content, and uses Amazon Polly to generate voice data such as "Start playback," which is played over the device's speaker.
[1250] 2. Terminology:
[1251] If a user says, "Please explain artificial intelligence," the audio is captured and converted to text using Google Speech-to-Text. Next, the IBM Watson Tone Analyzer evaluates the user's emotion as "interested." The server uses the Google Search API to obtain an explanation of "artificial intelligence" and provides it to the user using Amazon Polly.
[1252] Prompt Sentence Examples
[1253] "How can I start playback with a tone of voice that users find relaxing?"
[1254] "Explain how combining speech recognition and emotion recognition can respond to the user."
[1255] This system allows users to not only control audio content naturally through voice, but also receive personalized feedback based on their emotions. It also provides quick explanations of technical terms and new terms, greatly improving user convenience.
[1256] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1257] Step 1: Capture audio
[1258] Description: The device uses a microphone to capture what the user says.
[1259] Input: User's spoken utterance
[1260] Output: Audio data file
[1261] What it does: When the microphone detects the user speaking, audio data is recorded in real time and temporarily stored for later processing steps.
[1262] Step 2: Voice Recognition
[1263] Description: The device sends the captured voice data to the server, where it is converted into text data using speech recognition means.
[1264] Input: Audio data file
[1265] Output: Text data
[1266] Specific operation: The captured voice data is sent to the server and converted into text data using the Google Speech-to-Text API, a speech recognition engine. For example, if a user says "play," the speech recognition engine generates the text "play."
[1267] Step 3: Sentiment Analysis
[1268] Description: The server performs emotion recognition based on text and voice data.
[1269] Input: Audio data file, text data
[1270] Output: Emotion data
[1271] How it works: The voice data file is passed to an emotion recognition engine such as IBM Watson Tone Analyzer, which analyzes the user's voice for emotions (e.g., "excited" or "relaxed") and saves the results as emotion data.
[1272] Step 4: Command Parsing
[1273] Description: The server analyzes text and emotion data to determine the user's intent.
[1274] Input: Text data, emotion data
[1275] Output: Operation command
[1276] Specific operation: The server combines the converted text data with the analyzed emotional data to determine the user's intention. For example, if the user's text data is "play" and the emotional data is recognized as "excited," the server will determine that "the user wants to play audio content" and generate the corresponding operation command.
[1277] Step 5: Run a search
[1278] Description: When a term explanation is needed, the server uses a search tool to retrieve the relevant information.
[1279] Input: Text data of technical terms
[1280] Output: explanatory text data
[1281] Specific operation: The server uses the Google Search API to search online for the technical term the user is looking for (for example, "artificial intelligence") and obtains the results as explanatory text data.
[1282] Step 6: Audio adjustment and output
[1283] Description: The server generates and outputs voice based on the final text data and emotion data.
[1284] Input: explanatory text data (or text data related to operations), emotion data
[1285] Output: Audio data
[1286] How it works: The server uses Amazon Polly to convert text data into speech, adjusting the speech (e.g., tone and speech rate) based on emotion data, and then sending the resulting speech to the device and playing it back to the user through the speaker.
[1287] (Application example 2)
[1288] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1289] Conventional audio content manipulation systems operate based solely on the user's voice commands. This makes it difficult to provide feedback and suggest content that fully reflects the user's emotions and intentions. Furthermore, explanations of technical terms and new terms rely solely on voice commands, making it difficult to deepen understanding according to the situation. This makes it difficult for users to obtain an intuitive and personalized experience.
[1290] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a voice recognition means, a control means, and a search means. This enables intuitive and personalized operation and feedback of audio content based on the user's voice and emotions.
[1291] Understood. Below are the definitions of important words included in the rewritten patent claims.
[1292] "Audio input means" refers to a device or function that captures the user's voice.
[1293] "Speech recognition means" refers to a device or function for converting captured voice data into text.
[1294] "Controller" refers to a device or function that analyzes text data and emotional state to manipulate and adjust audio content.
[1295] "Search means" refers to a device or function for retrieving explanations of technical terms or new terms from a database.
[1296] "Speech synthesis means" refers to a device or function that converts acquired text information into voice data.
[1297] "Emotion recognition means" refers to a device or function that analyzes the user's emotional state from voice data.
[1298] The "playback position" refers to the current progress position during playback of audio content.
[1299] "Voice operation" refers to a method in which a user inputs operation commands by voice.
[1300] "Speech rate" refers to the speed at which audio content is played aloud.
[1301] "Tone" refers to the tone and intonation of a voice.
[1302] A system for implementing the present invention includes a voice input means, a voice recognition means, an emotion recognition means, a control means, a search means, and a voice synthesis means.
[1303] Program processing explanation
[1304] The server coordinates these means to manipulate and adjust audio content based on the user's voice input and emotional state.
[1305] 1. Voice Recognition Method
[1306] It uses the smartphone's microphone to capture the user's voice.
[1307] The captured voice data is converted into text data using a voice recognition service (Google Cloud Speech-to-Text API).
[1308] 2. Emotion recognition means
[1309] The voice data is sent to an emotion recognition service (IBM Watson Tone Analyzer) to analyze the user's emotional state, which is classified as "happiness," "sadness," "surprise," "anger," etc.
[1310] 3. Control Measures
[1311] The control means analyzes the text data and emotion recognition results to determine the user's intention. For example, if the user says, "Play the next song," the control means understands the intention and issues a command to start playing the next song.
[1312] Adjust reading speed and tone depending on emotional state.
[1313] 4. Search Methods
[1314] When an explanation of a technical term or a new term is needed, the search tool accesses a database (such as the Wikipedia API) to retrieve the relevant information, which is stored as text data.
[1315] 5. Speech synthesis means
[1316] The acquired text information is converted into voice data by a speech synthesis service (Amazon Polly) and provided to the user, with the tone and speed adjusted based on the user's emotional state.
[1317] Specific example explanation
[1318] Example 1: Relaxation scenario
[1319] User: "Play some relaxing music"
[1320] Server Action:
[1321] 1. The voice recognition means generates text data such as "Play some relaxing music."
[1322] 2. The emotion recognition means determines that the user's emotional state is one of seeking "relaxation."
[1323] 3. The control means instructs playback of a relaxation playlist.
[1324] 4. The speech synthesis means audibly conveys the relevant announcement to the user.
[1325] Prompt Sentence Examples
[1326] If the user says they want to relax, this prompt should suggest music to play.
[1327] [Relax, Music, Play]
[1328] Example 2: Scenario where terminology needs to be explained
[1329] User: "What is artificial intelligence?"
[1330] Server Action:
[1331] 1. The text data "What is artificial intelligence?" is generated by the speech recognition means.
[1332] 2. Emotion recognition means identifies the user's interests.
[1333] 3. The control means obtains the necessary information through the search means and analyzes its meaning.
[1334] 4. The search tool retrieves "artificial intelligence" information from the database.
[1335] 5. The speech synthesis means provides the acquired information to the user by speech.
[1336] Prompt Sentence Examples
[1337] When a user asks "What is artificial intelligence?", generate an explanation using the Wikipedia API.
[1338] The above is a specific embodiment for carrying out the invention. This system allows users to enjoy intuitive and personalized audio content manipulation and feedback through voice and emotion.
[1339] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1340] Step 1:
[1341] The user inputs voice data using the smartphone's microphone. For example, the user might say, "Play some relaxing music." This voice data becomes the input.
[1342] Step 2:
[1343] The device captures the voice input data and sends it to a speech recognition service (Google Cloud Speech-to-Text API). The speech recognition service converts the voice data into text data. The input is voice data, and the output is the converted text data.
[1344] Step 3:
[1345] The device sends the converted text data to an emotion recognition service (IBM Watson Tone Analyzer), which analyzes the user's emotional state. For example, it detects that the user wants to relax. The input is text data, and the output is emotion data.
[1346] Step 4:
[1347] The server sends the text data and emotional data to the control means for analysis. The control means determines the user's intention based on the text data. For example, based on the text "Play relaxing music" and the emotional data indicating a desire to relax, the control means selects an appropriate music playlist. The input is text data and emotional data, and the output is an operation command.
[1348] Step 5:
[1349] The server retrieves information from a database (such as Wikipedia API) using search tools to provide explanations of technical terms and new terms as needed. For example, this step is executed when a user asks "What is artificial intelligence?" The input is text data, and the output is the retrieved information.
[1350] Step 6:
[1351] The server passes the acquired information or the selected music playlist to a speech synthesis means (Amazon Polly) and converts it into voice data. For example, the title and summary of the playlist are announced to the user by voice. This voice data is the output. The input is text data, and the output is voice data.
[1352] Step 7:
[1353] The device provides the generated voice data to the user through a speaker. The output from the voice synthesis means is played back as is. For example, a voice notification such as "Relaxing music will be played" is given. The input is the voice data, and the output is the voice feedback provided to the user.
[1354] These processing steps enable users to have intuitive and personalized audio content manipulation and feedback based on their voice input and emotional state.
[1355] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1356] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1357] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1358] [Fourth embodiment]
[1359] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1360] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1361] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1362] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1363] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1364] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1365] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1366] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1367] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1368] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1369] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1370] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1371] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1372] Understood. Below is the description of the "Mode for Carrying Out the Invention" of the patent specification.
[1373] The present invention provides a system including a voice input unit, a voice recognition unit, a control unit, a search unit, and a voice synthesis unit. This system allows a user to easily operate audio content by voice and quickly obtain explanations of technical terms and new terms. Specific embodiments of the present invention are described in detail below.
[1374] System configuration:
[1375] 1. Voice input method:
[1376] The device is equipped with a microphone that captures the user's speech as voice data.
[1377] This voice input means provides a basis for responding to operational commands and questions uttered by the user.
[1378] 2. Voice recognition means:
[1379] The captured voice data is sent to a voice recognition means and converted into text data.
[1380] This speech recognizer uses advanced algorithms to extract accurate text from speech.
[1381] 3. Control measures:
[1382] The converted text data is analyzed by the control means and the corresponding operation is performed.
[1383] For example, if the user says "play," the control will begin playing audio content.
[1384] 4. Search methods:
[1385] When a user requests an explanation of a particular term, the control means uses the search means to quickly obtain the meaning of that term.
[1386] The search tool accesses an extensive database and retrieves detailed explanations of technical terms.
[1387] 5. Speech synthesis means:
[1388] The acquired text information is passed to a voice synthesis means and converted into voice data.
[1389] The speech synthesis means generates natural speech in real time and provides audible feedback to the user.
[1390] Explanation of program operation:
[1391] 1. Speech Recognition:
[1392] The device captures the user's speech and converts it into text using a speech recognition service.
[1393] The converted text data is passed to the next processing step.
[1394] 2. Command analysis:
[1395] The control means analyzes the converted text data and determines the user's intention.
[1396] Specific actions are determined in response to control commands (e.g., play, stop, rewind, etc.) and questions (e.g., "What does AI mean?").
[1397] 3. Perform a search:
[1398] When an explanation of a technical term is required, the control means calls the search means and retrieves the meaning of the relevant term from the database.
[1399] The information obtained is stored as text.
[1400] 4. Audio output:
[1401] The final information is passed to a speech synthesis means and converted into speech data.
[1402] The terminal plays the generated voice data to the user through a speaker.
[1403] Examples:
[1404] 1. Playback instructions:
[1405] When the user utters "play", the terminal captures this speech and converts it into text "play" through speech recognition means.
[1406] The control means recognizes the command "play" and starts playing the audio content.
[1407] 2. Terminology:
[1408] If a user says, "Please explain artificial intelligence," the device converts the speech into text and recognizes it as "Please explain artificial intelligence."
[1409] The control means determines that an explanation of "artificial intelligence" is desired, and uses the search means to obtain the meaning of "artificial intelligence."
[1410] The acquired commentary is converted into voice data by a voice synthesis means and provided to the user.
[1411] In this way, the system of the present invention can meet the diverse needs of users, improving operability and convenience.
[1412] The processing flow will be explained below.
[1413] Understood. Below, I will explain the process flow in concrete steps.
[1414] Step 1:
[1415] The user speaks into the device's microphone to give voice commands such as play, stop, rewind, skip, or explain technical terms.
[1416] Step 2:
[1417] The device captures the user's speech as voice data through a microphone.
[1418] Step 3:
[1419] The voice data captured by the device is sent to the voice recognition means and converted into text data using the Google voice recognition service.
[1420] Step 4:
[1421] The terminal receives the converted text data and passes it to the control means.
[1422] Step 5:
[1423] The control means of the terminal analyzes the text data to determine the user's intended command (eg, play, stop, rewind, skip, or commentary).
[1424] Step 6:
[1425] If the command is "play," "stop," "rewind," or "skip," the device's control invokes the corresponding audio player function (e.g., the play, pause, rewind, or skip method) to perform the audio content operation.
[1426] Step 7:
[1427] If the command requests an explanation of a technical term, the control means of the terminal utilizes the search means to obtain an explanation of the specified term from the database.
[1428] Step 8:
[1429] A search means on the server searches a database for an explanation of the requested technical term and returns it to the terminal in text format.
[1430] Step 9:
[1431] The explanatory text acquired by the terminal is passed to a speech synthesis means, which generates voice data using TTS (Text-To-Speech).
[1432] Step 10:
[1433] The device outputs the generated voice data to the user through a speaker, providing feedback in response to explanations and commands.
[1434] The above are the steps for processing voice instructions in the system of the present invention. This detailed processing procedure allows the user to intuitively operate audio content and receive explanations of technical terms.
[1435] Example 1
[1436] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1437] In conventional voice-operated systems, when users operate audio content by voice, feedback on operations and explanations of technical terms can be delayed, resulting in a lack of convenience. Additionally, there is a lack of a mechanism for rapid and accurate voice recognition and subsequent processing.
[1438] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1439] In this invention, the server includes a voice input means, a voice recognition means for converting voice data into text, a control means for analyzing the text data and operating the media content, a search means for providing explanations of technical terms and undefined terms, and a voice synthesis means for outputting the obtained explanations by voice, thereby enabling quick and accurate provision of information and operation of the media content by voice operation.
[1440] 1. "Voice input means" means any device or technology used to capture a user's speech as voice data.
[1441] 2. "Speech recognition means" means a technology or device for converting voice data captured by a voice input means into text data.
[1442] 3. "Control means" refers to the technology or device that analyzes the text data converted by the speech recognition means and controls the media content.
[1443] 4. "Search means" refers to the technology or device that allows users to obtain explanations of technical terms or undefined terms they are looking for.
[1444] 5. "Speech synthesis means" refers to technology or equipment for outputting the acquired commentary in audio form.
[1445] 6. "Media Content" refers to all content, including audio and video, that can be used through voice control.
[1446] 7. "Prompt" means a voice command that allows a user to give verbal instructions to a system.
[1447] The present invention provides a system including a voice input unit, a voice recognition unit, a control unit, a search unit, and a voice synthesis unit. This system allows users to easily operate media content by voice and quickly obtain explanations of technical terms and new terms. Specific embodiments of the present invention are described in detail below.
[1448] System configuration
[1449] 1. Voice input method:
[1450] The device is equipped with a microphone to capture the user's speech as voice data. This voice input means provides a basis for responding to user commands and questions.
[1451] 2. Voice recognition means:
[1452] The voice data captured by the device is sent to a server, which converts the voice data into text using a speech recognition service such as the Google Cloud Speech-to-Text API, which uses advanced algorithms to extract accurate text from the speech.
[1453] 3. Control measures:
[1454] The server analyzes the converted text data to determine the user's intent and determines the specific actions to take in response to commands (e.g., play, stop, rewind, etc.) or questions (e.g., "What does AI mean?").
[1455] 4. Search methods:
[1456] When a user requests an explanation of a particular term, the server uses a search engine to quickly retrieve the meaning of the term, accessing the Wikipedia API and other extensive databases to retrieve detailed explanations of technical terms.
[1457] 5. Speech synthesis means:
[1458] The text information acquired by the server is converted into voice data using a speech synthesis service such as Amazon Polly, which generates natural-sounding voices in real time and provides voice feedback to the user.
[1459] Specific examples
[1460] 1. Playback instructions:
[1461] When the user says "play," the terminal captures this speech and converts it into text through speech recognition means as "play." The server recognizes the command "play" and starts playing the media content.
[1462] 2. Terminology:
[1463] When a user says, "Please explain artificial intelligence," the device converts the speech into text and recognizes it as "Please explain artificial intelligence." The server determines that an explanation of "artificial intelligence" is being sought, and uses a search means to obtain the meaning of "artificial intelligence." The obtained explanation is converted into voice data by a voice synthesis means and provided to the user.
[1464] Prompt Sentence Examples
[1465] 1. Playback request prompt:
[1466] When someone says "play," use speech recognition to convert it to text and provide instructions for starting to play media content.
[1467] 2. Glossary prompt:
[1468] If someone says "Please explain artificial intelligence," please provide instructions on how to use the Wikipedia API to obtain an explanation of artificial intelligence, synthesize it into voice, and provide it to the user.
[1469] The system of the present invention is thus able to meet the diverse needs of users and greatly improves operability and convenience.
[1470] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1471] Step 1: Voice Input
[1472] Users speak voice commands to control media content and to explain technical terms.
[1473] The terminal uses a microphone to capture the user's speech as voice data.
[1474] Input: User speech
[1475] Output: Audio data
[1476] Step 2: Voice Recognition
[1477] The terminal transmits the captured audio data to the server.
[1478] The server converts the voice data into text data using the Google Cloud Speech-to-Text API.
[1479] Input: Audio data
[1480] Data processing: Converting voice data into text data
[1481] Output: Text data
[1482] Step 3: Command analysis
[1483] The server analyzes the text data obtained by the voice recognition means and determines whether the spoken content is an operation command or a question.
[1484] For example, text data such as "play" is determined to be an operation command, and text data such as "please explain artificial intelligence" is determined to be a question.
[1485] Input: Text data
[1486] Data operation: Content analysis and category determination of text data (operation commands or questions)
[1487] Output: The result of the operation command or question
[1488] Step 4: Execute the operation command
[1489] If the server recognizes the operation command, it performs the corresponding operation on the media content.
[1490] For example, the command "play" will call the play function of the media player.
[1491] Input: Operation command
[1492] Data operation: Executes media operations corresponding to commands.
[1493] Output: Playing, stopping, etc. media content
[1494] Step 5: Question Processing
[1495] When the server recognizes the question, it uses a search means to obtain explanations of technical terms and undefined terms.
[1496] For example, in response to the question "Please explain artificial intelligence," an explanation of "artificial intelligence" is retrieved from the Wikipedia API.
[1497] Input: Question text data
[1498] Data processing: Performing searches that correspond to the query
[1499] Output: Search results text data
[1500] Step 6: Text-to-Speech
[1501] The server converts the acquired text data of the explanation into audio data using Amazon Polly.
[1502] Input: Text data explaining the question
[1503] Data processing: Convert text data into audio data
[1504] Output: Audio data
[1505] Step 7: Audio Output
[1506] The terminal outputs the audio data transmitted from the server to the user through a speaker.
[1507] Input: Audio data
[1508] Output: Play audio
[1509] This completes the processing flow. The system of the present invention has the advantage that users can quickly and accurately use information and media content through voice operations.
[1510] (Application example 1)
[1511] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1512] Conventional content distribution services require users to physically operate the service and check search results, resulting in problems with usability and convenience. In particular, when requesting explanations of technical terms, users are often forced to search and check text information, which often interrupts the viewing experience. While systems that allow voice control exist, they lack the functionality to simultaneously operate viewing content and provide terminology explanations in real time. This degrades the user experience and makes it difficult to provide a comfortable viewing environment.
[1513] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1514] In this invention, the server includes a voice input means, a voice recognition means for converting voice data into text, a control means for analyzing the text data and operating the voice and audiovisual content, a search means for providing explanations of technical terms and new terms, a voice synthesis means for outputting the search results by voice, and a means for providing audio instructions for operating the audiovisual content and explanations of terms in real time. This allows the user to intuitively operate the audiovisual content by voice instructions while instantly obtaining audio explanations of technical terms, thereby creating an optimal usage environment.
[1515] "Voice input means" refers to a device or system, such as a microphone, for capturing a user's speech as voice data.
[1516] A "voice recognition means" is a technique or system for converting captured voice data into text data.
[1517] The "control means" is a system or method that analyzes text data and operates the audiovisual content based on the analysis.
[1518] A "search tool" is a system or method for quickly retrieving information about a particular term or terminology.
[1519] "Speech synthesis means" refers to a system or technology that converts text information into voice data and provides it to the user in voice form.
[1520] "Viewing content" refers to media content such as audio and video that a user views.
[1521] A "real-time audio feedback mechanism" is a system or method for providing immediate audio feedback to a user upon request.
[1522] System configuration
[1523] This invention provides a system that allows a user to operate viewing content through voice input and obtain audio explanations of technical terms in real time. The system includes the following main means.
[1524] 1. Voice input method:
[1525] - A microphone is used as a voice input means to capture the user's speech as voice data. A microphone built into a smartphone, smart glasses, a head-mounted display, or a robot is used.
[1526] 2. Voice recognition means:
[1527] - Use speech recognition software such as Google Cloud Speech-to-Text API to convert the captured voice data into text data, thereby capturing the user's voice instructions as text.
[1528] 3. Control measures:
[1529] - Analyze text data and manipulate viewing content. Use Python and natural language processing libraries (NLTK and spaCy) to analyze text data from speech recognition tools and perform operations such as play, stop, and rewind.
[1530] 4. Search methods:
[1531] - Use Elasticsearch or other database search technologies to provide explanations of technical terms and new terms. Search for explanations of terms captured as text and return them in text format.
[1532] 5. Speech synthesis means:
[1533] - Use speech synthesis technology such as Google Cloud Text-to-Speech API to convert text information obtained from search tools into audio data and provide feedback to users.
[1534] 6. Real-time feedback methods:
[1535] - Provide audio instructions and glossary of terms for the content in real time, ensuring seamless operation and commentary of the content.
[1536] Usage and Examples
[1537] In a real-world usage scenario, a user will interface with the system using a smartphone or other compatible device. The following example illustrates how to operate the system.
[1538] 1. Playback instructions:
[1539] - When the user says "play," the microphone captures the voice and converts it into text "play" using the Google Cloud Speech-to-Text API. The control means recognizes this and starts playing the content.
[1540] 2. Terminology:
[1541] - When a user says, "Please explain artificial intelligence," the speech is converted into text in the same way. This text is analyzed by the control means, and an explanation of "artificial intelligence" is retrieved from the Elasticsearch database through the search means. The retrieved text information is converted into audio data using the Google Cloud Text-to-Speech API and provided to the user in real time.
[1542] Example prompt sentence:
[1543] "Explain the following statement: 'Artificial intelligence refers to systems designed to have human intelligence.'"
[1544] These methods allow users to intuitively control content through voice commands and instantly receive explanations of technical terms, improving the user experience and ease of use.
[1545] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1546] Step 1:
[1547] The user speaks a voice command to the device (e.g., "play," "please explain with AI," etc.). The input is the user's voice, and the output is the audio data captured by the device's microphone.
[1548] Step 2:
[1549] The terminal captures the user's voice using the voice input means and obtains it as voice data. The input is the user's voice and the output is voice data.
[1550] Step 3:
[1551] The device sends voice data to the Google Cloud Speech-to-Text API, which converts the voice data into text. The input is voice data, and the output is text data. The server calls a Google Cloud service to convert the voice data.
[1552] Step 4:
[1553] The server uses a control means to analyze the received text data and determine the operation and explanation acquisition instructions based on that. The input is text data, and the output is the operation instructions or search query as the analysis result. The specific operation is to determine the user's intention using a text data analysis algorithm.
[1554] Step 5:
[1555] When the server operates the viewing content based on the analysis results, it executes operations such as play, stop, and rewind. The input is the operation command, and the output is the actual content operation. The specific operation is to call the media player's control API and execute the required operation.
[1556] Step 6:
[1557] When a term explanation is requested, the server uses a search mechanism to search for relevant information in a database. The input is a search query, and the output is text information as search results. The specific operation is to retrieve the appropriate information using Elasticsearch or other database search technologies.
[1558] Step 7:
[1559] The server sends the text information acquired through the search tool to the Google Cloud Text-to-Speech API and converts it into voice data. The input is text information and the output is voice data. The specific operation is to generate voice from text using the voice synthesis API.
[1560] Step 8:
[1561] The terminal plays back the voice data obtained from the voice synthesis means and provides real-time feedback to the user. The input is voice data, and the output is voice feedback to the user. The specific operation is to play back the voice data using the terminal's speaker.
[1562] Step 9:
[1563] The flow from step 1 to step 8 is repeated until the user utters a new command or performs the next operation. The input is the next voice command, and the output is the capture of the next voice data.
[1564] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1565] Understood. Below is a description of the "Mode for carrying out the invention" based on the invention that combines an emotion engine.
[1566] The present invention provides a system including a voice input unit, a voice recognition unit, a control unit, a search unit, a voice synthesis unit, and an emotion recognition unit. This system allows users to easily operate audio content by voice, quickly obtain explanations of technical terms and new terms, and also realizes more appropriate feedback and content suggestions based on the user's emotions. Specific embodiments of the present invention are described in detail below.
[1567] System configuration:
[1568] 1. Voice input method:
[1569] The device is equipped with a microphone that captures the user's speech as voice data.
[1570] The voice input means provides a basis for responding to operational commands and questions uttered by the user.
[1571] 2. Voice recognition means:
[1572] The captured voice data is sent to a voice recognition means and converted into text data.
[1573] The speech recognizer uses advanced algorithms to extract accurate text from speech.
[1574] 3. Emotion recognition means:
[1575] Analyze user emotions from captured voice data.
[1576] The emotion recognition means identifies the emotional state of the user based on information such as the tone, speed, and strength of the user's voice.
[1577] 4. Control measures:
[1578] The converted text data and the emotion data obtained by the emotion recognition means are analyzed, and a corresponding operation is performed.
[1579] For example, when the user says "play," the control means starts playing audio content and adjusts the speaking speed and tone based on the user's emotions.
[1580] 5. Search methods:
[1581] When a user requests an explanation of a particular term, the control means uses the search means to quickly obtain the meaning of that term.
[1582] The search tool accesses an extensive database and retrieves detailed explanations of technical terms.
[1583] 6. Speech synthesis means:
[1584] The acquired text information is passed to a voice synthesis means and converted into voice data.
[1585] The speech synthesis means generates natural speech in real time and provides audible feedback to the user.
[1586] Explanation of program operation:
[1587] 1. Speech Recognition:
[1588] The device captures the user's speech and converts it into text using a speech recognition service.
[1589] The converted text data is passed to the next processing step.
[1590] 2. Emotion analysis:
[1591] The emotion recognition means of the terminal analyzes the user's emotions from the captured voice data.
[1592] The results of the sentiment analysis are passed to the control means and used in the next processing step.
[1593] 3. Command Analysis:
[1594] The control means analyzes the converted text data and the emotion analysis data from the emotion recognition means to determine the user's intention.
[1595] Specific actions corresponding to operation commands (e.g., play, stop, rewind) or questions (e.g., "What does XX mean?") are determined.
[1596] 4. Perform a search:
[1597] When an explanation of a technical term is required, the control means calls the search means and retrieves the meaning of the relevant term from the database.
[1598] The information obtained is stored as text.
[1599] 5. Adjustment and Audio Output:
[1600] Based on the emotion recognition results, the reading speed and tone of the audio content are adjusted.
[1601] The final information is passed to a speech synthesis means and converted into speech data.
[1602] The device outputs the generated voice data to the user through a speaker, providing feedback in response to explanations and commands.
[1603] Examples:
[1604] 1. Playback instructions:
[1605] When the user utters "play", the terminal captures this speech and converts it into text "play" through speech recognition means.
[1606] The control means recognizes the command "play" and starts playing the audio content. Furthermore, the emotion recognition means analyzes the user's emotion, and if the user is nervous, for example, starts playing the audio content in a calm tone.
[1607] 2. Terminology:
[1608] If a user says, "Please explain artificial intelligence," the device converts the speech into text and recognizes it as "Please explain artificial intelligence."
[1609] The control means determines that an explanation of "artificial intelligence" is desired, and uses the search means to obtain the meaning of "artificial intelligence."
[1610] The acquired commentary is converted into voice data by a voice synthesis means and provided to the user. If the user shows interest, the emotion recognition means can detect this and continue with a more detailed commentary.
[1611] In this way, the system of the present invention can meet the diverse needs of users, not only improving operability and convenience but also providing a more personalized experience according to the user's emotions.
[1612] The processing flow will be explained below.
[1613] Understood. Below I will explain the process in concrete steps.
[1614] Step 1:
[1615] The user speaks into the device's microphone to give voice commands such as play, stop, rewind, skip, or explain technical terms.
[1616] Step 2:
[1617] The device captures the user's speech as voice data through a microphone.
[1618] Step 3:
[1619] The voice data captured by the terminal is sent to a voice recognition means, which converts the voice data into text data.
[1620] Step 4:
[1621] The terminal receives the converted text data and passes it to the control means.
[1622] Step 5:
[1623] An emotion recognition means of the terminal analyzes the user's emotions from the captured voice data.
[1624] Step 6:
[1625] The control means analyzes the text data and the emotion recognition results to determine the user's intention.
[1626] Step 7:
[1627] If the command is "play," "stop," "rewind," or "skip," the control invokes the corresponding audio player function (e.g., play, pause, rewind, skip method) to perform the operation on the audio content.
[1628] Step 8:
[1629] If the command requests an explanation of technical terms, the control means uses the search means to obtain an explanation of the specified term from the database.
[1630] Step 9:
[1631] A search means on the server searches a database for an explanation of the requested technical term and returns it to the terminal in text format.
[1632] Step 10:
[1633] The explanatory text acquired by the terminal is passed to a speech synthesis means, which generates voice data using TTS (Text-To-Speech).
[1634] Step 11:
[1635] Based on the emotion recognition results, the reading speed and tone of the audio content are adjusted.
[1636] Step 12:
[1637] The device outputs the generated voice data to the user through a speaker, providing feedback in response to explanations and commands.
[1638] Examples:
[1639] Playback instructions example
[1640] 1. Step 1: The user says "play."
[1641] 2. Step 2: The device captures audio data through the microphone.
[1642] 3. Step 3: The voice data is converted into text using a speech recognition tool.
[1643] 4. Step 4: The converted text "play" is sent to the control means.
[1644] 5. Step 5: The emotion recognition means analyzes the user's emotions and determines, for example, that the user is nervous.
[1645] 6. Step 6: The control means analyzes the "play" command and also takes into account the user's emotions.
[1646] 7. Step 7: The control means calls the play method of the audio player to begin playback.
[1647] 8. Step 11: Based on the emotion recognition results, start playing at a gentle speed and tone.
[1648] 9. Step 12: The terminal outputs the playback of the audio content to the user through the speaker.
[1649] When requesting an explanation of technical terms
[1650] 1. Step 1: The user says, "Please explain to me, AI."
[1651] 2. Step 2: The device captures audio data through the microphone.
[1652] 3. Step 3: The voice data is converted into text using a speech recognition tool.
[1653] 4. Step 4: The converted text "Please explain the AI" is sent to the control means.
[1654] 5. Step 5: The emotion recognition means analyzes the user's emotion and determines, for example, that the user is interested.
[1655] 6. Step 6: The control means analyzes the text "Please explain the AI" and determines the user's intention.
[1656] 7. Step 8: The control means instructs the search means to search for an explanation of "AI."
[1657] 8. Step 9: The server's search function retrieves explanations about "AI" from the database and returns them to the terminal in text format.
[1658] 9. Step 10: The terminal passes the explanatory text to the speech synthesis means to generate speech data.
[1659] 10. Step 11: Based on the emotion recognition results, a detailed commentary is generated in an appropriate tone of voice.
[1660] 11. Step 12: The terminal outputs the generated voice data to the user through the speaker to provide an explanation.
[1661] The above are the processing steps for voice instruction and emotion recognition in the system of the present invention. This detailed processing procedure allows users to operate audio content and receive technical terminology explanations in an intuitive and personalized manner.
[1662] Example 2
[1663] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1664] Conventional speech recognition systems have the ability to operate audio content and provide information based on user voice commands. However, they have the problem of being unable to provide feedback or suggest content based on the user's emotions, which limits the user experience. Another issue is that it is difficult to quickly provide explanations of technical terms and new terms, making it difficult to meet diverse user needs.
[1665] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes a voice input means, a voice recognition means for converting voice data into text data, an emotion recognition means for analyzing emotions from the voice data, a control means for analyzing the text data and the emotion data to operate the audio content, a search means for providing explanations of technical terms and new terms, and a voice synthesis means for converting search results into voice data and outputting the voice. This enables personalized feedback and content suggestions based on the user's emotions, and also makes it possible to quickly provide explanations of technical terms and new terms.
[1666] "Voice input means" refers to a device for capturing a user's speech, and primarily refers to a microphone or similar hardware.
[1667] A "voice recognition means" is a technique or device for converting captured voice data into text data, utilizing a voice recognition algorithm.
[1668] "Emotion recognition means" refers to a means for analyzing and identifying a user's emotions from voice data, and is a technology for determining an emotional state based on information such as voice tone, speed, and strength.
[1669] "Control means" refers to a device or system for analyzing the converted text data and emotion data and for performing manipulations on the audio content.
[1670] "Search tools" refers to a mechanism for retrieving technical terms or explanations of new terms from online or offline databases.
[1671] A "voice synthesis means" is a technique or device for converting text data into voice data and providing voice feedback to the user.
[1672] The present invention provides a system including a voice input unit, a voice recognition unit, a control unit, a search unit, a voice synthesis unit, and an emotion recognition unit. This system allows users to easily operate audio content by voice, and not only quickly obtain explanations of technical terms and new terms, but also realizes more appropriate feedback and content suggestions based on the user's emotions. Specific embodiments are described in detail below.
[1673] System configuration
[1674] 1. Voice input method:
[1675] The device is equipped with a microphone that captures the user's speech as voice data.
[1676] Example: Use a commercially available microphone.
[1677] 2. Voice recognition means:
[1678] The captured voice data is sent to a server and converted into text data by a voice recognition means.
[1679] Examples of software used: General speech recognition engines (e.g., Google Speech-to-Text, Amazon Transcribe)
[1680] 3. Emotion recognition means:
[1681] The user's emotions are analyzed based on the same voice data.
[1682] Example of software used: Emotion recognition engine (e.g. IBM Watson Tone Analyzer)
[1683] 4. Control measures:
[1684] The server analyzes the converted text data and the emotion recognition results and generates instructions to perform appropriate operations.
[1685] Example: When the user says "play", start playing audio content.
[1686] 5. Search methods:
[1687] When an explanation of technical terms or new terms is required, the server uses search means to retrieve relevant information from the database.
[1688] Example: General search engines (e.g. Google Search API)
[1689] 6. Speech synthesis means:
[1690] The obtained information and processing results are converted from text data to audio data and provided to the user through the device's speaker.
[1691] Examples of software used: Speech synthesis engines (e.g., Amazon Polly, Google Text-to-Speech)
[1692] Specific examples
[1693] 1. Playback instructions:
[1694] When a user says "play," the device's microphone captures this voice. The captured voice data is sent to the server, where it is converted into text data ("play") using Google Speech-to-Text. The IBM Watson Tone Analyzer then analyzes the user's emotion, rating it as "excited." Based on this, the server instructs playback of audio content, and uses Amazon Polly to generate voice data such as "Start playback," which is played over the device's speaker.
[1695] 2. Terminology:
[1696] If a user says, "Please explain artificial intelligence," the audio is captured and converted to text using Google Speech-to-Text. Next, the IBM Watson Tone Analyzer evaluates the user's emotion as "interested." The server uses the Google Search API to obtain an explanation of "artificial intelligence" and provides it to the user using Amazon Polly.
[1697] Prompt Sentence Examples
[1698] "How can I start playback with a tone of voice that users find relaxing?"
[1699] "Explain how combining speech recognition and emotion recognition can respond to the user."
[1700] This system allows users to not only control audio content naturally through voice, but also receive personalized feedback based on their emotions. It also provides quick explanations of technical terms and new terms, greatly improving user convenience.
[1701] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1702] Step 1: Capture audio
[1703] Description: The device uses a microphone to capture what the user says.
[1704] Input: User's spoken utterance
[1705] Output: Audio data file
[1706] What it does: When the microphone detects the user speaking, audio data is recorded in real time and temporarily stored for later processing steps.
[1707] Step 2: Voice Recognition
[1708] Description: The device sends the captured voice data to the server, where it is converted into text data using speech recognition means.
[1709] Input: Audio data file
[1710] Output: Text data
[1711] Specific operation: The captured voice data is sent to the server and converted into text data using the Google Speech-to-Text API, a speech recognition engine. For example, if a user says "play," the speech recognition engine generates the text "play."
[1712] Step 3: Sentiment Analysis
[1713] Description: The server performs emotion recognition based on text and voice data.
[1714] Input: Audio data file, text data
[1715] Output: Emotion data
[1716] How it works: The voice data file is passed to an emotion recognition engine such as IBM Watson Tone Analyzer, which analyzes the user's voice for emotions (e.g., "excited" or "relaxed") and saves the results as emotion data.
[1717] Step 4: Command Parsing
[1718] Description: The server analyzes text and emotion data to determine the user's intent.
[1719] Input: Text data, emotion data
[1720] Output: Operation command
[1721] Specific operation: The server combines the converted text data with the analyzed emotional data to determine the user's intention. For example, if the user's text data is "play" and the emotional data is recognized as "excited," the server will determine that "the user wants to play audio content" and generate the corresponding operation command.
[1722] Step 5: Run a search
[1723] Description: When a term explanation is needed, the server uses a search tool to retrieve the relevant information.
[1724] Input: Text data of technical terms
[1725] Output: explanatory text data
[1726] Specific operation: The server uses the Google Search API to search online for the technical term the user is looking for (for example, "artificial intelligence") and obtains the results as explanatory text data.
[1727] Step 6: Audio adjustment and output
[1728] Description: The server generates and outputs voice based on the final text data and emotion data.
[1729] Input: explanatory text data (or text data related to operations), emotion data
[1730] Output: Audio data
[1731] How it works: The server uses Amazon Polly to convert text data into speech, adjusting the speech (e.g., tone and speech rate) based on emotion data, and then sending the resulting speech to the device and playing it back to the user through the speaker.
[1732] (Application example 2)
[1733] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1734] Conventional audio content manipulation systems operate based solely on the user's voice commands. This makes it difficult to provide feedback and suggest content that fully reflects the user's emotions and intentions. Furthermore, explanations of technical terms and new terms rely solely on voice commands, making it difficult to deepen understanding according to the situation. This makes it difficult for users to obtain an intuitive and personalized experience.
[1735] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a voice recognition means, a control means, and a search means. This enables intuitive and personalized operation and feedback of audio content based on the user's voice and emotions.
[1736] Understood. Below are the definitions of important words included in the rewritten patent claims.
[1737] "Audio input means" refers to a device or function that captures the user's voice.
[1738] "Speech recognition means" refers to a device or function for converting captured voice data into text.
[1739] "Controller" refers to a device or function that analyzes text data and emotional state to manipulate and adjust audio content.
[1740] "Search means" refers to a device or function for retrieving explanations of technical terms or new terms from a database.
[1741] "Speech synthesis means" refers to a device or function that converts acquired text information into voice data.
[1742] "Emotion recognition means" refers to a device or function that analyzes the user's emotional state from voice data.
[1743] The "playback position" refers to the current progress position during playback of audio content.
[1744] "Voice operation" refers to a method in which a user inputs operation commands by voice.
[1745] "Speech rate" refers to the speed at which audio content is played aloud.
[1746] "Tone" refers to the tone and intonation of a voice.
[1747] A system for implementing the present invention includes a voice input means, a voice recognition means, an emotion recognition means, a control means, a search means, and a voice synthesis means.
[1748] Program processing explanation
[1749] The server coordinates these means to manipulate and adjust audio content based on the user's voice input and emotional state.
[1750] 1. Voice Recognition Method
[1751] It uses the smartphone's microphone to capture the user's voice.
[1752] The captured voice data is converted into text data using a voice recognition service (Google Cloud Speech-to-Text API).
[1753] 2. Emotion recognition means
[1754] The voice data is sent to an emotion recognition service (IBM Watson Tone Analyzer) to analyze the user's emotional state, which is classified as "happiness," "sadness," "surprise," "anger," etc.
[1755] 3. Control Measures
[1756] The control means analyzes the text data and emotion recognition results to determine the user's intention. For example, if the user says, "Play the next song," the control means understands the intention and issues a command to start playing the next song.
[1757] Adjust reading speed and tone depending on emotional state.
[1758] 4. Search Methods
[1759] When an explanation of a technical term or a new term is needed, the search tool accesses a database (such as the Wikipedia API) to retrieve the relevant information, which is stored as text data.
[1760] 5. Speech synthesis means
[1761] The acquired text information is converted into voice data by a speech synthesis service (Amazon Polly) and provided to the user, with the tone and speed adjusted based on the user's emotional state.
[1762] Specific example explanation
[1763] Example 1: Relaxation scenario
[1764] User: "Play some relaxing music"
[1765] Server Action:
[1766] 1. The voice recognition means generates text data such as "Play some relaxing music."
[1767] 2. The emotion recognition means determines that the user's emotional state is one of seeking "relaxation."
[1768] 3. The control means instructs playback of a relaxation playlist.
[1769] 4. The speech synthesis means audibly conveys the relevant announcement to the user.
[1770] Prompt Sentence Examples
[1771] If the user says they want to relax, this prompt should suggest music to play.
[1772] [Relax, Music, Play]
[1773] Example 2: Scenario where terminology needs to be explained
[1774] User: "What is artificial intelligence?"
[1775] Server Action:
[1776] 1. The text data "What is artificial intelligence?" is generated by the speech recognition means.
[1777] 2. Emotion recognition means identifies the user's interests.
[1778] 3. The control means obtains the necessary information through the search means and analyzes its meaning.
[1779] 4. The search tool retrieves "artificial intelligence" information from the database.
[1780] 5. The speech synthesis means provides the acquired information to the user by speech.
[1781] Prompt Sentence Examples
[1782] When a user asks "What is artificial intelligence?", generate an explanation using the Wikipedia API.
[1783] The above is a specific embodiment for carrying out the invention. This system allows users to enjoy intuitive and personalized audio content manipulation and feedback through voice and emotion.
[1784] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1785] Step 1:
[1786] The user inputs voice data using the smartphone's microphone. For example, the user might say, "Play some relaxing music." This voice data becomes the input.
[1787] Step 2:
[1788] The device captures the voice input data and sends it to a speech recognition service (Google Cloud Speech-to-Text API). The speech recognition service converts the voice data into text data. The input is voice data, and the output is the converted text data.
[1789] Step 3:
[1790] The device sends the converted text data to an emotion recognition service (IBM Watson Tone Analyzer), which analyzes the user's emotional state. For example, it detects that the user wants to relax. The input is text data, and the output is emotion data.
[1791] Step 4:
[1792] The server sends the text data and emotional data to the control means for analysis. The control means determines the user's intention based on the text data. For example, based on the text "Play relaxing music" and the emotional data indicating a desire to relax, the control means selects an appropriate music playlist. The input is text data and emotional data, and the output is an operation command.
[1793] Step 5:
[1794] The server retrieves information from a database (such as Wikipedia API) using search tools to provide explanations of technical terms and new terms as needed. For example, this step is executed when a user asks "What is artificial intelligence?" The input is text data, and the output is the retrieved information.
[1795] Step 6:
[1796] The server passes the acquired information or the selected music playlist to a speech synthesis means (Amazon Polly) and converts it into voice data. For example, the title and summary of the playlist are announced to the user by voice. This voice data is the output. The input is text data, and the output is voice data.
[1797] Step 7:
[1798] The device provides the generated voice data to the user through a speaker. The output from the voice synthesis means is played back as is. For example, a voice notification such as "Relaxing music will be played" is given. The input is the voice data, and the output is the voice feedback provided to the user.
[1799] These processing steps enable users to have intuitive and personalized audio content manipulation and feedback based on their voice input and emotional state.
[1800] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1801] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1802] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1803] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1804] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1805] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1806] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1807] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1808] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1809] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1810] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1811] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1812] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1813] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1814] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1815] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1816] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1817] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1818] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1819] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1820] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1821] The following is further disclosed regarding the above embodiment.
[1822] Understood. Below are the proposed claims:
[1823] (Claim 1)
[1824] A voice input means;
[1825] a speech recognition means for converting speech data into text;
[1826] a control means for analyzing the text data and manipulating the audio content;
[1827] A search tool that provides explanations of technical terms and new terms;
[1828] a speech synthesis means for outputting the search results by voice;
[1829] A system including:
[1830] (Claim 2)
[1831] 10. The system of claim 1, further comprising the ability to resume playback from a previous playback position.
[1832] (Claim 3)
[1833] 10. The system of claim 1, wherein the reading speed and tone can be adjusted by voice control.
[1834] "Example 1"
[1835] (Claim 1)
[1836] A voice input means;
[1837] a speech recognition means for converting speech data into text;
[1838] a control means for analyzing the text data and manipulating the media content;
[1839] A search tool that provides explanations of technical terms and undefined terms;
[1840] a voice synthesis means for outputting the acquired commentary by voice;
[1841] A system including:
[1842] (Claim 2)
[1843] 10. The system of claim 1, further comprising the capability of initiating playback based on a prompt to play media content.
[1844] (Claim 3)
[1845] 2. The system according to claim 1, further comprising a function of acquiring and outputting an explanation of a term in voice in response to a prompt for the explanation of the term.
[1846] "Application Example 1"
[1847] (Claim 1)
[1848] A voice input means;
[1849] a speech recognition means for converting speech data into text;
[1850] a control means for analyzing the text data and manipulating the audio and audio content;
[1851] A search tool that provides explanations of technical terms and new terms;
[1852] a speech synthesis means for outputting the search results by voice;
[1853] A means for providing audio instructions for operating the viewing content and explanations of terms in real time;
[1854] A system including:
[1855] (Claim 2)
[1856] 10. The system of claim 1, operating as an application installed on a smartphone, smart glasses, a head-mounted display, or a robot.
[1857] (Claim 3)
[1858] 10. The system according to claim 1, wherein content playback, stopping, rewinding, and fast forwarding can be controlled by voice operation.
[1859] "Example 2: Combining Emotion Engines"
[1860] (Claim 1)
[1861] A voice input means;
[1862] a speech recognition means for converting speech data into text data;
[1863] an emotion recognition means for analyzing emotions from voice data;
[1864] a control means for analyzing the text data and the emotion data to manipulate the audio content;
[1865] A search tool that provides explanations of technical terms and new terms;
[1866] a speech synthesis means for converting the search results into speech data and outputting the speech;
[1867] A system including:
[1868] (Claim 2)
[1869] 10. The system of claim 1, further comprising the ability to resume playback from a previous playback position.
[1870] (Claim 3)
[1871] 10. The system of claim 1, wherein the reading speed and tone can be adjusted by voice control.
[1872] "Application example 2 when combining emotion engines"
[1873] Understood. We will extract the technically novel aspects from the application examples and rewrite the claims accordingly.
[1874] Revised Claims:
[1875] (Claim 1)
[1876] A voice input means;
[1877] a speech recognition means for converting speech data into text;
[1878] control means for analyzing the text data and emotional state to manipulate and adjust the audio content;
[1879] A search tool that provides explanations of technical terms and new terms;
[1880] a speech synthesis means for outputting the search results by voice;
[1881] an emotion recognition means for analyzing an emotional state;
[1882] A system including:
[1883] (Claim 2)
[1884] 10. The system of claim 1, further comprising the ability to resume playback from a previous playback position.
[1885] (Claim 3)
[1886] 10. The system of claim 1, wherein the speech rate and tone can be adjusted based on voice control and emotional state. [Explanation of symbols]
[1887] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. A voice input means; a speech recognition means for converting speech data into text; a control means for analyzing the text data and manipulating the audio content; A search tool that provides explanations of technical terms and new terms; a speech synthesis means for outputting the search results by voice; A system including:
2. 2. The system of claim 1, further comprising the ability to resume playback from a previous playback position.
3. 10. The system of claim 1, wherein the reading speed and tone are adjustable by voice control.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A