system
A system for converting text inputs from hearing-impaired individuals into audio data that reflects their intended message and emotions addresses the challenge of limited verbal communication, enabling effective real-time interaction.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-09-30
- Publication Date
- 2026-04-09
AI Technical Summary
There is a challenge in achieving smooth real-time communication between hearing-impaired individuals and those with normal hearing, as the means for hearing-impaired individuals to convey their intentions verbally are limited.
A system comprising an input means for receiving user input, a natural language processing means for analyzing text data and generating parameters for speech synthesis, a speech synthesis means for generating speech data, a transmission means for sending the speech data to the user's terminal, and a playback means for playing the speech data, enabling real-time communication.
Enables hearing-impaired individuals to communicate quickly and effectively with others by converting their text inputs into audio data that reflects their intended message and emotions, facilitating smoother interactions.
Smart Images

Figure 2026062206000001_ABST
Abstract
Description
Technical Field
[0001] The technology of this disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] There is a problem that it is difficult to achieve smooth communication between hearing-impaired people and healthy people. In particular, since the means for hearing-impaired people to convey their intentions verbally are limited, it is difficult to communicate in real time, which is an issue.
Means for Solving the Problems
[0005] The present invention solves the above problems by the following means: a system that includes an input means for receiving user input, a natural language processing means for analyzing text data received by the input means and generating parameters for speech synthesis, a speech synthesis means for generating speech data using the parameters generated by the natural language processing means, a transmission means for sending the speech data generated by the speech synthesis means to the user's terminal, and a playback means for playing the speech data on the user's terminal, thereby enabling real-time communication even for people with hearing impairments.
[0006] An "input method" is an interface for receiving text input from the user.
[0007] A "natural language processing means" is a means that analyzes text data input by an input means and generates the necessary parameters for speech synthesis.
[0008] "Speech synthesis means" refers to means for generating speech data using parameters generated by natural language processing means.
[0009] "Transmission means" refers to a communication means for sending the generated audio data to the user's terminal.
[0010] "Playback means" refers to means for playing back audio data received on the user's terminal.
[0011] "Adjustment means" refers to means for adjusting parameters such as language, pitch, and speed when generating audio data. [Brief explanation of the drawing]
[0012] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3]It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which a plurality of emotions are mapped. [Figure 10] It shows an emotion map to which a plurality of emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when an emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when an emotion engine is combined.
MODE FOR CARRYING OUT THE INVENTION
[0013] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0014] First, the language used in the following description will be explained.
[0015] In the following embodiments, the labeled processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0016] In the following embodiments, the labeled RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0017] In the following embodiments, the labeled storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, and the like.
[0018] In the following embodiments, the labeled communication I / F (Interface) is an interface including a communication processor and an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark), and the like.
[0019] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0020] [First Embodiment]
[0021] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0022] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0023] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0024] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0025] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0026] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0027] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0028] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0029] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0030] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0031] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0032] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0033] This system allows hearing-impaired individuals to input what they want to say as text, analyze that text to generate audio data, and then transmit it to hearing individuals. The configuration for implementing this system is described below.
[0034] System Overview
[0035] This system includes a terminal for user text input, a server that analyzes the input text and generates audio data, and a terminal that plays the generated audio data. The specific roles and processing flow of each component will be explained below.
[0036] User input and terminal operation
[0037] User: The user enters what they want to say in text format. For example, the user enters "Hello, how are you?" into the text box on their device.
[0038] Terminal: Provides a text input field, receives input from the user, and sends the content to the server when the submit button is pressed.
[0039] Server operation
[0040] Server: Receives text data sent from the terminal. Next, it analyzes the text using natural language processing and generates parameters for speech synthesis. For example, for the text "Hello, how are you?", it sets parameters such as language, pitch, and speed.
[0041] Natural language processing means: Analyzes the received text and derives appropriate speech synthesis parameters. Here, language models and sentiment analysis are used to determine parameters for generating more natural speech.
[0042] Speech synthesis means: Using parameters generated by natural language processing means, input text is converted into speech data. This speech data can then be converted into any desired audio format, such as a WAV file or MP3 file.
[0043] Transmission method: The generated audio data is sent back to the terminal. The terminal is configured to receive data via HTTP requests.
[0044] Device playback operation
[0045] Terminal: Receives audio data sent back from the server and plays it for the user. Uses software or devices with audio playback capabilities to transmit text input as audio.
[0046] Specific example
[0047] Input and Submission: The user enters "Hello, how are you?" into the text box and presses the submit button.
[0048] Text parsing: The server receives the text and converts "Hello, how are you?" into speech synthesis parameters.
[0049] Speech generation: A speech synthesis device converts the text "Hello, how are you?" into speech data.
[0050] Playback: The device receives and plays the audio data. The played audio contains the message, "Hello, how are you?"
[0051] This system enables people with hearing impairments to speak using their voices, facilitating smooth communication with hearing individuals. Details of each processing step are described below.
[0052] The following describes the processing flow.
[0053] Step 1: The user enters what they want to say in text format.
[0054] User: Enter "Hello, how are you?" into the text input field and click the send button.
[0055] Step 2: The terminal receives user input and sends the data to the server.
[0056] Terminal: Sends the entered text data "Hello, how are you?" to the server in POST request format.
[0057] Step 3: The server sends the received data to a natural language processing system for analysis.
[0058] Server: Passes the received text data to a natural language processing engine for analysis.
[0059] Step 4: The natural language processing engine analyzes the text data and generates speech synthesis parameters.
[0060] Server: The natural language processing engine analyzes the received text "Hello, how are you?" and generates parameters for speech synthesis (language, pitch, speed, etc.).
[0061] Step 5: Generate audio data using a speech synthesis method based on the analysis results.
[0062] Server: The speech synthesis engine generates speech data using parameters. The generated speech data contains the message, "Hello, how are you?"
[0063] Step 6: Send the generated audio data to the device.
[0064] Server: Sends the generated audio data to the terminal.
[0065] Step 7: The device plays the audio data received from the server.
[0066] Terminal: Passes the audio data received from the server to the playback function, which then plays the audio.
[0067] Step 8: The user communicates using the generated voice.
[0068] User: Uses the audio "Hello, how are you?" played from the device to communicate with a healthy person. During this process, the user confirms that their words have been transmitted as audio.
[0069] (Example 1)
[0070] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0071] In recent years, methods such as writing notes and using tablet devices have become common ways for people with hearing impairments to communicate smoothly with those without hearing impairments. However, these methods make real-time communication difficult, which is particularly inconvenient in emergencies or situations requiring a quick response. There is a need for a system that allows people with hearing impairments to communicate smoothly using voice.
[0072] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0073] In this invention, the server includes an input means for receiving user input, a natural language processing means for analyzing the text data received by the input means and generating parameters for speech synthesis, a speech synthesis means for generating speech data using the parameters generated by the natural language processing means, a transmission means for transmitting the speech data generated by the speech synthesis means to the user's terminal, and a playback means for playing the speech data on the user's terminal. This makes it possible for people with hearing impairments to quickly communicate their intentions through speech.
[0074] "An input means for receiving user input" refers to a device or system for users to input what they want to say in text format.
[0075] "Natural language processing means" refers to an algorithm or software that analyzes text data received by an input means and generates parameters for speech synthesis.
[0076] "Speech synthesis means" refers to an algorithm or software that converts input text into speech data using parameters generated by natural language processing means.
[0077] "Transmission means" refers to a device or system for transmitting audio data generated by speech synthesis means to a user's terminal.
[0078] "Playback means" refers to a device or software for playing audio data on the user's terminal.
[0079] "Adjustment means" refers to a device or software used to adjust parameters such as language, pitch, speed, and emotion when generating audio data.
[0080] A "terminal" is a device that allows a user to input text using an input device, send it to a server, and receive and play back the generated audio data.
[0081] This invention relates to a system for enabling hearing-impaired individuals to input what they wish to express as text, analyze that text to generate audio data, and communicate it to hearing individuals. The system includes a terminal for user text input, a server that analyzes the input text and generates audio data, and a terminal that plays the generated audio data.
[0082] System Overview
[0083] This system consists of the following elements:
[0084] 1. Input means for receiving user input
[0085] Users enter what they want to say in text format using an application on their smartphone or computer, or an input field on a webpage. For example, they might type, "Hello, how are you?"
[0086] 2. Transmission of text data via input means
[0087] Once the user completes the input and presses the submit button, the device sends this text data to the server. An HTTP POST request is used for submission.
[0088] 3. Text analysis by the server
[0089] The server receives text data sent from the terminal and analyzes it using natural language processing tools. This analysis uses language models such as Google's TENSORFLOW® and OpenAI's GPT-3®. The server analyzes the sentiment and grammar of the received text and generates optimal parameters for speech synthesis.
[0090] 4. Generating speech synthesis parameters
[0091] The server's natural language processing system generates speech synthesis parameters such as pitch, speed, and emotion based on the analysis results. For example, it extracts appropriate intonation and emotion from the text "Hello, how are you?".
[0092] 5. Server-driven generation of audio data
[0093] Based on parameters generated by natural language processing, the server uses speech synthesis to generate actual audio data. This speech synthesis utilizes services such as Amazon Polly or Google Cloud Text-to-Speech. The generated audio data is saved as WAV or MP3 files.
[0094] 6. Sending and playing audio data
[0095] The server sends the generated audio data to the terminal as an HTTP response. The terminal plays the received audio data using a playback device. Playback requires HTML5. <audio>Tags or dedicated audio playback applications are used.
[0096] Specific example
[0097] For example, the process would be as follows:
[0098] Input and Sending: The user types "Hello, how are you?" into the text box on the device and presses the send button.
[0099] Text analysis: The server receives text, analyzes it, and converts it into speech synthesis parameters. The natural language processing model used is OpenAI's GPT-3, among others.
[0100] Speech generation: The server uses Amazon Polly to convert the text "Hello, how are you?" into speech data.
[0101] Playback: The device receives audio data and HTML5 <audio>Play using tags.
[0102] Example of a prompt
[0103] "Please use GPT-3 to convert the text 'Hello, how are you?' into speech synthesis parameters. Then, use a speech synthesis service to generate the speech data and send it back to the terminal."
[0104] This system enables hearing-impaired individuals to communicate quickly through voice. The specific processing steps at each stage are carried out based on the means described in the claims.
[0105] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0106] Step 1:
[0107] User: The user enters what they want to say in text format using a smartphone or computer. For example, they might type, "Hello, how are you?"
[0108] Input: Text "Hello, how are you?"
[0109] Output: Text data entered into the user's terminal
[0110] Step 2:
[0111] Terminal: When the user presses the send button, the terminal sends the entered text data to the server. An HTTP POST request is used for transmission. Specifically, the text data is sent to the server in JSON format.
[0112] Input: Text data "Hello, how are you?"
[0113] Output: JSON data sent as an HTTP POST request
[0114] Step 3:
[0115] Server: The server receives an HTTP POST request and extracts text data. The received text data is then passed to a natural language processing system.
[0116] Input: JSON data sent to the server
[0117] Output: Text data passed to the natural language processing system.
[0118] Step 4:
[0119] Server: The natural language processing system analyzes text data and generates parameters for speech synthesis (e.g., language, pitch, speed, emotion, etc.). Examples of natural language processing models used include TensorFlow and GPT-3. For example, it performs emotion analysis on the text "Hello, how are you?" and generates appropriate intonation parameters.
[0120] Input: Text data "Hello, how are you?"
[0121] Output: Speech synthesis parameters (language, pitch, speed, emotion, etc.)
[0122] Step 5:
[0123] Server: Using the generated parameters, the speech synthesis system generates speech data. This system may use Amazon Polly or Google Cloud Text-to-Speech. For example, based on the parameters, speech data for "Hello, how are you?" is generated.
[0124] Input: Speech synthesis parameters
[0125] Output: Audio data (such as WAV or MP3 files)
[0126] Step 6:
[0127] Server: Sends the generated audio data to the terminal as an HTTP response. At this time, the appropriate MIME type (e.g., audio / mpeg) is set in the response header.
[0128] Input: Audio data
[0129] Output: HTTP response sent to the terminal
[0130] Step 7:
[0131] Terminal: The terminal analyzes the received audio data and plays it back using its audio playback function. Specifically, it uses HTML5 <audio>Tags or dedicated audio playback applications are used. For example, received audio data is played back to the user as "Hello, how are you?".
[0132] Input: Audio data received as an HTTP response
[0133] Output: Audio played to the user
[0134] This process allows hearing-impaired individuals to input what they want to say as text and transmit it to hearing individuals as audio.
[0135] (Application Example 1)
[0136] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0137] People with hearing impairments often have difficulty communicating with store staff in physical stores. Therefore, support is needed in situations where smooth communication is required, such as when asking for product information or completing a purchase. Furthermore, existing systems have limited capabilities for converting text data to audio data, making them unsuitable for use in physical stores.
[0138] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0139] In this invention, the server includes an input means for receiving user input, a natural language processing means for analyzing the text data received by the input means and generating parameters for speech synthesis, and a speech synthesis means for generating speech data using the parameters generated by the natural language processing means. This makes it possible to generate speech data and transmit it to other people in a physical store in order to support communication between people with hearing impairments and other people in the store.
[0140] A "user" is a person who uses this system to input text and convert that text into speech.
[0141] "Input means" refers to a device or function for a user to input text data. For example, this includes terminals such as smartphones and smart glasses.
[0142] "Natural language processing means" refers to a device or program that has the function of analyzing text data received by an input means and generating parameters for speech synthesis.
[0143] "Speech synthesis means" refers to a device or function that generates speech data using parameters generated by natural language processing means.
[0144] "Transmission means" refers to a device or function for transmitting generated audio data to the user's terminal.
[0145] "Playback means" refers to a device or function for playing audio data on a user terminal.
[0146] "A means of generating audio data and transmitting it to other people in a physical store to support communication between people with hearing impairments and other people in a physical store" refers to a means of converting text data entered by a person with hearing impairment into audio data and transmitting that audio to other people in a physical store.
[0147] This system is designed to facilitate communication between people with hearing impairments and other people in physical stores. A specific implementation of this system is described below.
[0148] System Configuration
[0149] User: A person with a hearing impairment performs text input.
[0150] Terminal: Provides a text input field and accepts text input from the user. This includes devices such as smartphones and smart glasses.
[0151] Server: Analyzes text data sent from the terminal, generates speech synthesis parameters, and creates speech data.
[0152] Natural language processing means: Analyzes text data and derives appropriate speech synthesis parameters.
[0153] Speech synthesis means: Generates speech data based on parameters.
[0154] Transmission method: The generated audio data is sent to the terminal.
[0155] Playback method: The audio data is played on the device and conveyed to other people, such as store clerks.
[0156] Hardware and software used
[0157] Hardware: Smartphones, smart glasses, servers.
[0158] Software: Python, Google Text-to-Speech (gTTS) library, Playsound library.
[0159] Processing flow and data manipulation / calculation
[0160] 1. The user enters text.
[0161] Enter text into the input field on your smartphone or smart glasses.
[0162] 2. The terminal sends text data to the server.
[0163] The text data is sent to the server in the format of an HTTP POST request.
[0164] 3. The server analyzes the text data and generates parameters.
[0165] The system analyzes text using natural language processing techniques and generates parameters such as language, pitch, and speed.
[0166] 4. The server generates the audio data.
[0167] Speech data is generated using parameters produced by a speech synthesis system.
[0168] 5. The server sends the audio data to the terminal.
[0169] The generated audio data is sent to the device.
[0170] 6. The device plays the audio data.
[0171] Audio data is played back using a playback device (smartphone or smart glasses) to convey the intentions of the hearing-impaired person to store staff, etc.
[0172] Specific example
[0173] For example, suppose a hearing-impaired person enters "Where is this product?" into smart glasses in a physical store. This text input is sent to a server, which analyzes the text and generates audio data saying "Where is this product?". This generated audio data is sent to the smart glasses and played back, allowing the store clerk to understand that the question is from a hearing-impaired person and guide them to the product's location.
[0174] Example of a prompt
[0175] Prompt: The user typed "Where can I find this product?". Convert this text into audio data and return the URL.
[0176] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0177] Step 1:
[0178] The user enters text.
[0179] Input: A person with a hearing impairment types "Where can I find this product?" into the text input field on their smartphone or smart glasses.
[0180] Action: The user types text into the input field and presses the submit button.
[0181] Output: The entered text data is recorded on the terminal.
[0182] Step 2:
[0183] The terminal sends text data to the server.
[0184] Input: Text data entered by the user.
[0185] Operation: The device sends text data to the server using an HTTP POST request.
[0186] Output: The server receives text data.
[0187] Step 3:
[0188] The server analyzes the text data and generates parameters.
[0189] Input: Text data received by the server: "Where is this product located?"
[0190] Operation: The server's natural language processing system analyzes the text and generates speech synthesis parameters such as language, pitch, and speed.
[0191] Output: Generated speech synthesis parameters.
[0192] Step 4:
[0193] The server generates the audio data.
[0194] Input: Speech synthesis parameters generated by the server.
[0195] Operation: The server's speech synthesis system generates speech data based on the parameters.
[0196] Output: Generated audio data (e.g., MP3 file).
[0197] Step 5:
[0198] The server sends the audio data to the terminal.
[0199] Input: Generated audio data.
[0200] Operation: The server's transmission method sends the generated audio data to the terminal in HTTP response format.
[0201] Output: The device receives the audio data.
[0202] Step 6:
[0203] The device plays audio data
[0204] Input: Received audio data.
[0205] Operation: The device's playback mechanism plays the received audio data. Audio is played from the smartphone or smart glasses, and the store clerk receives the question from the hearing-impaired person.
[0206] Output: The audio is played in the physical store, allowing the intentions of the hearing-impaired person to be conveyed to others.
[0207] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0208] This invention combines an emotion recognition function with a system that allows hearing-impaired individuals to input what they want to say as text, analyzes that text to generate audio data, and transmits it to hearing individuals. This allows the system to recognize the user's emotions from the input text data and reflect those emotions during speech synthesis.
[0209] System Overview
[0210] This system includes a terminal for user text input, a server that analyzes the input text and generates audio data, and a terminal that plays the generated audio data. In particular, it incorporates an emotion engine that recognizes the user's emotions from the text data.
[0211] User input and terminal operation
[0212] User: The user enters what they want to say in text format. For example, the user enters "Hello, how are you?" into the text box on their device.
[0213] Terminal: Provides a text input field, receives input from the user, and sends the content to the server when the submit button is pressed.
[0214] Server operation
[0215] Server: Receives text data sent from the terminal. Next, natural language processing means analyze the text and generate parameters for speech synthesis. In particular, emotion recognition means recognize the user's emotions from the text data and reflect that emotion information in the speech synthesis parameters.
[0216] Natural language processing means: Analyzes received text data and uses emotion recognition means to analyze the user's emotions. For example, it recognizes a positive emotion in response to the text "Hello, how are you?". Based on this, it sets parameters to make the voice sound brighter.
[0217] Emotion Engine: Recognizes user emotions from text data and generates emotion classification results. These results are sent to a speech synthesis system and used to generate speech data with natural-sounding expressions.
[0218] Speech synthesis means: Using parameters generated by emotion recognition means, input text is converted into speech data. Pitch, speed, intonation, etc., are added, adjusted based on emotion.
[0219] Transmission method: The generated audio data is sent back to the terminal. The terminal is configured to receive data via HTTP requests.
[0220] Device playback operation
[0221] Terminal: Receives audio data sent back from the server and plays it for the user. Using software or devices with audio playback capabilities, it transmits text input as audio. Because the audio data reflects the user's emotions, more natural communication is possible.
[0222] Specific example
[0223] Input and Submission: The user enters "Hello, how are you?" into the text box and presses the submit button.
[0224] Text analysis and sentiment recognition: The server receives text, analyzes it to see if it says "Hello, how are you?", and recognizes a positive sentiment.
[0225] Speech generation: The speech synthesis system sets parameters based on emotion recognition and converts the text "Hello, how are you?" into speech data in a cheerful tone.
[0226] Playback: The device receives and plays the audio data. The played audio conveys the message, "Hello, how are you?" with a positive tone.
[0227] This system enables hearing-impaired individuals to not only speak using their voices but also to convey their emotions, thereby achieving richer communication. Details of each processing step will be explained next.
[0228] The following describes the processing flow.
[0229] Step 1: The user enters what they want to say in text format.
[0230] User: Enter "Hello, how are you?" into the text input field and click the send button.
[0231] Step 2: The terminal receives user input and sends the data to the server.
[0232] Terminal: Sends the entered text data "Hello, how are you?" to the server in POST request format.
[0233] Step 3: The server sends the received data to a natural language processing system for analysis.
[0234] Server: Passes the received text data to a natural language processing engine for analysis.
[0235] Step 4: The natural language processing engine analyzes the text data and sends it to the sentiment recognition engine.
[0236] Server: The natural language processing engine analyzes the received text "Hello, how are you?" and passes the result to the emotion recognition engine.
[0237] Step 5: The emotion recognition engine extracts emotions from the text and generates emotional information.
[0238] Server: The emotion recognition engine extracts positive emotions from the text "Hello, how are you?" and generates emotion information.
[0239] Step 6: Set the speech synthesis parameters along with the emotional information.
[0240] Server: Based on the emotion information obtained from the emotion recognition engine, it sets the parameters used by the speech synthesis engine (e.g., pitch, speed, intonation).
[0241] Step 7: The speech synthesis engine generates speech data that reflects the emotion parameters.
[0242] Server: The speech synthesis engine converts the text "Hello, how are you?" into voice data with a cheerful tone, based on the configured parameters.
[0243] Step 8: Send the generated audio data to the device.
[0244] Server: Prepares to send the generated audio data to the terminal and returns it via an HTTP response.
[0245] Step 9: Play the audio data received by the device.
[0246] Terminal: Passes the audio data received from the server to the playback function, and plays the audio.
[0247] Step 10: The user communicates using the generated voice.
[0248] User: Uses a voice message containing positive emotions, such as "Hello, how are you?", played from the device to communicate with a healthy person.
[0249] This detailed processing step generates audio data that reflects not only the text content entered by the user, but also their emotions, thereby supporting more effective communication with people without disabilities.
[0250] (Example 2)
[0251] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0252] Conventional speech synthesis systems can convert text input by hearing-impaired individuals into speech data, but they cannot reflect the user's emotions, thus limiting their communication capabilities. In particular, the inability to convey the emotions behind speech makes accurate and rich communication difficult. Solving this problem is essential.
[0253] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0254] In this invention, the server includes an input means for receiving user input, a natural language processing means for analyzing the text data received by the input means, recognizing emotions, and generating parameters for speech synthesis, and a speech synthesis means for generating speech data using the parameters based on the emotions recognized by the natural language processing means. This makes it possible to recognize emotions from text data entered by a hearing-impaired person and generate speech data that reflects those emotions.
[0255] An "input means for receiving user input" is an interface for users to input what they want to say in text format, and a device that has the function of sending the entered text to the system.
[0256] "Natural language processing means" refers to technology for analyzing user input and understanding its content, and in particular, a device that has the function of recognizing emotions and generating parameters for speech synthesis.
[0257] An "emotion recognition device" is a device that analyzes the user's emotions from input text data and classifies those emotions into specific categories.
[0258] A "speech synthesis means" is a technology that converts input text into speech data using parameters generated by a natural language processing means, and is a device that has the function of reflecting emotion-based adjustments.
[0259] A "transmission means" is a device that has the function of transmitting the generated audio data to the user's terminal.
[0260] "Playback means" refers to a device or software for playing back audio data received on a user terminal.
[0261] A "modification device" is a device that has the function of adjusting parameters such as language, pitch, speed, and intonation based on emotion when generating audio data.
[0262] A "terminal" is a device used by a user to input and send text, and is a device that has the function of sending data to a server in the format of an HTTP request.
[0263] This invention combines an emotion recognition function with a system that allows hearing-impaired individuals to input what they want to say as text, analyzes that text to generate audio data, and transmits it to hearing individuals. This allows the system to recognize the user's emotions from the input text data and reflect those emotions during speech synthesis.
[0264] This system includes a terminal for user text input, a server that analyzes the input text and generates audio data, and a terminal that plays the generated audio data. In particular, it incorporates an emotion engine that recognizes the user's emotions from the text data.
[0265] User input and terminal operation
[0266] User:
[0267] The user enters what they want to say in text format. For example, the user might type "Hello, how are you?" into the text box on their device.
[0268] Terminal:
[0269] The terminal provides a text input field, receives input from the user, and sends the content to the server when the submit button is pressed. Specifically, it sends data to the server using an HTTP request.
[0270] Server operation
[0271] server:
[0272] The server receives text data sent from the terminal. The received text data is passed to a natural language processing system for analysis. Next, an emotion engine is used to recognize the user's emotions from the text data.
[0273] Natural language processing methods:
[0274] This technology analyzes received text data to understand the meaning of what the user is saying. For example, it analyzes and understands the meaning of the text, "Hello, how are you?"
[0275] Emotional engine:
[0276] Based on the analysis results from the natural language processing means, recognize the user's emotion. For example, recognize a positive emotion from the text "Hello, how are you?".
[0277] Speech synthesis means:
[0278] Using the parameters generated by the emotion recognition means, convert the input text into voice data. The pitch, speed, intonation, etc. are adjusted. For example, generate voice data in a bright tone that reflects a positive emotion.
[0279] Transmission means:
[0280] Return the voice data generated by the speech synthesis means to the terminal. Transmit the voice data using an HTTP response.
[0281] Playback operation of the terminal
[0282] Terminal:
[0283] The terminal receives the voice data returned from the server and uses the voice playback function to play the voice data to the user. As a result, the content of the input text is played as voice reflecting the emotion. Specifically, transmit the content "Hello, how are you?" with a positive emotion.
[0284] Specific example
[0285] Input and transmission:
[0286] The user enters "Hello, how are you?" in the text box and presses the send button.
[0287] Analysis of text and emotion recognition:
[0288] The server receives and analyzes "Hello, how are you?" and recognizes a positive emotion.
[0289] Voice generation:
[0290] Server: The speech synthesis system sets parameters to match positive emotions and converts "Hello, how are you?" into voice data in a cheerful tone.
[0291] reproduction:
[0292] Terminal: Receives audio data and plays it back using its audio playback function. The played audio conveys the message, "Hello, how are you?" along with a positive tone.
[0293] Concrete examples of prompt sentences for generative AI models
[0294] Prompt example:
[0295] "Analyze the following text data to recognize the user's emotions. Generate emotion-based speech synthesis parameters and then generate the speech data."
[0296] This system allows emotions to be reflected in the text entered by people with hearing impairments, enabling more natural and enriching communication.
[0297] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0298] Step 1:
[0299] User:
[0300] Inputs and outputs:
[0301] Input: The user enters "Hello, how are you?" into the text input field on the device.
[0302] Output: The entered text "Hello, how are you?" is stored in the terminal.
[0303] Specific actions:
[0304] The user enters the content they want to say in the text input field of the terminal and presses the send button.
[0305] Step 2:
[0306] Terminal:
[0307] Input and Output:
[0308] Input: The text "Hello, how are you?" entered in Step 1
[0309] Output: The text data is sent to the server as an HTTP request.
[0310] Specific Operations:
[0311] The terminal retrieves the content of the text input field and sends it to the server using an HTTP POST request.
[0312] Step 3:
[0313] Server:
[0314] Input and Output:
[0315] Input: The text data "Hello, how are you?" sent from the terminal
[0316] Output: The text data is received by the server.
[0317] Specific Operations:
[0318] The server receives the POST request sent from the terminal and extracts the text data from the request body.
[0319] Step 4:
[0320] Server:
[0321] Input and Output:
[0322] Input: Text data received in Step 3: "Hello, how are you?"
[0323] Output: Analysis results of text data (meaning and content of the statement)
[0324] Specific actions:
[0325] The server passes text data to a natural language processing system, which then analyzes the text and understands its meaning.
[0326] Step 5:
[0327] server:
[0328] Inputs and outputs:
[0329] Input: Meaning of the text data "Hello, how are you?" analyzed in Step 4
[0330] Output: Emotion recognition result (e.g., positive)
[0331] Specific actions:
[0332] The server uses an emotion engine to recognize the user's emotions from the analysis results. For example, it recognizes positive emotions from "Hello, how are you?".
[0333] Step 6:
[0334] server:
[0335] Inputs and outputs:
[0336] Input: Emotions recognized in Step 5 and text data analyzed in Step 4
[0337] Output: Speech synthesis parameters (e.g., bright tone, appropriate speed and intonation)
[0338] Specific actions:
[0339] The server sets parameters for generating speech data using speech synthesis based on the emotion recognition results. For example, it sets parameters for a bright tone based on positive emotions.
[0340] Step 7:
[0341] server:
[0342] Inputs and outputs:
[0343] Input: Speech synthesis parameters generated in step 6
[0344] Output: Audio data (Example: "Hello, how are you?" in a cheerful tone)
[0345] Specific actions:
[0346] The server uses speech synthesis to convert the input text into speech data based on the configured parameters.
[0347] Step 8:
[0348] server:
[0349] Inputs and outputs:
[0350] Input: Audio data generated in Step 7
[0351] Output: The audio data is sent to the terminal as an HTTP response.
[0352] Specific actions:
[0353] The server sends the generated audio data back to the terminal as an HTTP response.
[0354] Step 9:
[0355] Terminal:
[0356] Inputs and outputs:
[0357] Input: Audio data sent from the server
[0358] Output: The voice played by the user (e.g., "Hello, how are you?" in a cheerful tone)
[0359] Specific actions:
[0360] The device receives audio data sent from the server and plays it back using its audio playback function. Because the played audio reflects the user's emotions, more natural communication is possible.
[0361] (Application Example 2)
[0362] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0363] Traditionally, there has been a lack of adequate means for hearing-impaired individuals to quickly generate appropriate voice messages, including emotions, in emergencies and notify security guards and emergency services. This has led to delays in emergency communication and the potential for misunderstandings. This invention aims to solve these problems by recognizing emotions based on text input and generating voice data that includes those emotions to notify emergency services.
[0364] In Application Example 2, the specific processing performed by the specific processing unit 290 of the data processing device 12 is realized by the following means. In this invention, the server includes an input means for receiving user input, a natural language processing means for analyzing the received text data and generating parameters for speech synthesis, a speech synthesis means for generating speech data using the parameters generated by the natural language processing means, an emotion recognition means for recognizing emotions from the text data using the natural language processing means and reflecting that emotion information in the speech synthesis parameters, a transmission means for transmitting the generated speech data to a user terminal, a playback means for playing the speech data on the user terminal, and a notification means for analyzing the input text data in an emergency, generating a speech message including emotions, and notifying security guards or emergency services. This makes it possible for hearing-impaired individuals to quickly and accurately generate a speech message including emotions in an emergency and to respond appropriately to the emergency.
[0365] "Input means" refers to a device or interface for a user to input text data.
[0366] "Natural language processing means" refers to systems and algorithms that analyze received text data and generate parameters for speech synthesis.
[0367] "Speech synthesis means" refers to a device or software that generates speech data using parameters generated by natural language processing means.
[0368] "Emotion recognition means" refers to a system or algorithm that recognizes a user's emotions from text data and reflects that emotional information in speech synthesis parameters.
[0369] "Transmission means" refers to the communication means used to send the generated audio data to the user's terminal.
[0370] "Playback means" refers to a device or software that plays back audio data transmitted to the user's terminal.
[0371] "Notification means" refers to a mechanism or system that analyzes text data entered in an emergency, generates an audio message including emotions, and notifies security guards or emergency services.
[0372] The system according to the present invention is a system for enabling hearing-impaired individuals to quickly and accurately generate emotionally charged voice messages in emergencies and notify security guards and emergency services. This system mainly consists of an input means, a natural language processing means, a speech synthesis means, an emotion recognition means, a transmission means, a playback means, and a notification means.
[0373] Hardware and software configuration
[0374] Hardware used: Smartphone or robot (with voice output function)
[0375] Software used:
[0376] text2emotion: Emotion Recognition Library
[0377] gTTS (Google Text-to-Speech): A speech synthesis library.
[0378] requests: A library for sending HTTP requests.
[0379] Operational description of each means
[0380] 1. Input method:
[0381] This provides an interface for users to enter text data in emergencies. This interface will be implemented as a text input field on smartphones or robots. Users will enter emergency messages here.
[0382] 2. Natural language processing methods:
[0383] The system analyzes text data received through the input means and generates parameters for speech synthesis. This step is crucial as it involves understanding the content of the text data and converting it into speech data.
[0384] 3. Speech synthesis means:
[0385] Speech data is generated using parameters produced by natural language processing. During this process, pitch, speed, intonation, and other parameters are adjusted to produce natural-sounding speech that also conveys emotion.
[0386] 4. Emotion recognition means:
[0387] The system analyzes text data received from a natural language processing system to recognize the user's emotions. This recognized emotion information is then reflected in the parameters of the speech synthesis system. The text2emotion library is used for this emotion recognition.
[0388] 5. Transmission method:
[0389] This is a communication method for sending generated audio data to the user's device. This method uses the requests library to send audio data in HTTP request format.
[0390] 6. Regeneration means:
[0391] This is a device or software that plays back audio data transmitted to a user's terminal. It provides audio output on smartphones and robots.
[0392] 7. Means of notification:
[0393] In emergencies, the system analyzes entered text data, generates an emotionally charged voice message, and immediately notifies security personnel or emergency services. This notification is sent to a server by transmitting the generated voice data.
[0394] Specific example
[0395] User actions:
[0396] The user types "Help me, there's a fire" into the text input field of their smartphone or robot. Then they press the send button.
[0397] Server operation:
[0398] The server analyzes the received text data using natural language processing and recognizes the emotion of fear from the content of "Fire!". As a result, it generates audio data that includes the emotion of urgency and fear using speech synthesis.
[0399] Speech synthesis:
[0400] The speech synthesis means uses speech parameters that reflect the emotion of fear recognized by the emotion recognition means to generate an urgent voice message saying, "Help me, there's a fire."
[0401] Emergency notification:
[0402] The generated audio data is transmitted to security guards and emergency services via a transmission device to facilitate a rapid response.
[0403] Example of a prompt
[0404] Towards emotion recognition AI:
[0405] Text: "Help me, there's a fire."
[0406] Please return the emotion recognition results in JSON format (e.g., {"happy": 0.1, "sad": 0.3, "angry": 0.1, "fear": 0.5}).
[0407] Towards speech synthesis AI:
[0408] Text: "Help me, there's a fire."
[0409] Recognized emotion: fear
[0410] Please generate natural-sounding voice data that reflects emotions.
[0411] In this way, the present invention enables communication that includes the emotions of hearing-impaired individuals in emergencies, and facilitates a rapid and appropriate emergency response.
[0412] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0413] Step 1:
[0414] The user enters an emergency message into a text input field on their smartphone or robot and presses the send button. The entered text data contains emergency information, such as "Help, there's a fire." This input is then sent to the server via the input device.
[0415] Input: User-entered text data "Help, there's a fire"
[0416] Output: Text data sent to the server
[0417] Step 2:
[0418] The server passes the received text data to a natural language processing system, which analyzes the text's content. The natural language processing system understands the meaning from the input text and generates basic parameters for speech synthesis.
[0419] Input: Received text data
[0420] Data Processing / Calculation: Analyze text content using natural language processing techniques and generate speech synthesis parameters.
[0421] Output: Generated speech synthesis parameters
[0422] Step 3:
[0423] The natural language processing system passes text data to the emotion recognition system to recognize the user's emotions. Here, the text2emotion library is used to analyze the emotions implied in the text and identify the main emotions.
[0424] Input: Text data passed from a natural language processing system.
[0425] Data processing / calculation: Recognizing emotions using the text2emotion library.
[0426] Output: Recognized emotion information (Example: {"fear": 0.5, "sad": 0.3, "happy": 0.1, "angry": 0.1})
[0427] Step 4:
[0428] Based on the emotion information recognized by the emotion recognition means, the speech synthesis means generates speech data. Using the gTTS library, speech data with speech parameters that reflect the emotion is created.
[0429] Input: Sentimental information and speech synthesis parameters
[0430] Data processing / calculation: Generate audio data using the gTTS library.
[0431] Output: Generated audio data file
[0432] Step 5:
[0433] The generated audio data file is sent via HTTP request to the security guard or emergency service system, which is the recipient of the emergency notification. The requests library is used.
[0434] Input: Generated audio data file
[0435] Data processing / calculation: Data transmission via HTTP requests
[0436] Output: Completion of voice data transmission to security guards and emergency services.
[0437] Step 6:
[0438] Security and emergency service systems play back transmitted audio data and initiate a rapid response. The playback mechanism uses an audio output device to play back emergency messages that reflect emotions.
[0439] Input: Received audio data file
[0440] Data processing / calculation: Audio data playback
[0441] Output: Initiate emergency response and play voice message.
[0442] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0443] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0444] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0445] [Second Embodiment]
[0446] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0447] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0448] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0449] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0450] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0451] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0452] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0453] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0454] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0455] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0456] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0457] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0458] This system allows hearing-impaired individuals to input what they want to say as text, analyze that text to generate audio data, and transmit it to hearing individuals. The configuration for implementing this system is described below.
[0459] System Overview
[0460] This system includes a terminal for user text input, a server that analyzes the input text and generates audio data, and a terminal that plays the generated audio data. The specific roles and processing flow of each component will be explained below.
[0461] User input and terminal operation
[0462] User: The user enters what they want to say in text format. For example, the user enters "Hello, how are you?" into the text box on their device.
[0463] Terminal: Provides a text input field, receives input from the user, and sends the content to the server when the submit button is pressed.
[0464] Server operation
[0465] Server: Receives text data sent from the terminal. Next, it analyzes the text using natural language processing and generates parameters for speech synthesis. For example, for the text "Hello, how are you?", it sets parameters such as language, pitch, and speed.
[0466] Natural language processing means: Analyzes the received text and derives appropriate speech synthesis parameters. Here, language models and sentiment analysis are used to determine parameters for generating more natural speech.
[0467] Speech synthesis means: Using parameters generated by natural language processing means, input text is converted into speech data. This speech data can then be converted into any desired audio format, such as a WAV file or MP3 file.
[0468] Transmission method: The generated audio data is sent back to the terminal. The terminal is configured to receive data via HTTP requests.
[0469] Device playback operation
[0470] Terminal: Receives audio data sent back from the server and plays it for the user. Uses software or devices with audio playback capabilities to transmit text input as audio.
[0471] Specific example
[0472] Input and Submission: The user enters "Hello, how are you?" into the text box and presses the submit button.
[0473] Text parsing: The server receives the text and converts "Hello, how are you?" into speech synthesis parameters.
[0474] Speech generation: A speech synthesis device converts the text "Hello, how are you?" into speech data.
[0475] Playback: The device receives and plays the audio data. The played audio contains the message, "Hello, how are you?"
[0476] This system enables hearing-impaired individuals to speak using voice, facilitating smooth communication with hearing individuals. Details of each processing step are described below.
[0477] The following describes the processing flow.
[0478] Step 1: The user enters what they want to say in text format.
[0479] User: Enter "Hello, how are you?" into the text input field and click the send button.
[0480] Step 2: The terminal receives user input and sends the data to the server.
[0481] Terminal: Sends the entered text data "Hello, how are you?" to the server in POST request format.
[0482] Step 3: The server sends the received data to a natural language processing system for analysis.
[0483] Server: Passes the received text data to a natural language processing engine for analysis.
[0484] Step 4: The natural language processing engine analyzes the text data and generates speech synthesis parameters.
[0485] Server: The natural language processing engine analyzes the received text "Hello, how are you?" and generates parameters for speech synthesis (language, pitch, speed, etc.).
[0486] Step 5: Generate audio data using a speech synthesis method based on the analysis results.
[0487] Server: The speech synthesis engine generates speech data using parameters. The generated speech data contains the message, "Hello, how are you?"
[0488] Step 6: Send the generated audio data to the device.
[0489] Server: Sends the generated audio data to the terminal.
[0490] Step 7: The device plays the audio data received from the server.
[0491] Terminal: Passes the audio data received from the server to the playback function, which then plays the audio.
[0492] Step 8: The user communicates using the generated voice.
[0493] User: Uses the audio "Hello, how are you?" played from the device to communicate with a healthy person via voice. During this process, the user confirms that their words have been transmitted as audio.
[0494] (Example 1)
[0495] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0496] In recent years, methods such as writing notes and using tablet devices have become common ways for people with hearing impairments to communicate smoothly with those without hearing impairments. However, these methods make real-time communication difficult, which is particularly inconvenient in emergencies or situations requiring a quick response. There is a need for a system that allows people with hearing impairments to communicate smoothly using voice.
[0497] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0498] In this invention, the server includes an input means for receiving user input, a natural language processing means for analyzing the text data received by the input means and generating parameters for speech synthesis, a speech synthesis means for generating speech data using the parameters generated by the natural language processing means, a transmission means for transmitting the speech data generated by the speech synthesis means to the user's terminal, and a playback means for playing the speech data on the user's terminal. This makes it possible for people with hearing impairments to quickly communicate their intentions through speech.
[0499] "An input means for receiving user input" refers to a device or system for users to input what they want to say in text format.
[0500] "Natural language processing means" refers to an algorithm or software that analyzes text data received by an input means and generates parameters for speech synthesis.
[0501] "Speech synthesis means" refers to an algorithm or software that converts input text into speech data using parameters generated by natural language processing means.
[0502] "Transmission means" refers to a device or system for transmitting audio data generated by speech synthesis means to a user's terminal.
[0503] "Playback means" refers to a device or software for playing audio data on the user's terminal.
[0504] "Adjustment means" refers to a device or software used to adjust parameters such as language, pitch, speed, and emotion when generating audio data.
[0505] A "terminal" is a device that allows a user to input text using an input device, send it to a server, and receive and play back the generated audio data.
[0506] This invention relates to a system for enabling hearing-impaired individuals to input what they wish to express as text, analyze that text to generate audio data, and communicate it to hearing individuals. The system includes a terminal for user text input, a server that analyzes the input text and generates audio data, and a terminal that plays the generated audio data.
[0507] System Overview
[0508] This system consists of the following elements:
[0509] 1. Input means for receiving user input
[0510] Users enter what they want to say in text format using an application on their smartphone or computer, or an input field on a webpage. For example, they might type, "Hello, how are you?"
[0511] 2. Transmission of text data via input means
[0512] Once the user completes the input and presses the submit button, the device sends this text data to the server. An HTTP POST request is used for submission.
[0513] 3. Text analysis by the server
[0514] The server receives text data sent from the terminal and analyzes it using natural language processing tools. This analysis uses language models such as Google's TensorFlow and OpenAI's GPT-3. It analyzes the sentiment and grammar of the received text and generates optimal parameters for speech synthesis.
[0515] 4. Generating speech synthesis parameters
[0516] The server's natural language processing system generates speech synthesis parameters such as pitch, speed, and emotion based on the analysis results. For example, it extracts appropriate intonation and emotion from the text "Hello, how are you?".
[0517] 5. Server-driven generation of audio data
[0518] Based on parameters generated by natural language processing, the server uses speech synthesis to generate actual audio data. This speech synthesis utilizes services such as Amazon Polly or Google Cloud Text-to-Speech. The generated audio data is saved as WAV or MP3 files.
[0519] 6. Sending and playing audio data
[0520] The server sends the generated audio data to the terminal as an HTTP response. The terminal plays the received audio data using a playback device. Playback requires HTML5. <audio>Tags or dedicated audio playback applications are used.
[0521] Specific example
[0522] For example, the process would be as follows:
[0523] Input and Sending: The user types "Hello, how are you?" into the text box on the device and presses the send button.
[0524] Text analysis: The server receives text, analyzes it, and converts it into speech synthesis parameters. The natural language processing model used is OpenAI's GPT-3, among others.
[0525] Speech generation: The server uses Amazon Polly to convert the text "Hello, how are you?" into speech data.
[0526] Playback: The device receives audio data and HTML5 <audio>Play using tags.
[0527] Example of a prompt
[0528] "Please use GPT-3 to convert the text 'Hello, how are you?' into speech synthesis parameters. Then, use a speech synthesis service to generate the speech data and send it back to the terminal."
[0529] This system enables hearing-impaired individuals to communicate quickly through voice. The specific processing steps at each stage are carried out based on the means described in the claims.
[0530] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0531] Step 1:
[0532] User: The user enters what they want to say in text format using a smartphone or computer. For example, they might type, "Hello, how are you?"
[0533] Input: Text "Hello, how are you?"
[0534] Output: Text data entered into the user's terminal
[0535] Step 2:
[0536] Terminal: When the user presses the send button, the terminal sends the entered text data to the server. An HTTP POST request is used for transmission. Specifically, the text data is sent to the server in JSON format.
[0537] Input: Text data "Hello, how are you?"
[0538] Output: JSON data sent as an HTTP POST request
[0539] Step 3:
[0540] Server: The server receives an HTTP POST request and extracts text data. The received text data is then passed to a natural language processing system.
[0541] Input: JSON data sent to the server
[0542] Output: Text data passed to the natural language processing system.
[0543] Step 4:
[0544] Server: The natural language processing system analyzes text data and generates parameters for speech synthesis (e.g., language, pitch, speed, emotion, etc.). Examples of natural language processing models used include TensorFlow and GPT-3. For example, it performs emotion analysis on the text "Hello, how are you?" and generates appropriate intonation parameters.
[0545] Input: Text data "Hello, how are you?"
[0546] Output: Speech synthesis parameters (language, pitch, speed, emotion, etc.)
[0547] Step 5:
[0548] Server: Using the generated parameters, the speech synthesis system generates speech data. This system may use Amazon Polly or Google Cloud Text-to-Speech. For example, based on the parameters, speech data for "Hello, how are you?" is generated.
[0549] Input: Speech synthesis parameters
[0550] Output: Audio data (such as WAV or MP3 files)
[0551] Step 6:
[0552] Server: Sends the generated audio data to the terminal as an HTTP response. At this time, the appropriate MIME type (e.g., audio / mpeg) is set in the response header.
[0553] Input: Audio data
[0554] Output: HTTP response sent to the terminal
[0555] Step 7:
[0556] Terminal: The terminal analyzes the received audio data and plays it back using its audio playback function. Specifically, it uses HTML5 <audio>Tags or dedicated audio playback applications are used. For example, received audio data is played back to the user as "Hello, how are you?".
[0557] Input: Audio data received as an HTTP response
[0558] Output: Audio played to the user
[0559] This process allows hearing-impaired individuals to input what they want to say as text and transmit it to hearing individuals as audio.
[0560] (Application Example 1)
[0561] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0562] People with hearing impairments often have difficulty communicating with store staff in physical stores. Therefore, support is needed in situations where smooth communication is required, such as when asking for product information or completing a purchase. Furthermore, existing systems have limited capabilities for converting text data to audio data, making them unsuitable for use in physical stores.
[0563] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0564] In this invention, the server includes an input means for receiving user input, a natural language processing means for analyzing the text data received by the input means and generating parameters for speech synthesis, and a speech synthesis means for generating speech data using the parameters generated by the natural language processing means. This makes it possible to generate speech data and transmit it to other people in a physical store in order to support communication between people with hearing impairments and other people in the store.
[0565] A "user" is a person who uses this system to input text and convert that text into speech.
[0566] "Input means" refers to a device or function for a user to input text data. For example, this includes terminals such as smartphones and smart glasses.
[0567] "Natural language processing means" refers to a device or program that has the function of analyzing text data received by an input means and generating parameters for speech synthesis.
[0568] "Speech synthesis means" refers to a device or function that generates speech data using parameters generated by natural language processing means.
[0569] "Transmission means" refers to a device or function for transmitting generated audio data to the user's terminal.
[0570] "Playback means" refers to a device or function for playing audio data on a user terminal.
[0571] "A means of generating audio data and transmitting it to other people in a physical store to support communication between people with hearing impairments and other people in a physical store" refers to a means of converting text data entered by a person with hearing impairment into audio data and transmitting that audio to other people in a physical store.
[0572] This system is designed to facilitate communication between people with hearing impairments and other people in physical stores. A specific implementation of this system is described below.
[0573] System Configuration
[0574] User: A person with a hearing impairment performs text input.
[0575] Terminal: Provides a text input field and accepts text input from the user. This includes devices such as smartphones and smart glasses.
[0576] Server: Analyzes text data sent from the terminal, generates speech synthesis parameters, and creates speech data.
[0577] Natural language processing means: Analyzes text data and derives appropriate speech synthesis parameters.
[0578] Speech synthesis means: Generates speech data based on parameters.
[0579] Transmission method: The generated audio data is sent to the terminal.
[0580] Playback method: The audio data is played on the device and conveyed to other people, such as store clerks.
[0581] Hardware and software used
[0582] Hardware: Smartphones, smart glasses, servers.
[0583] Software: Python, Google Text-to-Speech (gTTS) library, Playsound library.
[0584] Processing flow and data manipulation / calculation
[0585] 1. The user enters text.
[0586] Enter text into the input field on your smartphone or smart glasses.
[0587] 2. The terminal sends text data to the server.
[0588] The text data is sent to the server in the format of an HTTP POST request.
[0589] 3. The server analyzes the text data and generates parameters.
[0590] The system analyzes text using natural language processing techniques and generates parameters such as language, pitch, and speed.
[0591] 4. The server generates the audio data.
[0592] Speech data is generated using parameters produced by a speech synthesis system.
[0593] 5. The server sends the audio data to the terminal.
[0594] The generated audio data is sent to the device.
[0595] 6. The device plays the audio data.
[0596] Audio data is played back using a playback device (smartphone or smart glasses) to convey the intentions of the hearing-impaired person to store staff, etc.
[0597] Specific example
[0598] For example, suppose a person with hearing impairment enters the question "Where is this product?" into smart glasses in a physical store. This text input is sent to a server, which analyzes the text and generates audio data saying "Where is this product?". This generated audio data is sent to the smart glasses and played back, allowing the store clerk to understand that the question is from someone with hearing impairment and guide them to the product's location.
[0599] Example of a prompt
[0600] Prompt: The user typed "Where can I find this product?". Convert this text into audio data and return the URL.
[0601] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0602] Step 1:
[0603] The user enters text.
[0604] Input: A person with a hearing impairment types "Where can I find this product?" into the text input field on their smartphone or smart glasses.
[0605] Action: The user types text into the input field and presses the submit button.
[0606] Output: The entered text data is recorded on the terminal.
[0607] Step 2:
[0608] The terminal sends text data to the server.
[0609] Input: Text data entered by the user.
[0610] Operation: The device sends text data to the server using an HTTP POST request.
[0611] Output: The server receives text data.
[0612] Step 3:
[0613] The server analyzes the text data and generates parameters.
[0614] Input: Text data received by the server: "Where is this product located?"
[0615] Operation: The server's natural language processing system analyzes the text and generates speech synthesis parameters such as language, pitch, and speed.
[0616] Output: Generated speech synthesis parameters.
[0617] Step 4:
[0618] The server generates the audio data.
[0619] Input: Speech synthesis parameters generated by the server.
[0620] Operation: The server's speech synthesis system generates speech data based on the parameters.
[0621] Output: Generated audio data (e.g., MP3 file).
[0622] Step 5:
[0623] The server sends the audio data to the terminal.
[0624] Input: Generated audio data.
[0625] Operation: The server's transmission method sends the generated audio data to the terminal in HTTP response format.
[0626] Output: The device receives the audio data.
[0627] Step 6:
[0628] The device plays audio data
[0629] Input: Received audio data.
[0630] Operation: The device's playback mechanism plays the received audio data. Audio is played from the smartphone or smart glasses, and the store clerk receives the question from the hearing-impaired person.
[0631] Output: The audio is played in the physical store, allowing the intentions of the hearing-impaired person to be conveyed to others.
[0632] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0633] This invention combines an emotion recognition function with a system that allows hearing-impaired individuals to input what they want to say as text, analyzes that text to generate audio data, and transmits it to hearing individuals. This allows the system to recognize the user's emotions from the input text data and reflect those emotions during speech synthesis.
[0634] System Overview
[0635] This system includes a terminal for user text input, a server that analyzes the input text and generates audio data, and a terminal that plays the generated audio data. In particular, it incorporates an emotion engine that recognizes the user's emotions from the text data.
[0636] User input and terminal operation
[0637] User: The user enters what they want to say in text format. For example, the user enters "Hello, how are you?" into the text box on their device.
[0638] Terminal: Provides a text input field, receives input from the user, and sends the content to the server when the submit button is pressed.
[0639] Server operation
[0640] Server: Receives text data sent from the terminal. Next, natural language processing means analyze the text and generate parameters for speech synthesis. In particular, emotion recognition means recognize the user's emotions from the text data and reflect that emotion information in the speech synthesis parameters.
[0641] Natural language processing means: Analyzes received text data and uses emotion recognition means to analyze the user's emotions. For example, it recognizes a positive emotion in response to the text "Hello, how are you?". Based on this, it sets parameters to make the voice sound brighter.
[0642] Emotion Engine: Recognizes user emotions from text data and generates emotion classification results. These results are sent to a speech synthesis system and used to generate speech data with natural-sounding expressions.
[0643] Speech synthesis means: Using parameters generated by emotion recognition means, it converts input text into speech data. It adds pitch, speed, intonation, etc., adjusted based on emotion.
[0644] Transmission method: The generated audio data is sent back to the terminal. The terminal is configured to receive data via HTTP requests.
[0645] Device playback operation
[0646] Terminal: Receives audio data sent back from the server and plays it for the user. Using software or devices with audio playback capabilities, it transmits text input as audio. Because the audio data reflects the user's emotions, more natural communication is possible.
[0647] Specific example
[0648] Input and Submission: The user enters "Hello, how are you?" into the text box and presses the submit button.
[0649] Text analysis and sentiment recognition: The server receives text, analyzes it to see if it says "Hello, how are you?", and recognizes a positive sentiment.
[0650] Speech generation: The speech synthesis system sets parameters based on emotion recognition and converts the text "Hello, how are you?" into speech data in a cheerful tone.
[0651] Playback: The device receives and plays the audio data. The played audio conveys the message, "Hello, how are you?" with a positive tone.
[0652] This system enables hearing-impaired individuals to not only speak using their voices but also to convey their emotions, thereby achieving richer communication. Details of each processing step will be explained next.
[0653] The following describes the processing flow.
[0654] Step 1: The user enters what they want to say in text format.
[0655] User: Enter "Hello, how are you?" into the text input field and click the send button.
[0656] Step 2: The terminal receives user input and sends the data to the server.
[0657] Terminal: Sends the entered text data "Hello, how are you?" to the server in POST request format.
[0658] Step 3: The server sends the received data to a natural language processing system for analysis.
[0659] Server: Passes the received text data to a natural language processing engine for analysis.
[0660] Step 4: The natural language processing engine analyzes the text data and sends it to the sentiment recognition engine.
[0661] Server: The natural language processing engine analyzes the received text "Hello, how are you?" and passes the result to the emotion recognition engine.
[0662] Step 5: The emotion recognition engine extracts emotions from the text and generates emotional information.
[0663] Server: The emotion recognition engine extracts positive emotions from the text "Hello, how are you?" and generates emotion information.
[0664] Step 6: Set the speech synthesis parameters along with the emotional information.
[0665] Server: Based on the emotion information obtained from the emotion recognition engine, it sets the parameters used by the speech synthesis engine (e.g., pitch, speed, intonation).
[0666] Step 7: The speech synthesis engine generates speech data that reflects the emotion parameters.
[0667] Server: The speech synthesis engine converts the text "Hello, how are you?" into voice data with a cheerful tone, based on the configured parameters.
[0668] Step 8: Send the generated audio data to the device.
[0669] Server: Prepares to send the generated audio data to the terminal and returns it via an HTTP response.
[0670] Step 9: Play the audio data received by the device.
[0671] Terminal: Passes the audio data received from the server to the playback function, and plays the audio.
[0672] Step 10: The user communicates using the generated voice.
[0673] User: Uses a voice message containing positive emotions, such as "Hello, how are you?", played from the device to communicate with a healthy person.
[0674] This detailed processing step allows for the generation of audio data that reflects not only the text content entered by the user, but also their emotions, thereby supporting more effective communication with individuals without cognitive impairment.
[0675] (Example 2)
[0676] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0677] Conventional speech synthesis systems can convert text input by hearing-impaired individuals into speech data, but they cannot reflect the user's emotions, thus limiting their communication capabilities. In particular, the inability to convey the emotions behind speech makes accurate and rich communication difficult. Solving this problem is essential.
[0678] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0679] In this invention, the server includes an input means for receiving user input, a natural language processing means for analyzing the text data received by the input means, recognizing emotions, and generating parameters for speech synthesis, and a speech synthesis means for generating speech data using the parameters based on the emotions recognized by the natural language processing means. This makes it possible to recognize emotions from text data entered by a hearing-impaired person and generate speech data that reflects those emotions.
[0680] An "input means for receiving user input" is an interface for users to input what they want to say in text format, and a device that has the function of sending the entered text to the system.
[0681] "Natural language processing means" refers to technology for analyzing user input and understanding its content, and in particular, a device that has the function of recognizing emotions and generating parameters for speech synthesis.
[0682] An "emotion recognition device" is a device that analyzes the user's emotions from input text data and classifies those emotions into specific categories.
[0683] A "speech synthesis means" is a technology that converts input text into speech data using parameters generated by a natural language processing means, and is a device that has the function of reflecting emotion-based adjustments.
[0684] A "transmission means" is a device that has the function of transmitting the generated audio data to the user's terminal.
[0685] "Playback means" refers to a device or software for playing back audio data received on a user terminal.
[0686] A "modification device" is a device that has the function of adjusting parameters such as language, pitch, speed, and intonation based on emotion when generating audio data.
[0687] A "terminal" is a device used by a user to input and send text, and is a device that has the function of sending data to a server in the format of an HTTP request.
[0688] This invention combines an emotion recognition function with a system that allows hearing-impaired individuals to input what they want to say as text, analyzes that text to generate audio data, and transmits it to hearing individuals. This allows the system to recognize the user's emotions from the input text data and reflect those emotions during speech synthesis.
[0689] This system includes a terminal for user text input, a server that analyzes the input text and generates audio data, and a terminal that plays the generated audio data. In particular, it incorporates an emotion engine that recognizes the user's emotions from the text data.
[0690] User input and terminal operation
[0691] User:
[0692] The user enters what they want to say in text format. For example, the user might type "Hello, how are you?" into the text box on their device.
[0693] Terminal:
[0694] The terminal provides a text input field, receives input from the user, and sends the content to the server when the submit button is pressed. Specifically, it sends data to the server using an HTTP request.
[0695] Server operation
[0696] server:
[0697] The server receives text data sent from the terminal. The received text data is passed to a natural language processing system for analysis. Next, an emotion engine is used to recognize the user's emotions from the text data.
[0698] Natural language processing methods:
[0699] This technology analyzes received text data to understand the meaning of what the user is saying. For example, it analyzes and understands the meaning of the text, "Hello, how are you?"
[0700] Emotional engine:
[0701] Based on the analysis results from natural language processing tools, the system recognizes the user's emotions. For example, it can recognize positive emotions from the text "Hello, how are you?".
[0702] Speech synthesis method:
[0703] Using parameters generated by emotion recognition, the input text is converted into audio data. Pitch, speed, intonation, and other parameters are adjusted. For example, audio data is generated with a bright tone that reflects positive emotions.
[0704] Transmission method:
[0705] The speech synthesis system sends the generated audio data back to the terminal. The audio data is transmitted using an HTTP response.
[0706] Device playback operation
[0707] Terminal:
[0708] The device receives the audio data sent back from the server and plays it for the user using its audio playback function. This allows the entered text content to be played back as audio, reflecting the emotions it conveys. Specifically, it might say, "Hello, how are you?" with positive emotions.
[0709] Specific example
[0710] Input and submission:
[0711] User: Type "Hello, how are you?" into the text box and press the send button.
[0712] Text analysis and sentiment recognition:
[0713] Server: Receives and analyzes the message "Hello, how are you?" and recognizes positive emotions.
[0714] Speech generation:
[0715] Server: The speech synthesis system sets parameters to match positive emotions and converts "Hello, how are you?" into voice data in a cheerful tone.
[0716] reproduction:
[0717] Terminal: Receives audio data and plays it back using its audio playback function. The played audio conveys the message, "Hello, how are you?" along with a positive tone.
[0718] Concrete examples of prompt sentences for generative AI models
[0719] Prompt example:
[0720] "Analyze the following text data to recognize the user's emotions. Generate emotion-based speech synthesis parameters and then generate the speech data."
[0721] This system allows emotions to be reflected in the text entered by people with hearing impairments, enabling more natural and enriching communication.
[0722] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0723] Step 1:
[0724] User:
[0725] Inputs and outputs:
[0726] Input: The user enters "Hello, how are you?" into the text input field on the device.
[0727] Output: The entered text "Hello, how are you?" is stored in the terminal.
[0728] Specific actions:
[0729] The user enters what they want to say into the text input field on their device and presses the send button.
[0730] Step 2:
[0731] Terminal:
[0732] Inputs and outputs:
[0733] Input: The text entered in Step 1: "Hello, how are you?"
[0734] Output: Text data is sent to the server as an HTTP request.
[0735] Specific actions:
[0736] The terminal retrieves the contents of the text input field and sends it to the server using an HTTP POST request.
[0737] Step 3:
[0738] server:
[0739] Inputs and outputs:
[0740] Input: Text data sent from the device: "Hello, how are you?"
[0741] Output: Text data is received by the server.
[0742] Specific actions:
[0743] The server receives a POST request sent from the terminal and extracts text data from the request body.
[0744] Step 4:
[0745] server:
[0746] Inputs and outputs:
[0747] Input: Text data received in Step 3: "Hello, how are you?"
[0748] Output: Analysis results of text data (meaning and content of the statement)
[0749] Specific actions:
[0750] The server passes text data to a natural language processing system, which then analyzes the text and understands its meaning.
[0751] Step 5:
[0752] server:
[0753] Inputs and outputs:
[0754] Input: Meaning of the text data "Hello, how are you?" analyzed in Step 4
[0755] Output: Emotion recognition result (e.g., positive)
[0756] Specific actions:
[0757] The server uses an emotion engine to recognize the user's emotions from the analysis results. For example, it recognizes positive emotions from "Hello, how are you?".
[0758] Step 6:
[0759] server:
[0760] Inputs and outputs:
[0761] Input: Emotions recognized in Step 5 and text data analyzed in Step 4
[0762] Output: Speech synthesis parameters (e.g., bright tone, appropriate speed and intonation)
[0763] Specific actions:
[0764] The server sets parameters for generating speech data using speech synthesis based on the emotion recognition results. For example, it sets parameters for a bright tone based on positive emotions.
[0765] Step 7:
[0766] server:
[0767] Inputs and outputs:
[0768] Input: Speech synthesis parameters generated in step 6
[0769] Output: Audio data (Example: "Hello, how are you?" in a cheerful tone)
[0770] Specific actions:
[0771] The server uses speech synthesis to convert the input text into speech data based on the configured parameters.
[0772] Step 8:
[0773] server:
[0774] Inputs and outputs:
[0775] Input: Audio data generated in Step 7
[0776] Output: The audio data is sent to the terminal as an HTTP response.
[0777] Specific actions:
[0778] The server sends the generated audio data back to the terminal as an HTTP response.
[0779] Step 9:
[0780] Terminal:
[0781] Inputs and outputs:
[0782] Input: Audio data sent from the server
[0783] Output: The voice played by the user (e.g., "Hello, how are you?" in a cheerful tone)
[0784] Specific actions:
[0785] The device receives audio data sent from the server and plays it back using its audio playback function. Because the played audio reflects the user's emotions, more natural communication is possible.
[0786] (Application Example 2)
[0787] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0788] Traditionally, there has been a lack of adequate means for hearing-impaired individuals to quickly generate appropriate voice messages, including emotions, in emergencies and notify security guards and emergency services. This has led to delays in emergency communication and the potential for misunderstandings. This invention aims to solve these problems by recognizing emotions based on text input and generating voice data that includes those emotions to notify emergency services.
[0789] In Application Example 2, the specific processing performed by the specific processing unit 290 of the data processing device 12 is realized by the following means. In this invention, the server includes an input means for receiving user input, a natural language processing means for analyzing the received text data and generating parameters for speech synthesis, a speech synthesis means for generating speech data using the parameters generated by the natural language processing means, an emotion recognition means for recognizing emotions from the text data using the natural language processing means and reflecting that emotion information in the speech synthesis parameters, a transmission means for transmitting the generated speech data to a user terminal, a playback means for playing the speech data on the user terminal, and a notification means for analyzing the input text data in an emergency, generating a speech message including emotions, and notifying security guards or emergency services. This makes it possible for hearing-impaired individuals to quickly and accurately generate a speech message including emotions in an emergency and to respond appropriately to the emergency.
[0790] "Input means" refers to a device or interface for a user to input text data.
[0791] "Natural language processing means" refers to systems and algorithms that analyze received text data and generate parameters for speech synthesis.
[0792] "Speech synthesis means" refers to a device or software that generates speech data using parameters generated by natural language processing means.
[0793] "Emotion recognition means" refers to a system or algorithm that recognizes a user's emotions from text data and reflects that emotional information in speech synthesis parameters.
[0794] "Transmission means" refers to the communication means used to send the generated audio data to the user's terminal.
[0795] "Playback means" refers to a device or software that plays back audio data transmitted to the user's terminal.
[0796] "Notification means" refers to a mechanism or system that analyzes text data entered in an emergency, generates an audio message including emotions, and notifies security guards or emergency services.
[0797] The system according to the present invention is a system for enabling hearing-impaired individuals to quickly and accurately generate emotionally charged voice messages in emergencies and notify security guards and emergency services. This system mainly consists of an input means, a natural language processing means, a speech synthesis means, an emotion recognition means, a transmission means, a playback means, and a notification means.
[0798] Hardware and software configuration
[0799] Hardware used: Smartphone or robot (with voice output function)
[0800] Software used:
[0801] text2emotion: Emotion Recognition Library
[0802] gTTS (Google Text-to-Speech): A speech synthesis library.
[0803] requests: A library for sending HTTP requests.
[0804] Operational description of each means
[0805] 1. Input method:
[0806] This provides an interface for users to enter text data in emergencies. This interface will be implemented as a text input field on smartphones or robots. Users will enter emergency messages here.
[0807] 2. Natural language processing methods:
[0808] The system analyzes text data received through the input means and generates parameters for speech synthesis. This step is crucial as it involves understanding the content of the text data and converting it into speech data.
[0809] 3. Speech synthesis means:
[0810] Speech data is generated using parameters produced by natural language processing. During this process, pitch, speed, intonation, and other parameters are adjusted to produce natural-sounding speech that also conveys emotion.
[0811] 4. Emotion recognition means:
[0812] The system analyzes text data received from a natural language processing system to recognize the user's emotions. This recognized emotion information is then reflected in the parameters of the speech synthesis system. The text2emotion library is used for this emotion recognition.
[0813] 5. Transmission method:
[0814] This is a communication method for sending generated audio data to the user's device. This method uses the requests library to send audio data in HTTP request format.
[0815] 6. Regeneration means:
[0816] This is a device or software that plays back audio data transmitted to a user's terminal. It provides audio output on smartphones and robots.
[0817] 7. Means of notification:
[0818] In emergencies, the system analyzes entered text data, generates an emotionally charged voice message, and immediately notifies security personnel or emergency services. This notification is sent to a server by transmitting the generated voice data.
[0819] Specific example
[0820] User actions:
[0821] The user types "Help me, there's a fire" into the text input field of their smartphone or robot. Then they press the send button.
[0822] Server operation:
[0823] The server analyzes the received text data using natural language processing and recognizes the emotion of fear from the content of "Fire!". As a result, it generates audio data that includes the emotion of urgency and fear using speech synthesis.
[0824] Speech synthesis:
[0825] The speech synthesis means uses speech parameters that reflect the emotion of fear recognized by the emotion recognition means to generate an urgent voice message saying, "Help me, there's a fire."
[0826] Emergency notification:
[0827] The generated audio data is transmitted to security guards and emergency services via a transmission device to facilitate a rapid response.
[0828] Example of a prompt
[0829] Towards emotion recognition AI:
[0830] Text: "Help me, there's a fire."
[0831] Please return the emotion recognition results in JSON format (e.g., {"happy": 0.1, "sad": 0.3, "angry": 0.1, "fear": 0.5}).
[0832] Towards speech synthesis AI:
[0833] Text: "Help me, there's a fire."
[0834] Recognized emotion: fear
[0835] Please generate natural-sounding voice data that reflects emotions.
[0836] In this way, the present invention enables communication that includes the emotions of hearing-impaired individuals in emergencies, and facilitates a rapid and appropriate emergency response.
[0837] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0838] Step 1:
[0839] The user enters an emergency message into a text input field on their smartphone or robot and presses the send button. The entered text data contains emergency information, such as "Help, there's a fire." This input is then sent to the server via the input device.
[0840] Input: User-entered text data "Help, there's a fire"
[0841] Output: Text data sent to the server
[0842] Step 2:
[0843] The server passes the received text data to a natural language processing system, which analyzes the text's content. The natural language processing system understands the meaning from the input text and generates basic parameters for speech synthesis.
[0844] Input: Received text data
[0845] Data Processing / Calculation: Analyze text content using natural language processing techniques and generate speech synthesis parameters.
[0846] Output: Generated speech synthesis parameters
[0847] Step 3:
[0848] The natural language processing system passes text data to the emotion recognition system to recognize the user's emotions. Here, the text2emotion library is used to analyze the emotions implied in the text and identify the main emotions.
[0849] Input: Text data passed from a natural language processing system.
[0850] Data processing / calculation: Recognizing emotions using the text2emotion library.
[0851] Output: Recognized emotion information (Example: {"fear": 0.5, "sad": 0.3, "happy": 0.1, "angry": 0.1})
[0852] Step 4:
[0853] Based on the emotion information recognized by the emotion recognition means, the speech synthesis means generates speech data. Using the gTTS library, speech data with speech parameters that reflect the emotion is created.
[0854] Input: Sentimental information and speech synthesis parameters
[0855] Data processing / calculation: Generate audio data using the gTTS library.
[0856] Output: Generated audio data file
[0857] Step 5:
[0858] The generated audio data file is sent via HTTP request to the security guard or emergency service system, which is the recipient of the emergency notification. The requests library is used.
[0859] Input: Generated audio data file
[0860] Data processing / calculation: Data transmission via HTTP requests
[0861] Output: Completion of voice data transmission to security guards and emergency services.
[0862] Step 6:
[0863] Security and emergency service systems play back transmitted audio data and initiate a rapid response. The playback mechanism uses an audio output device to play back emergency messages that reflect emotions.
[0864] Input: Received audio data file
[0865] Data processing / calculation: Audio data playback
[0866] Output: Initiate emergency response and play voice message.
[0867] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0868] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0869] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0870] [Third Embodiment]
[0871] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0872] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0873] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0874] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0875] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0876] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0877] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0878] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0879] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0880] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0881] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0882] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0883] This system allows hearing-impaired individuals to input what they want to say as text, analyze that text to generate audio data, and then transmit it to hearing individuals. The configuration for implementing this system is described below.
[0884] System Overview
[0885] This system includes a terminal for user text input, a server that analyzes the input text and generates audio data, and a terminal that plays the generated audio data. The specific roles and processing flow of each component will be explained below.
[0886] User input and terminal operation
[0887] User: The user enters what they want to say in text format. For example, the user enters "Hello, how are you?" into the text box on their device.
[0888] Terminal: Provides a text input field, receives input from the user, and sends the content to the server when the submit button is pressed.
[0889] Server operation
[0890] Server: Receives text data sent from the terminal. Next, it analyzes the text using natural language processing and generates parameters for speech synthesis. For example, for the text "Hello, how are you?", it sets parameters such as language, pitch, and speed.
[0891] Natural language processing means: Analyzes the received text and derives appropriate speech synthesis parameters. Here, language models and sentiment analysis are used to determine parameters for generating more natural speech.
[0892] Speech synthesis means: Using parameters generated by natural language processing means, input text is converted into speech data. This speech data can then be converted into any desired audio format, such as a WAV file or MP3 file.
[0893] Transmission method: The generated audio data is sent back to the terminal. The terminal is configured to receive data via HTTP requests.
[0894] Device playback operation
[0895] Terminal: Receives audio data sent back from the server and plays it for the user. Uses software or devices with audio playback capabilities to transmit text input as audio.
[0896] Specific example
[0897] Input and Submission: The user enters "Hello, how are you?" into the text box and presses the submit button.
[0898] Text parsing: The server receives the text and converts "Hello, how are you?" into speech synthesis parameters.
[0899] Speech generation: A speech synthesis device converts the text "Hello, how are you?" into speech data.
[0900] Playback: The device receives and plays the audio data. The played audio contains the message, "Hello, how are you?"
[0901] This system enables people with hearing impairments to speak using their voices, facilitating smooth communication with hearing individuals. Details of each processing step are described below.
[0902] The following describes the processing flow.
[0903] Step 1: The user enters what they want to say in text format.
[0904] User: Enter "Hello, how are you?" into the text input field and click the send button.
[0905] Step 2: The terminal receives user input and sends the data to the server.
[0906] Terminal: Sends the entered text data "Hello, how are you?" to the server in POST request format.
[0907] Step 3: The server sends the received data to a natural language processing system for analysis.
[0908] Server: Passes the received text data to a natural language processing engine for analysis.
[0909] Step 4: The natural language processing engine analyzes the text data and generates speech synthesis parameters.
[0910] Server: The natural language processing engine analyzes the received text "Hello, how are you?" and generates parameters for speech synthesis (language, pitch, speed, etc.).
[0911] Step 5: Generate audio data using a speech synthesis method based on the analysis results.
[0912] Server: The speech synthesis engine generates speech data using parameters. The generated speech data contains the message, "Hello, how are you?"
[0913] Step 6: Send the generated audio data to the device.
[0914] Server: Sends the generated audio data to the terminal.
[0915] Step 7: The device plays the audio data received from the server.
[0916] Terminal: Passes the audio data received from the server to the playback function, which then plays the audio.
[0917] Step 8: The user communicates using the generated voice.
[0918] User: Uses the audio "Hello, how are you?" played from the device to communicate with a healthy person. During this process, the user confirms that their words have been transmitted as audio.
[0919] (Example 1)
[0920] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0921] In recent years, methods such as writing notes and using tablet devices have become common ways for people with hearing impairments to communicate smoothly with those without hearing impairments. However, these methods make real-time communication difficult, which is particularly inconvenient in emergencies or situations requiring a quick response. There is a need for a system that allows people with hearing impairments to communicate smoothly using voice.
[0922] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0923] In this invention, the server includes an input means for receiving user input, a natural language processing means for analyzing the text data received by the input means and generating parameters for speech synthesis, a speech synthesis means for generating speech data using the parameters generated by the natural language processing means, a transmission means for transmitting the speech data generated by the speech synthesis means to the user's terminal, and a playback means for playing the speech data on the user's terminal. This makes it possible for people with hearing impairments to quickly communicate their intentions through speech.
[0924] "An input means for receiving user input" refers to a device or system for users to input what they want to say in text format.
[0925] "Natural language processing means" refers to an algorithm or software that analyzes text data received by an input means and generates parameters for speech synthesis.
[0926] "Speech synthesis means" refers to an algorithm or software that converts input text into speech data using parameters generated by natural language processing means.
[0927] "Transmission means" refers to a device or system for transmitting audio data generated by speech synthesis means to a user's terminal.
[0928] "Playback means" refers to a device or software for playing audio data on the user's terminal.
[0929] "Adjustment means" refers to a device or software used to adjust parameters such as language, pitch, speed, and emotion when generating audio data.
[0930] A "terminal" is a device that allows a user to input text using an input device, send it to a server, and receive and play back the generated audio data.
[0931] This invention relates to a system for enabling hearing-impaired individuals to input what they wish to express as text, analyze that text to generate audio data, and communicate it to hearing individuals. The system includes a terminal for user text input, a server that analyzes the input text and generates audio data, and a terminal that plays the generated audio data.
[0932] System Overview
[0933] This system consists of the following elements:
[0934] 1. Input means for receiving user input
[0935] Users enter what they want to say in text format using an application on their smartphone or computer, or an input field on a webpage. For example, they might type, "Hello, how are you?"
[0936] 2. Transmission of text data via input means
[0937] Once the user completes the input and presses the submit button, the device sends this text data to the server. An HTTP POST request is used for submission.
[0938] 3. Text analysis by the server
[0939] The server receives text data sent from the terminal and analyzes it using natural language processing tools. This analysis uses language models such as Google's TensorFlow and OpenAI's GPT-3. It analyzes the sentiment and grammar of the received text and generates optimal parameters for speech synthesis.
[0940] 4. Generating speech synthesis parameters
[0941] The server's natural language processing system generates speech synthesis parameters such as pitch, speed, and emotion based on the analysis results. For example, it extracts appropriate intonation and emotion from the text "Hello, how are you?".
[0942] 5. Server-driven generation of audio data
[0943] Based on parameters generated by natural language processing, the server uses speech synthesis to generate actual audio data. This speech synthesis utilizes services such as Amazon Polly or Google Cloud Text-to-Speech. The generated audio data is saved as WAV or MP3 files.
[0944] 6. Sending and playing audio data
[0945] The server sends the generated audio data to the terminal as an HTTP response. The terminal plays the received audio data using a playback device. Playback requires HTML5. <audio>Tags or dedicated audio playback applications are used.
[0946] Specific example
[0947] For example, the process would be as follows:
[0948] Input and Sending: The user types "Hello, how are you?" into the text box on the device and presses the send button.
[0949] Text analysis: The server receives text, analyzes it, and converts it into speech synthesis parameters. The natural language processing model used is OpenAI's GPT-3, among others.
[0950] Speech generation: The server uses Amazon Polly to convert the text "Hello, how are you?" into speech data.
[0951] Playback: The device receives audio data and HTML5 <audio>Play using tags.
[0952] Example of a prompt
[0953] "Please use GPT-3 to convert the text 'Hello, how are you?' into speech synthesis parameters. Then, use a speech synthesis service to generate the speech data and send it back to the terminal."
[0954] This system enables hearing-impaired individuals to communicate quickly through voice. The specific processing steps at each stage are carried out based on the means described in the claims.
[0955] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0956] Step 1:
[0957] User: The user enters what they want to say in text format using a smartphone or computer. For example, they might type, "Hello, how are you?"
[0958] Input: Text "Hello, how are you?"
[0959] Output: Text data entered into the user's terminal
[0960] Step 2:
[0961] Terminal: When the user presses the send button, the terminal sends the entered text data to the server. An HTTP POST request is used for transmission. Specifically, the text data is sent to the server in JSON format.
[0962] Input: Text data "Hello, how are you?"
[0963] Output: JSON data sent as an HTTP POST request
[0964] Step 3:
[0965] Server: The server receives an HTTP POST request and extracts text data. The received text data is then passed to a natural language processing system.
[0966] Input: JSON data sent to the server
[0967] Output: Text data passed to the natural language processing system.
[0968] Step 4:
[0969] Server: The natural language processing system analyzes text data and generates parameters for speech synthesis (e.g., language, pitch, speed, emotion, etc.). Examples of natural language processing models used include TensorFlow and GPT-3. For example, it performs emotion analysis on the text "Hello, how are you?" and generates appropriate intonation parameters.
[0970] Input: Text data "Hello, how are you?"
[0971] Output: Speech synthesis parameters (language, pitch, speed, emotion, etc.)
[0972] Step 5:
[0973] Server: Using the generated parameters, the speech synthesis system generates speech data. This system may use Amazon Polly or Google Cloud Text-to-Speech. For example, speech data for "Hello, how are you?" is generated based on the parameters.
[0974] Input: Speech synthesis parameters
[0975] Output: Audio data (such as WAV or MP3 files)
[0976] Step 6:
[0977] Server: Sends the generated audio data to the terminal as an HTTP response. At this time, the appropriate MIME type (e.g., audio / mpeg) is set in the response header.
[0978] Input: Audio data
[0979] Output: HTTP response sent to the terminal
[0980] Step 7:
[0981] Terminal: The terminal analyzes the received audio data and plays it back using its audio playback function. Specifically, it uses HTML5 <audio>Tags or dedicated audio playback applications are used. For example, the received audio data is played back to the user as "Hello, how are you?".
[0982] Input: Audio data received as an HTTP response
[0983] Output: Audio played to the user
[0984] This process allows hearing-impaired individuals to input what they want to say as text and transmit it to hearing individuals as audio.
[0985] (Application Example 1)
[0986] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0987] People with hearing impairments often have difficulty communicating with store staff in physical stores. Therefore, support is needed in situations where smooth communication is required, such as when asking for product information or completing a purchase. Furthermore, existing systems have limited capabilities for converting text data to audio data, making them unsuitable for use in physical stores.
[0988] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0989] In this invention, the server includes an input means for receiving user input, a natural language processing means for analyzing the text data received by the input means and generating parameters for speech synthesis, and a speech synthesis means for generating speech data using the parameters generated by the natural language processing means. This makes it possible to generate speech data and transmit it to other people in a physical store in order to support communication between people with hearing impairments and other people in the store.
[0990] A "user" is a person who uses this system to input text and convert that text into speech.
[0991] "Input means" refers to a device or function for a user to input text data. For example, this includes terminals such as smartphones and smart glasses.
[0992] "Natural language processing means" refers to a device or program that has the function of analyzing text data received by an input means and generating parameters for speech synthesis.
[0993] "Speech synthesis means" refers to a device or function that generates speech data using parameters generated by natural language processing means.
[0994] "Transmission means" refers to a device or function for transmitting generated audio data to the user's terminal.
[0995] "Playback means" refers to a device or function for playing audio data on a user terminal.
[0996] "A means of generating audio data and transmitting it to other people in a physical store to support communication between people with hearing impairments and other people in a physical store" refers to a means of converting text data entered by a person with hearing impairment into audio data and transmitting that audio to other people in a physical store.
[0997] This system is designed to facilitate communication between people with hearing impairments and other people in physical stores. A specific implementation of this system is described below.
[0998] System Configuration
[0999] User: A person with a hearing impairment performs text input.
[1000] Terminal: Provides a text input field and accepts text input from the user. This includes devices such as smartphones and smart glasses.
[1001] Server: Analyzes text data sent from the terminal, generates speech synthesis parameters, and creates speech data.
[1002] Natural language processing means: Analyzes text data and derives appropriate speech synthesis parameters.
[1003] Speech synthesis means: Generates speech data based on parameters.
[1004] Transmission method: The generated audio data is sent to the terminal.
[1005] Playback method: The audio data is played on the device and conveyed to other people, such as store clerks.
[1006] Hardware and software used
[1007] Hardware: Smartphones, smart glasses, servers.
[1008] Software: Python, Google Text-to-Speech (gTTS) library, Playsound library.
[1009] Processing flow and data manipulation / calculation
[1010] 1. The user enters text.
[1011] Enter text into the input field on your smartphone or smart glasses.
[1012] 2. The terminal sends text data to the server.
[1013] The text data is sent to the server in the format of an HTTP POST request.
[1014] 3. The server analyzes the text data and generates parameters.
[1015] The system analyzes text using natural language processing techniques and generates parameters such as language, pitch, and speed.
[1016] 4. The server generates the audio data.
[1017] Speech data is generated using parameters produced by a speech synthesis system.
[1018] 5. The server sends the audio data to the terminal.
[1019] The generated audio data is sent to the device.
[1020] 6. The device plays the audio data.
[1021] Audio data is played back using a playback device (smartphone or smart glasses) to convey the intentions of the hearing-impaired person to store staff, etc.
[1022] Specific example
[1023] For example, suppose a hearing-impaired person enters "Where is this product?" into smart glasses in a physical store. This text input is sent to a server, which analyzes the text and generates audio data saying "Where is this product?". This generated audio data is sent to the smart glasses and played back, allowing the store clerk to understand that the question is from a hearing-impaired person and guide them to the product's location.
[1024] Example of a prompt
[1025] Prompt: The user typed "Where can I find this product?". Convert this text into audio data and return the URL.
[1026] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[1027] Step 1:
[1028] The user enters text.
[1029] Input: A person with a hearing impairment types "Where can I find this product?" into the text input field on their smartphone or smart glasses.
[1030] Action: The user types text into the input field and presses the submit button.
[1031] Output: The entered text data is recorded on the terminal.
[1032] Step 2:
[1033] The terminal sends text data to the server.
[1034] Input: Text data entered by the user.
[1035] Operation: The device sends text data to the server using an HTTP POST request.
[1036] Output: The server receives text data.
[1037] Step 3:
[1038] The server analyzes the text data and generates parameters.
[1039] Input: Text data received by the server: "Where is this product located?"
[1040] Operation: The server's natural language processing system analyzes the text and generates speech synthesis parameters such as language, pitch, and speed.
[1041] Output: Generated speech synthesis parameters.
[1042] Step 4:
[1043] The server generates the audio data.
[1044] Input: Speech synthesis parameters generated by the server.
[1045] Operation: The server's speech synthesis system generates speech data based on the parameters.
[1046] Output: Generated audio data (e.g., MP3 file).
[1047] Step 5:
[1048] The server sends the audio data to the terminal.
[1049] Input: Generated audio data.
[1050] Operation: The server's transmission method sends the generated audio data to the terminal in HTTP response format.
[1051] Output: The device receives the audio data.
[1052] Step 6:
[1053] The device plays audio data
[1054] Input: Received audio data.
[1055] Operation: The device's playback mechanism plays the received audio data. Audio is played from the smartphone or smart glasses, and the store clerk receives the question from the hearing-impaired person.
[1056] Output: The audio is played in the physical store, allowing the intentions of the hearing-impaired person to be conveyed to others.
[1057] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[1058] This invention combines an emotion recognition function with a system that allows hearing-impaired individuals to input what they want to say as text, analyzes that text to generate audio data, and transmits it to hearing individuals. This allows the system to recognize the user's emotions from the input text data and reflect those emotions during speech synthesis.
[1059] System Overview
[1060] This system includes a terminal for user text input, a server that analyzes the input text and generates audio data, and a terminal that plays the generated audio data. In particular, it incorporates an emotion engine that recognizes the user's emotions from the text data.
[1061] User input and terminal operation
[1062] User: The user enters what they want to say in text format. For example, the user enters "Hello, how are you?" into the text box on their device.
[1063] Terminal: Provides a text input field, receives input from the user, and sends the content to the server when the submit button is pressed.
[1064] Server operation
[1065] Server: Receives text data sent from the terminal. Next, natural language processing means analyze the text and generate parameters for speech synthesis. In particular, emotion recognition means recognize the user's emotions from the text data and reflect that emotion information in the speech synthesis parameters.
[1066] Natural language processing means: Analyzes received text data and uses emotion recognition means to analyze the user's emotions. For example, it recognizes a positive emotion in response to the text "Hello, how are you?". Based on this, it sets parameters to make the voice sound brighter.
[1067] Emotion Engine: Recognizes user emotions from text data and generates emotion classification results. These results are sent to a speech synthesis system and used to generate speech data with natural-sounding expressions.
[1068] Speech synthesis means: Using parameters generated by emotion recognition means, input text is converted into speech data. Pitch, speed, intonation, etc., are added, adjusted based on emotion.
[1069] Transmission method: The generated audio data is sent back to the terminal. The terminal is configured to receive data via HTTP requests.
[1070] Device playback operation
[1071] Terminal: Receives audio data sent back from the server and plays it for the user. Using software or devices with audio playback capabilities, it transmits text input as audio. Because the audio data reflects the user's emotions, more natural communication is possible.
[1072] Specific example
[1073] Input and Submission: The user enters "Hello, how are you?" into the text box and presses the submit button.
[1074] Text analysis and sentiment recognition: The server receives text, analyzes it to see if it says "Hello, how are you?", and recognizes a positive sentiment.
[1075] Speech generation: The speech synthesis system sets parameters based on emotion recognition and converts the text "Hello, how are you?" into speech data in a cheerful tone.
[1076] Playback: The device receives and plays the audio data. The played audio conveys the message, "Hello, how are you?" with a positive tone.
[1077] This system enables hearing-impaired individuals to not only speak using their voices but also to convey their emotions, thereby achieving richer communication. Details of each processing step will be explained next.
[1078] The following describes the processing flow.
[1079] Step 1: The user enters what they want to say in text format.
[1080] User: Enter "Hello, how are you?" into the text input field and click the send button.
[1081] Step 2: The terminal receives user input and sends the data to the server.
[1082] Terminal: Sends the entered text data "Hello, how are you?" to the server in POST request format.
[1083] Step 3: The server sends the received data to a natural language processing system for analysis.
[1084] Server: Passes the received text data to a natural language processing engine for analysis.
[1085] Step 4: The natural language processing engine analyzes the text data and sends it to the sentiment recognition engine.
[1086] Server: The natural language processing engine analyzes the received text "Hello, how are you?" and passes the result to the emotion recognition engine.
[1087] Step 5: The emotion recognition engine extracts emotions from the text and generates emotional information.
[1088] Server: The emotion recognition engine extracts positive emotions from the text "Hello, how are you?" and generates emotion information.
[1089] Step 6: Set the speech synthesis parameters along with the emotional information.
[1090] Server: Based on the emotion information obtained from the emotion recognition engine, it sets the parameters used by the speech synthesis engine (e.g., pitch, speed, intonation).
[1091] Step 7: The speech synthesis engine generates speech data that reflects the emotion parameters.
[1092] Server: The speech synthesis engine converts the text "Hello, how are you?" into voice data with a cheerful tone, based on the configured parameters.
[1093] Step 8: Send the generated audio data to the device.
[1094] Server: Prepares to send the generated audio data to the terminal and returns it via an HTTP response.
[1095] Step 9: Play the audio data received by the device.
[1096] Terminal: Passes the audio data received from the server to the playback function, and plays the audio.
[1097] Step 10: The user communicates using the generated voice.
[1098] User: Uses a voice message containing positive emotions, such as "Hello, how are you?", played from the device to communicate with a healthy person.
[1099] This detailed processing step allows for the generation of audio data that reflects not only the text content entered by the user, but also their emotions, thereby supporting more effective communication with individuals without cognitive impairment.
[1100] (Example 2)
[1101] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[1102] Conventional speech synthesis systems can convert text input by hearing-impaired individuals into speech data, but they cannot reflect the user's emotions, thus limiting their communication capabilities. In particular, the inability to convey the emotions behind speech makes accurate and rich communication difficult. Solving this problem is essential.
[1103] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[1104] In this invention, the server includes an input means for receiving user input, a natural language processing means for analyzing the text data received by the input means, recognizing emotions, and generating parameters for speech synthesis, and a speech synthesis means for generating speech data using the parameters based on the emotions recognized by the natural language processing means. This makes it possible to recognize emotions from text data entered by a hearing-impaired person and generate speech data that reflects those emotions.
[1105] An "input means for receiving user input" is an interface for users to input what they want to say in text format, and a device that has the function of sending the entered text to the system.
[1106] "Natural language processing means" refers to technology for analyzing user input and understanding its content, and in particular, a device that has the function of recognizing emotions and generating parameters for speech synthesis.
[1107] An "emotion recognition device" is a device that analyzes the user's emotions from input text data and classifies those emotions into specific categories.
[1108] A "speech synthesis means" is a technology that converts input text into speech data using parameters generated by a natural language processing means, and is a device that has the function of reflecting emotion-based adjustments.
[1109] A "transmission means" is a device that has the function of transmitting the generated audio data to the user's terminal.
[1110] "Playback means" refers to a device or software for playing back audio data received on a user terminal.
[1111] A "modification device" is a device that has the function of adjusting parameters such as language, pitch, speed, and intonation based on emotion when generating audio data.
[1112] A "terminal" is a device used by a user to input and send text, and is a device that has the function of sending data to a server in the format of an HTTP request.
[1113] This invention combines an emotion recognition function with a system that allows hearing-impaired individuals to input what they want to say as text, analyzes that text to generate audio data, and transmits it to hearing individuals. This allows the system to recognize the user's emotions from the input text data and reflect those emotions during speech synthesis.
[1114] This system includes a terminal for user text input, a server that analyzes the input text and generates audio data, and a terminal that plays the generated audio data. In particular, it incorporates an emotion engine that recognizes the user's emotions from the text data.
[1115] User input and terminal operation
[1116] User:
[1117] The user enters what they want to say in text format. For example, the user might type "Hello, how are you?" into the text box on their device.
[1118] Terminal:
[1119] The terminal provides a text input field, receives input from the user, and sends the content to the server when the submit button is pressed. Specifically, it sends data to the server using an HTTP request.
[1120] Server operation
[1121] server:
[1122] The server receives text data sent from the terminal. The received text data is passed to a natural language processing system for analysis. Next, an emotion engine is used to recognize the user's emotions from the text data.
[1123] Natural language processing methods:
[1124] This technology analyzes received text data to understand the meaning of what the user is saying. For example, it analyzes and understands the meaning of the text, "Hello, how are you?"
[1125] Emotional engine:
[1126] Based on the analysis results from natural language processing tools, the system recognizes the user's emotions. For example, it can recognize positive emotions from the text "Hello, how are you?".
[1127] Speech synthesis method:
[1128] Using parameters generated by emotion recognition, the input text is converted into audio data. Pitch, speed, intonation, and other parameters are adjusted. For example, audio data is generated with a bright tone that reflects positive emotions.
[1129] Transmission method:
[1130] The speech synthesis system sends the generated audio data back to the terminal. The audio data is transmitted using an HTTP response.
[1131] Device playback operation
[1132] Terminal:
[1133] The device receives the audio data sent back from the server and plays it for the user using its audio playback function. This allows the entered text content to be played back as audio, reflecting the emotions it conveys. Specifically, it might say, "Hello, how are you?" with positive emotions.
[1134] Specific example
[1135] Input and submission:
[1136] User: Type "Hello, how are you?" into the text box and press the send button.
[1137] Text analysis and sentiment recognition:
[1138] Server: Receives and analyzes the message "Hello, how are you?" and recognizes positive emotions.
[1139] Speech generation:
[1140] Server: The speech synthesis system sets parameters to match positive emotions and converts "Hello, how are you?" into voice data in a cheerful tone.
[1141] reproduction:
[1142] Terminal: Receives audio data and plays it back using its audio playback function. The played audio conveys the message, "Hello, how are you?" along with a positive tone.
[1143] Concrete examples of prompt sentences for generative AI models
[1144] Prompt example:
[1145] "Analyze the following text data to recognize the user's emotions. Generate emotion-based speech synthesis parameters and then generate the speech data."
[1146] This system allows emotions to be reflected in the text entered by people with hearing impairments, enabling more natural and enriching communication.
[1147] The flow of the specific processing in Example 2 will be explained using Figure 13.
[1148] Step 1:
[1149] User:
[1150] Inputs and outputs:
[1151] Input: The user enters "Hello, how are you?" into the text input field on the device.
[1152] Output: The entered text "Hello, how are you?" is stored in the terminal.
[1153] Specific actions:
[1154] The user enters what they want to say into the text input field on their device and presses the send button.
[1155] Step 2:
[1156] Terminal:
[1157] Inputs and outputs:
[1158] Input: The text entered in Step 1: "Hello, how are you?"
[1159] Output: Text data is sent to the server as an HTTP request.
[1160] Specific actions:
[1161] The terminal retrieves the contents of the text input field and sends it to the server using an HTTP POST request.
[1162] Step 3:
[1163] server:
[1164] Inputs and outputs:
[1165] Input: Text data sent from the device: "Hello, how are you?"
[1166] Output: Text data is received by the server.
[1167] Specific actions:
[1168] The server receives a POST request sent from the terminal and extracts text data from the request body.
[1169] Step 4:
[1170] server:
[1171] Inputs and outputs:
[1172] Input: Text data received in Step 3: "Hello, how are you?"
[1173] Output: Analysis results of text data (meaning and content of the statement)
[1174] Specific actions:
[1175] The server passes text data to a natural language processing system, which then analyzes the text and understands its meaning.
[1176] Step 5:
[1177] server:
[1178] Inputs and outputs:
[1179] Input: Meaning of the text data "Hello, how are you?" analyzed in Step 4
[1180] Output: Emotion recognition result (e.g., positive)
[1181] Specific actions:
[1182] The server uses an emotion engine to recognize the user's emotions from the analysis results. For example, it recognizes positive emotions from "Hello, how are you?".
[1183] Step 6:
[1184] server:
[1185] Inputs and outputs:
[1186] Input: Emotions recognized in Step 5 and text data analyzed in Step 4
[1187] Output: Speech synthesis parameters (e.g., bright tone, appropriate speed and intonation)
[1188] Specific actions:
[1189] The server sets parameters for generating speech data using speech synthesis based on the emotion recognition results. For example, it sets parameters for a bright tone based on positive emotions.
[1190] Step 7:
[1191] server:
[1192] Inputs and outputs:
[1193] Input: Speech synthesis parameters generated in step 6
[1194] Output: Audio data (Example: "Hello, how are you?" in a cheerful tone)
[1195] Specific actions:
[1196] The server uses speech synthesis to convert the input text into speech data based on the configured parameters.
[1197] Step 8:
[1198] server:
[1199] Inputs and outputs:
[1200] Input: Audio data generated in Step 7
[1201] Output: The audio data is sent to the terminal as an HTTP response.
[1202] Specific actions:
[1203] The server sends the generated audio data back to the terminal as an HTTP response.
[1204] Step 9:
[1205] Terminal:
[1206] Inputs and outputs:
[1207] Input: Audio data sent from the server
[1208] Output: The voice played by the user (e.g., "Hello, how are you?" in a cheerful tone)
[1209] Specific actions:
[1210] The device receives audio data sent from the server and plays it using its audio playback function. Because the played audio reflects the user's emotions, more natural communication is possible.
[1211] (Application Example 2)
[1212] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[1213] Traditionally, there has been a lack of adequate means for hearing-impaired individuals to quickly generate appropriate voice messages, including emotions, in emergencies and notify security guards and emergency services. This has led to delays in emergency communication and the potential for misunderstandings. This invention aims to solve these problems by recognizing emotions based on text input and generating voice data that includes those emotions to notify emergency services.
[1214] In Application Example 2, the specific processing performed by the specific processing unit 290 of the data processing device 12 is realized by the following means. In this invention, the server includes an input means for receiving user input, a natural language processing means for analyzing the received text data and generating parameters for speech synthesis, a speech synthesis means for generating speech data using the parameters generated by the natural language processing means, an emotion recognition means for recognizing emotions from the text data using the natural language processing means and reflecting that emotion information in the speech synthesis parameters, a transmission means for transmitting the generated speech data to a user terminal, a playback means for playing the speech data on the user terminal, and a notification means for analyzing the input text data in an emergency, generating a speech message including emotions, and notifying security guards or emergency services. This makes it possible for hearing-impaired individuals to quickly and accurately generate a speech message including emotions in an emergency and to respond appropriately to the emergency.
[1215] "Input means" refers to a device or interface for a user to input text data.
[1216] "Natural language processing means" refers to systems and algorithms that analyze received text data and generate parameters for speech synthesis.
[1217] "Speech synthesis means" refers to a device or software that generates speech data using parameters generated by natural language processing means.
[1218] "Emotion recognition means" refers to a system or algorithm that recognizes a user's emotions from text data and reflects that emotional information in speech synthesis parameters.
[1219] "Transmission means" refers to the communication means used to send the generated audio data to the user's terminal.
[1220] "Playback means" refers to a device or software that plays back audio data transmitted to the user's terminal.
[1221] "Notification means" refers to a mechanism or system that analyzes text data entered in an emergency, generates an audio message including emotions, and notifies security guards or emergency services.
[1222] The system according to the present invention is a system for enabling hearing-impaired individuals to quickly and accurately generate emotionally charged voice messages in emergencies and notify security guards and emergency services. This system mainly consists of an input means, a natural language processing means, a speech synthesis means, an emotion recognition means, a transmission means, a playback means, and a notification means.
[1223] Hardware and software configuration
[1224] Hardware used: Smartphone or robot (with voice output function)
[1225] Software used:
[1226] text2emotion: Emotion Recognition Library
[1227] gTTS (Google Text-to-Speech): A speech synthesis library.
[1228] requests: Library for sending HTTP requests
[1229] Operational description of each means
[1230] 1. Input method:
[1231] This provides an interface for users to enter text data in emergencies. This interface will be implemented as a text input field on smartphones or robots. Users will enter emergency messages here.
[1232] 2. Natural language processing tools:
[1233] The system analyzes text data received through an input means and generates parameters for speech synthesis. This step is crucial as it involves understanding the content of the text data and converting it into speech data.
[1234] 3. Speech synthesis means:
[1235] Speech data is generated using parameters produced by natural language processing. During this process, pitch, speed, intonation, and other parameters are adjusted to produce natural-sounding speech that also conveys emotion.
[1236] 4. Emotion recognition means:
[1237] The system analyzes text data received from a natural language processing system to recognize the user's emotions. This recognized emotion information is then reflected in the parameters of the speech synthesis system. The text2emotion library is used for this emotion recognition.
[1238] 5. Transmission method:
[1239] This is a communication method for sending generated audio data to the user's device. This method uses the requests library to send audio data in HTTP request format.
[1240] 6. Regeneration means:
[1241] This is a device or software that plays back audio data transmitted to a user's terminal. It provides audio output on smartphones and robots.
[1242] 7. Means of notification:
[1243] In an emergency, the system analyzes the entered text data, generates an emotionally charged voice message, and immediately notifies security personnel or emergency services. This notification is sent to a server by transmitting the generated voice data.
[1244] Specific example
[1245] User actions:
[1246] The user types "Help me, there's a fire" into the text input field of their smartphone or robot. Then they press the send button.
[1247] Server operation:
[1248] The server analyzes the received text data using natural language processing and recognizes the emotion of fear from the content of "Fire!". As a result, it generates audio data that includes the feeling of urgency and fear using speech synthesis.
[1249] Speech synthesis:
[1250] The speech synthesis means uses speech parameters that reflect the emotion of fear recognized by the emotion recognition means to generate an urgent voice message saying, "Help me, there's a fire."
[1251] Emergency notification:
[1252] The generated audio data is transmitted to security guards and emergency services via a transmission device to facilitate a rapid response.
[1253] Example of a prompt
[1254] Towards emotion recognition AI:
[1255] Text: "Help me, there's a fire."
[1256] Please return the emotion recognition results in JSON format (e.g., {"happy": 0.1, "sad": 0.3, "angry": 0.1, "fear": 0.5}).
[1257] Towards speech synthesis AI:
[1258] Text: "Help me, there's a fire."
[1259] Recognized emotion: fear
[1260] Please generate natural-sounding voice data that reflects emotions.
[1261] In this way, the present invention enables communication that includes the emotions of hearing-impaired individuals in emergencies, and facilitates a rapid and appropriate emergency response.
[1262] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[1263] Step 1:
[1264] The user enters an emergency message into a text input field on their smartphone or robot and presses the send button. The entered text data contains emergency information, such as "Help, there's a fire." This input is then sent to the server via the input device.
[1265] Input: User-entered text data "Help, there's a fire"
[1266] Output: Text data sent to the server
[1267] Step 2:
[1268] The server passes the received text data to a natural language processing system, which analyzes the text's content. The natural language processing system understands the meaning from the input text and generates basic parameters for speech synthesis.
[1269] Input: Received text data
[1270] Data Processing / Calculation: Analyze text content using natural language processing techniques and generate speech synthesis parameters.
[1271] Output: Generated speech synthesis parameters
[1272] Step 3:
[1273] The natural language processing system passes text data to the emotion recognition system to recognize the user's emotions. Here, the text2emotion library is used to analyze the emotions implied in the text and identify the main emotions.
[1274] Input: Text data passed from a natural language processing system.
[1275] Data processing / calculation: Recognizing emotions using the text2emotion library.
[1276] Output: Recognized emotion information (Example: {"fear": 0.5, "sad": 0.3, "happy": 0.1, "angry": 0.1})
[1277] Step 4:
[1278] Based on the emotion information recognized by the emotion recognition means, the speech synthesis means generates speech data. Using the gTTS library, speech data with speech parameters that reflect the emotion is created.
[1279] Input: Sentimental information and speech synthesis parameters
[1280] Data processing / calculation: Generate audio data using the gTTS library.
[1281] Output: Generated audio data file
[1282] Step 5:
[1283] The generated audio data file is sent via HTTP request to the security guard or emergency service system, which is the recipient of the emergency notification. The requests library is used.
[1284] Input: Generated audio data file
[1285] Data processing / calculation: Data transmission via HTTP requests
[1286] Output: Completion of voice data transmission to security guards and emergency services.
[1287] Step 6:
[1288] Security and emergency service systems play back transmitted audio data and initiate a rapid response. The playback mechanism uses an audio output device to play back emergency messages that reflect emotions.
[1289] Input: Received audio data file
[1290] Data processing / calculation: Audio data playback
[1291] Output: Initiate emergency response and play voice message.
[1292] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[1293] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1294] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[1295] [Fourth Embodiment]
[1296] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[1297] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1298] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1299] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[1300] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[1301] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[1302] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[1303] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[1304] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[1305] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1306] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1307] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[1308] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1309] This system allows hearing-impaired individuals to input what they want to say as text, analyze that text to generate audio data, and then transmit it to hearing individuals. The configuration for implementing this system is described below.
[1310] System Overview
[1311] This system includes a terminal for user text input, a server that analyzes the input text and generates audio data, and a terminal that plays the generated audio data. The specific roles and processing flow of each component will be explained below.
[1312] User input and terminal operation
[1313] User: The user enters what they want to say in text format. For example, the user enters "Hello, how are you?" into the text box on their device.
[1314] Terminal: Provides a text input field, receives input from the user, and sends the content to the server when the submit button is pressed.
[1315] Server operation
[1316] Server: Receives text data sent from the terminal. Next, it analyzes the text using natural language processing and generates parameters for speech synthesis. For example, for the text "Hello, how are you?", it sets parameters such as language, pitch, and speed.
[1317] Natural language processing means: Analyzes the received text and derives appropriate speech synthesis parameters. Here, language models and sentiment analysis are used to determine parameters for generating more natural speech.
[1318] Speech synthesis means: Using parameters generated by natural language processing means, input text is converted into speech data. This speech data can then be converted into any desired audio format, such as a WAV file or MP3 file.
[1319] Transmission method: The generated audio data is sent back to the terminal. The terminal is configured to receive data via HTTP requests.
[1320] Device playback operation
[1321] Terminal: Receives audio data sent back from the server and plays it for the user. Uses software or devices with audio playback capabilities to transmit text input as audio.
[1322] Specific example
[1323] Input and Submission: The user enters "Hello, how are you?" into the text box and presses the submit button.
[1324] Text parsing: The server receives the text and converts "Hello, how are you?" into speech synthesis parameters.
[1325] Speech generation: A speech synthesis device converts the text "Hello, how are you?" into speech data.
[1326] Playback: The device receives and plays the audio data. The played audio contains the message, "Hello, how are you?"
[1327] This system enables people with hearing impairments to speak using their voices, facilitating smooth communication with hearing individuals. Details of each processing step are described below.
[1328] The following describes the processing flow.
[1329] Step 1: The user enters what they want to say in text format.
[1330] User: Enter "Hello, how are you?" into the text input field and click the send button.
[1331] Step 2: The terminal receives user input and sends the data to the server.
[1332] Terminal: Sends the entered text data "Hello, how are you?" to the server in POST request format.
[1333] Step 3: The server sends the received data to a natural language processing system for analysis.
[1334] Server: Passes the received text data to a natural language processing engine for analysis.
[1335] Step 4: The natural language processing engine analyzes the text data and generates speech synthesis parameters.
[1336] Server: The natural language processing engine analyzes the received text "Hello, how are you?" and generates parameters for speech synthesis (language, pitch, speed, etc.).
[1337] Step 5: Generate audio data using a speech synthesis method based on the analysis results.
[1338] Server: The speech synthesis engine generates speech data using parameters. The generated speech data contains the message, "Hello, how are you?"
[1339] Step 6: Send the generated audio data to the device.
[1340] Server: Sends the generated audio data to the terminal.
[1341] Step 7: The device plays the audio data received from the server.
[1342] Terminal: Passes the audio data received from the server to the playback function, which then plays the audio.
[1343] Step 8: The user communicates using the generated voice.
[1344] User: Uses the audio "Hello, how are you?" played from the device to communicate with a healthy person. During this process, the user confirms that their words have been transmitted as audio.
[1345] (Example 1)
[1346] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1347] In recent years, methods such as writing notes and using tablet devices have become common ways for people with hearing impairments to communicate smoothly with those without hearing impairments. However, these methods make real-time communication difficult, which is particularly inconvenient in emergencies or situations requiring a quick response. There is a need for a system that allows people with hearing impairments to communicate smoothly using voice.
[1348] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[1349] In this invention, the server includes an input means for receiving user input, a natural language processing means for analyzing the text data received by the input means and generating parameters for speech synthesis, a speech synthesis means for generating speech data using the parameters generated by the natural language processing means, a transmission means for transmitting the speech data generated by the speech synthesis means to the user's terminal, and a playback means for playing the speech data on the user's terminal. This makes it possible for people with hearing impairments to quickly communicate their intentions through speech.
[1350] "An input means for receiving user input" refers to a device or system for users to input what they want to say in text format.
[1351] "Natural language processing means" refers to an algorithm or software that analyzes text data received by an input means and generates parameters for speech synthesis.
[1352] "Speech synthesis means" refers to an algorithm or software that converts input text into speech data using parameters generated by natural language processing means.
[1353] "Transmission means" refers to a device or system for transmitting audio data generated by speech synthesis means to a user's terminal.
[1354] "Playback means" refers to a device or software for playing audio data on the user's terminal.
[1355] "Adjustment means" refers to a device or software used to adjust parameters such as language, pitch, speed, and emotion when generating audio data.
[1356] A "terminal" is a device that allows a user to input text using an input device, send it to a server, and receive and play back the generated audio data.
[1357] This invention relates to a system for enabling hearing-impaired individuals to input what they wish to express as text, analyze that text to generate audio data, and communicate it to hearing individuals. The system includes a terminal for user text input, a server that analyzes the input text and generates audio data, and a terminal that plays the generated audio data.
[1358] System Overview
[1359] This system consists of the following elements:
[1360] 1. Input means for receiving user input
[1361] Users enter what they want to say in text format using an application on their smartphone or computer, or an input field on a webpage. For example, they might type, "Hello, how are you?"
[1362] 2. Transmission of text data via input means
[1363] Once the user completes the input and presses the submit button, the device sends this text data to the server. An HTTP POST request is used for submission.
[1364] 3. Text analysis by the server
[1365] The server receives text data sent from the terminal and analyzes it using natural language processing tools. This analysis uses language models such as Google's TensorFlow and OpenAI's GPT-3. It analyzes the sentiment and grammar of the received text and generates optimal parameters for speech synthesis.
[1366] 4. Generating speech synthesis parameters
[1367] The server's natural language processing system generates speech synthesis parameters such as pitch, speed, and emotion based on the analysis results. For example, it extracts appropriate intonation and emotion from the text "Hello, how are you?".
[1368] 5. Server-driven generation of audio data
[1369] Based on parameters generated by natural language processing, the server uses speech synthesis to generate actual audio data. This speech synthesis utilizes services such as Amazon Polly or Google Cloud Text-to-Speech. The generated audio data is saved as WAV or MP3 files.
[1370] 6. Sending and playing audio data
[1371] The server sends the generated audio data to the terminal as an HTTP response. The terminal plays the received audio data using a playback device. Playback requires HTML5. <audio>Tags or dedicated audio playback applications are used.
[1372] Specific example
[1373] For example, the process would be as follows:
[1374] Input and Sending: The user types "Hello, how are you?" into the text box on the device and presses the send button.
[1375] Text analysis: The server receives text, analyzes it, and converts it into speech synthesis parameters. The natural language processing model used is OpenAI's GPT-3, among others.
[1376] Speech generation: The server uses Amazon Polly to convert the text "Hello, how are you?" into speech data.
[1377] Playback: The device receives audio data and HTML5 <audio>Play using tags.
[1378] Example of a prompt
[1379] "Please use GPT-3 to convert the text 'Hello, how are you?' into speech synthesis parameters. Then, use a speech synthesis service to generate the speech data and send it back to the terminal."
[1380] This system enables hearing-impaired individuals to communicate quickly through voice. The specific processing steps at each stage are carried out based on the means described in the claims.
[1381] The flow of the specific processing in Example 1 will be explained using Figure 11.
[1382] Step 1:
[1383] User: The user enters what they want to say in text format using a smartphone or computer. For example, they might type, "Hello, how are you?"
[1384] Input: Text "Hello, how are you?"
[1385] Output: Text data entered into the user's terminal
[1386] Step 2:
[1387] Terminal: When the user presses the send button, the terminal sends the entered text data to the server. An HTTP POST request is used for transmission. Specifically, the text data is sent to the server in JSON format.
[1388] Input: Text data "Hello, how are you?"
[1389] Output: JSON data sent as an HTTP POST request
[1390] Step 3:
[1391] Server: The server receives an HTTP POST request and extracts text data. The received text data is then passed to a natural language processing system.
[1392] Input: JSON data sent to the server
[1393] Output: Text data passed to the natural language processing system.
[1394] Step 4:
[1395] Server: The natural language processing system analyzes text data and generates parameters for speech synthesis (e.g., language, pitch, speed, emotion, etc.). Examples of natural language processing models used include TensorFlow and GPT-3. For example, it performs emotion analysis on the text "Hello, how are you?" and generates appropriate intonation parameters.
[1396] Input: Text data "Hello, how are you?"
[1397] Output: Speech synthesis parameters (language, pitch, speed, emotion, etc.)
[1398] Step 5:
[1399] Server: Using the generated parameters, the speech synthesis system generates speech data. This system may use Amazon Polly or Google Cloud Text-to-Speech. For example, based on the parameters, speech data for "Hello, how are you?" is generated.
[1400] Input: Speech synthesis parameters
[1401] Output: Audio data (such as WAV or MP3 files)
[1402] Step 6:
[1403] Server: Sends the generated audio data to the terminal as an HTTP response. At this time, the appropriate MIME type (e.g., audio / mpeg) is set in the response header.
[1404] Input: Audio data
[1405] Output: HTTP response sent to the terminal
[1406] Step 7:
[1407] Terminal: The terminal analyzes the received audio data and plays it back using its audio playback function. Specifically, it uses HTML5 <audio>Tags or dedicated audio playback applications are used. For example, received audio data is played back to the user as "Hello, how are you?".
[1408] Input: Audio data received as an HTTP response
[1409] Output: Audio played to the user
[1410] This process allows hearing-impaired individuals to input what they want to say as text and transmit it to hearing individuals as audio.
[1411] (Application Example 1)
[1412] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1413] People with hearing impairments often have difficulty communicating with store staff in physical stores. Therefore, support is needed in situations where smooth communication is required, such as when asking for product information or completing a purchase. Furthermore, existing systems have limited capabilities for converting text data to audio data, making them unsuitable for use in physical stores.
[1414] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[1415] In this invention, the server includes an input means for receiving user input, a natural language processing means for analyzing the text data received by the input means and generating parameters for speech synthesis, and a speech synthesis means for generating speech data using the parameters generated by the natural language processing means. This makes it possible to generate speech data and transmit it to other people in a physical store in order to support communication between people with hearing impairments and other people in the store.
[1416] A "user" is a person who uses this system to input text and convert that text into speech.
[1417] "Input means" refers to a device or function for a user to input text data. For example, this includes terminals such as smartphones and smart glasses.
[1418] "Natural language processing means" refers to a device or program that has the function of analyzing text data received by an input means and generating parameters for speech synthesis.
[1419] "Speech synthesis means" refers to a device or function that generates speech data using parameters generated by natural language processing means.
[1420] "Transmission means" refers to a device or function for transmitting generated audio data to the user's terminal.
[1421] "Playback means" refers to a device or function for playing audio data on a user terminal.
[1422] "A means of generating audio data and transmitting it to other people in a physical store to support communication between people with hearing impairments and other people in a physical store" refers to a means of converting text data entered by a person with hearing impairment into audio data and transmitting that audio to other people in a physical store.
[1423] This system is designed to facilitate communication between people with hearing impairments and other people in physical stores. A specific implementation of this system is described below.
[1424] System Configuration
[1425] User: A person with a hearing impairment performs text input.
[1426] Terminal: Provides a text input field and accepts text input from the user. This includes devices such as smartphones and smart glasses.
[1427] Server: Analyzes text data sent from the terminal, generates speech synthesis parameters, and creates speech data.
[1428] Natural language processing means: Analyzes text data and derives appropriate speech synthesis parameters.
[1429] Speech synthesis means: Generates speech data based on parameters.
[1430] Transmission method: The generated audio data is sent to the terminal.
[1431] Playback method: The audio data is played on the device and conveyed to other people, such as store clerks.
[1432] Hardware and software used
[1433] Hardware: Smartphones, smart glasses, servers.
[1434] Software: Python, Google Text-to-Speech (gTTS) library, Playsound library.
[1435] Processing flow and data manipulation / calculation
[1436] 1. The user enters text.
[1437] Enter text into the input field on your smartphone or smart glasses.
[1438] 2. The terminal sends text data to the server.
[1439] The text data is sent to the server in the format of an HTTP POST request.
[1440] 3. The server analyzes the text data and generates parameters.
[1441] The system analyzes text using natural language processing techniques and generates parameters such as language, pitch, and speed.
[1442] 4. The server generates the audio data.
[1443] Speech data is generated using parameters produced by a speech synthesis system.
[1444] 5. The server sends the audio data to the terminal.
[1445] The generated audio data is sent to the device.
[1446] 6. The device plays the audio data.
[1447] Audio data is played back using a playback device (smartphone or smart glasses) to convey the intentions of the hearing-impaired person to store staff, etc.
[1448] Specific example
[1449] For example, suppose a hearing-impaired person enters "Where is this product?" into smart glasses in a physical store. This text input is sent to a server, which analyzes the text and generates audio data saying "Where is this product?". This generated audio data is sent to the smart glasses and played back, allowing the store clerk to understand that the question is from a hearing-impaired person and guide them to the product's location.
[1450] Example of a prompt
[1451] Prompt: The user typed "Where can I find this product?". Convert this text into audio data and return the URL.
[1452] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[1453] Step 1:
[1454] The user enters text.
[1455] Input: A person with a hearing impairment types "Where can I find this product?" into the text input field on their smartphone or smart glasses.
[1456] Action: The user types text into the input field and presses the submit button.
[1457] Output: The entered text data is recorded on the terminal.
[1458] Step 2:
[1459] The terminal sends text data to the server.
[1460] Input: Text data entered by the user.
[1461] Operation: The device sends text data to the server using an HTTP POST request.
[1462] Output: The server receives text data.
[1463] Step 3:
[1464] The server analyzes the text data and generates parameters.
[1465] Input: Text data received by the server: "Where is this product located?"
[1466] Operation: The server's natural language processing system analyzes the text and generates speech synthesis parameters such as language, pitch, and speed.
[1467] Output: Generated speech synthesis parameters.
[1468] Step 4:
[1469] The server generates the audio data.
[1470] Input: Speech synthesis parameters generated by the server.
[1471] Operation: The server's speech synthesis system generates speech data based on the parameters.
[1472] Output: Generated audio data (e.g., MP3 file).
[1473] Step 5:
[1474] The server sends the audio data to the terminal.
[1475] Input: Generated audio data.
[1476] Operation: The server's transmission method sends the generated audio data to the terminal in HTTP response format.
[1477] Output: The device receives the audio data.
[1478] Step 6:
[1479] The device plays audio data
[1480] Input: Received audio data.
[1481] Operation: The device's playback mechanism plays the received audio data. Audio is played from the smartphone or smart glasses, and the store clerk receives the question from the hearing-impaired person.
[1482] Output: The audio is played in the physical store, allowing the intentions of the hearing-impaired person to be conveyed to others.
[1483] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[1484] This invention combines an emotion recognition function with a system that allows hearing-impaired individuals to input what they want to say as text, analyzes that text to generate audio data, and transmits it to hearing individuals. This allows the system to recognize the user's emotions from the input text data and reflect those emotions during speech synthesis.
[1485] System Overview
[1486] This system includes a terminal for user text input, a server that analyzes the input text and generates audio data, and a terminal that plays the generated audio data. In particular, it incorporates an emotion engine that recognizes the user's emotions from the text data.
[1487] User input and terminal operation
[1488] User: The user enters what they want to say in text format. For example, the user enters "Hello, how are you?" into the text box on their device.
[1489] Terminal: Provides a text input field, receives input from the user, and sends the content to the server when the submit button is pressed.
[1490] Server operation
[1491] Server: Receives text data sent from the terminal. Next, natural language processing means analyze the text and generate parameters for speech synthesis. In particular, emotion recognition means recognize the user's emotions from the text data and reflect that emotion information in the speech synthesis parameters.
[1492] Natural language processing means: Analyzes received text data and uses emotion recognition means to analyze the user's emotions. For example, it recognizes a positive emotion in response to the text "Hello, how are you?". Based on this, it sets parameters to make the voice sound brighter.
[1493] Emotion Engine: Recognizes user emotions from text data and generates emotion classification results. These results are sent to a speech synthesis system and used to generate speech data with natural-sounding expressions.
[1494] Speech synthesis means: Using parameters generated by emotion recognition means, input text is converted into speech data. Pitch, speed, intonation, etc., are added, adjusted based on emotion.
[1495] Transmission method: The generated audio data is sent back to the terminal. The terminal is configured to receive data via HTTP requests.
[1496] Device playback operation
[1497] Terminal: Receives audio data sent back from the server and plays it for the user. Using software or devices with audio playback capabilities, it transmits text input as audio. Because the audio data reflects the user's emotions, more natural communication is possible.
[1498] Specific example
[1499] Input and Submission: The user enters "Hello, how are you?" into the text box and presses the submit button.
[1500] Text analysis and sentiment recognition: The server receives text, analyzes it to see if it says "Hello, how are you?", and recognizes a positive sentiment.
[1501] Speech generation: The speech synthesis system sets parameters based on emotion recognition and converts the text "Hello, how are you?" into speech data in a cheerful tone.
[1502] Playback: The device receives and plays the audio data. The played audio conveys the message, "Hello, how are you?" with a positive tone.
[1503] This system enables hearing-impaired individuals to not only speak using their voices but also to convey their emotions, thereby achieving richer communication. Details of each processing step will be explained next.
[1504] The following describes the processing flow.
[1505] Step 1: The user enters what they want to say in text format.
[1506] User: Enter "Hello, how are you?" into the text input field and click the send button.
[1507] Step 2: The terminal receives user input and sends the data to the server.
[1508] Terminal: Sends the entered text data "Hello, how are you?" to the server in POST request format.
[1509] Step 3: The server sends the received data to a natural language processing system for analysis.
[1510] Server: Passes the received text data to a natural language processing engine for analysis.
[1511] Step 4: The natural language processing engine analyzes the text data and sends it to the sentiment recognition engine.
[1512] Server: The natural language processing engine analyzes the received text "Hello, how are you?" and passes the result to the emotion recognition engine.
[1513] Step 5: The emotion recognition engine extracts emotions from the text and generates emotional information.
[1514] Server: The emotion recognition engine extracts positive emotions from the text "Hello, how are you?" and generates emotion information.
[1515] Step 6: Set the speech synthesis parameters along with the emotional information.
[1516] Server: Based on the emotion information obtained from the emotion recognition engine, it sets the parameters used by the speech synthesis engine (e.g., pitch, speed, intonation).
[1517] Step 7: The speech synthesis engine generates speech data that reflects the emotion parameters.
[1518] Server: The speech synthesis engine converts the text "Hello, how are you?" into voice data with a cheerful tone, based on the configured parameters.
[1519] Step 8: Send the generated audio data to the device.
[1520] Server: Prepares to send the generated audio data to the terminal and returns it via an HTTP response.
[1521] Step 9: Play the audio data received by the device.
[1522] Terminal: Passes the audio data received from the server to the playback function, and plays the audio.
[1523] Step 10: The user communicates using the generated voice.
[1524] User: Uses a voice message containing positive emotions, such as "Hello, how are you?", played from the device to communicate with a healthy person.
[1525] This detailed processing step generates audio data that reflects not only the text content entered by the user, but also their emotions, thereby supporting more effective communication with people without disabilities.
[1526] (Example 2)
[1527] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1528] Conventional speech synthesis systems can convert text input by hearing-impaired individuals into speech data, but they cannot reflect the user's emotions, thus limiting their communication capabilities. In particular, the inability to convey the emotions behind speech makes accurate and rich communication difficult. Solving this problem is essential.
[1529] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[1530] In this invention, the server includes an input means for receiving user input, a natural language processing means for analyzing the text data received by the input means, recognizing emotions, and generating parameters for speech synthesis, and a speech synthesis means for generating speech data using the parameters based on the emotions recognized by the natural language processing means. This makes it possible to recognize emotions from text data entered by a hearing-impaired person and generate speech data that reflects those emotions.
[1531] An "input means for receiving user input" is an interface for users to input what they want to say in text format, and a device that has the function of sending the entered text to the system.
[1532] "Natural language processing means" refers to technology for analyzing user input and understanding its content, and in particular, a device that has the function of recognizing emotions and generating parameters for speech synthesis.
[1533] An "emotion recognition device" is a device that analyzes the user's emotions from input text data and classifies those emotions into specific categories.
[1534] A "speech synthesis means" is a technology that converts input text into speech data using parameters generated by a natural language processing means, and is a device that has the function of reflecting emotion-based adjustments.
[1535] A "transmission means" is a device that has the function of transmitting the generated audio data to the user's terminal.
[1536] "Playback means" refers to a device or software for playing back audio data received on a user terminal.
[1537] A "modification device" is a device that has the function of adjusting parameters such as language, pitch, speed, and intonation based on emotion when generating audio data.
[1538] A "terminal" is a device used by a user to input and send text, and is a device that has the function of sending data to a server in the format of an HTTP request.
[1539] This invention combines an emotion recognition function with a system that allows hearing-impaired individuals to input what they want to say as text, analyzes that text to generate audio data, and transmits it to hearing individuals. This allows the system to recognize the user's emotions from the input text data and reflect those emotions during speech synthesis.
[1540] This system includes a terminal for user text input, a server that analyzes the input text and generates audio data, and a terminal that plays the generated audio data. In particular, it incorporates an emotion engine that recognizes the user's emotions from the text data.
[1541] User input and terminal operation
[1542] User:
[1543] The user enters what they want to say in text format. For example, the user might type "Hello, how are you?" into the text box on their device.
[1544] Terminal:
[1545] The terminal provides a text input field, receives input from the user, and sends the content to the server when the submit button is pressed. Specifically, it sends data to the server using an HTTP request.
[1546] Server operation
[1547] server:
[1548] The server receives text data sent from the terminal. The received text data is passed to a natural language processing system for analysis. Next, an emotion engine is used to recognize the user's emotions from the text data.
[1549] Natural language processing methods:
[1550] This technology analyzes received text data to understand the meaning of what the user is saying. For example, it analyzes and understands the meaning of the text, "Hello, how are you?"
[1551] Emotional engine:
[1552] Based on the analysis results from natural language processing tools, the system recognizes the user's emotions. For example, it can recognize positive emotions from the text "Hello, how are you?".
[1553] Speech synthesis method:
[1554] Using parameters generated by emotion recognition, the input text is converted into audio data. Pitch, speed, intonation, and other parameters are adjusted. For example, audio data is generated with a bright tone that reflects positive emotions.
[1555] Transmission method:
[1556] The speech synthesis system sends the generated audio data back to the terminal. The audio data is transmitted using an HTTP response.
[1557] Device playback operation
[1558] Terminal:
[1559] The device receives the audio data sent back from the server and uses its audio playback function to play the audio data for the user. This allows the entered text content to be played back as audio, reflecting emotions. Specifically, it might say, "Hello, how are you?" with positive emotions.
[1560] Specific example
[1561] Input and submission:
[1562] User: Type "Hello, how are you?" into the text box and press the send button.
[1563] Text analysis and sentiment recognition:
[1564] Server: Receives and analyzes the message "Hello, how are you?" and recognizes positive emotions.
[1565] Speech generation:
[1566] Server: The speech synthesis system sets parameters to match positive emotions and converts "Hello, how are you?" into voice data in a cheerful tone.
[1567] reproduction:
[1568] Terminal: Receives audio data and plays it back using its audio playback function. The played audio conveys the message, "Hello, how are you?" along with a positive tone.
[1569] Concrete examples of prompt sentences for generative AI models
[1570] Prompt example:
[1571] "Analyze the following text data to recognize the user's emotions. Generate emotion-based speech synthesis parameters and then generate the speech data."
[1572] This system allows emotions to be reflected in the text entered by people with hearing impairments, enabling more natural and enriching communication.
[1573] The flow of the specific processing in Example 2 will be explained using Figure 13.
[1574] Step 1:
[1575] User:
[1576] Inputs and outputs:
[1577] Input: The user enters "Hello, how are you?" into the text input field on the device.
[1578] Output: The entered text "Hello, how are you?" is stored in the terminal.
[1579] Specific actions:
[1580] The user enters what they want to say into the text input field on their device and presses the send button.
[1581] Step 2:
[1582] Terminal:
[1583] Inputs and outputs:
[1584] Input: The text entered in Step 1: "Hello, how are you?"
[1585] Output: Text data is sent to the server as an HTTP request.
[1586] Specific actions:
[1587] The terminal retrieves the contents of the text input field and sends it to the server using an HTTP POST request.
[1588] Step 3:
[1589] server:
[1590] Inputs and outputs:
[1591] Input: Text data sent from the device: "Hello, how are you?"
[1592] Output: Text data is received by the server.
[1593] Specific actions:
[1594] The server receives a POST request sent from the terminal and extracts text data from the request body.
[1595] Step 4:
[1596] server:
[1597] Inputs and outputs:
[1598] Input: Text data received in Step 3: "Hello, how are you?"
[1599] Output: Analysis results of text data (meaning and content of the statement)
[1600] Specific actions:
[1601] The server passes text data to a natural language processing system, which then analyzes the text and understands its meaning.
[1602] Step 5:
[1603] server:
[1604] Inputs and outputs:
[1605] Input: Meaning of the text data "Hello, how are you?" analyzed in Step 4
[1606] Output: Emotion recognition result (e.g., positive)
[1607] Specific actions:
[1608] The server uses an emotion engine to recognize the user's emotions from the analysis results. For example, it recognizes positive emotions from "Hello, how are you?".
[1609] Step 6:
[1610] server:
[1611] Inputs and outputs:
[1612] Input: Emotions recognized in Step 5 and text data analyzed in Step 4
[1613] Output: Speech synthesis parameters (e.g., bright tone, appropriate speed and intonation)
[1614] Specific actions:
[1615] The server sets parameters for generating speech data using speech synthesis based on the emotion recognition results. For example, it sets parameters for a bright tone based on positive emotions.
[1616] Step 7:
[1617] server:
[1618] Inputs and outputs:
[1619] Input: Speech synthesis parameters generated in step 6
[1620] Output: Audio data (Example: "Hello, how are you?" in a cheerful tone)
[1621] Specific actions:
[1622] The server uses speech synthesis to convert the input text into speech data based on the configured parameters.
[1623] Step 8:
[1624] server:
[1625] Inputs and outputs:
[1626] Input: Audio data generated in Step 7
[1627] Output: The audio data is sent to the terminal as an HTTP response.
[1628] Specific actions:
[1629] The server sends the generated audio data back to the terminal as an HTTP response.
[1630] Step 9:
[1631] Terminal:
[1632] Inputs and outputs:
[1633] Input: Audio data sent from the server
[1634] Output: The voice played by the user (e.g., "Hello, how are you?" in a cheerful tone)
[1635] Specific actions:
[1636] The device receives audio data sent from the server and plays it back using its audio playback function. Because the played audio reflects the user's emotions, more natural communication is possible.
[1637] (Application Example 2)
[1638] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1639] Traditionally, there has been a lack of adequate means for hearing-impaired individuals to quickly generate appropriate voice messages, including emotions, in emergencies and notify security guards and emergency services. This has led to delays in emergency communication and the potential for misunderstandings. This invention aims to solve these problems by recognizing emotions based on text input and generating voice data that includes those emotions to notify emergency services.
[1640] In Application Example 2, the specific processing performed by the specific processing unit 290 of the data processing device 12 is realized by the following means. In this invention, the server includes an input means for receiving user input, a natural language processing means for analyzing the received text data and generating parameters for speech synthesis, a speech synthesis means for generating speech data using the parameters generated by the natural language processing means, an emotion recognition means for recognizing emotions from the text data using the natural language processing means and reflecting that emotion information in the speech synthesis parameters, a transmission means for transmitting the generated speech data to a user terminal, a playback means for playing the speech data on the user terminal, and a notification means for analyzing the input text data in an emergency, generating a speech message including emotions, and notifying security guards or emergency services. This makes it possible for hearing-impaired individuals to quickly and accurately generate a speech message including emotions in an emergency and to respond appropriately to the emergency.
[1641] "Input means" refers to a device or interface for a user to input text data.
[1642] "Natural language processing means" refers to systems and algorithms that analyze received text data and generate parameters for speech synthesis.
[1643] "Speech synthesis means" refers to a device or software that generates speech data using parameters generated by natural language processing means.
[1644] "Emotion recognition means" refers to a system or algorithm that recognizes a user's emotions from text data and reflects that emotional information in speech synthesis parameters.
[1645] "Transmission means" refers to the communication means used to send the generated audio data to the user's terminal.
[1646] "Playback means" refers to a device or software that plays back audio data transmitted to the user's terminal.
[1647] "Notification means" refers to a mechanism or system that analyzes text data entered in an emergency, generates an audio message including emotions, and notifies security guards or emergency services.
[1648] The system according to the present invention is a system for enabling hearing-impaired individuals to quickly and accurately generate emotionally charged voice messages in emergencies and notify security guards and emergency services. This system mainly consists of an input means, a natural language processing means, a speech synthesis means, an emotion recognition means, a transmission means, a playback means, and a notification means.
[1649] Hardware and software configuration
[1650] Hardware used: Smartphone or robot (with voice output function)
[1651] Software used:
[1652] text2emotion: Emotion Recognition Library
[1653] gTTS (Google Text-to-Speech): A speech synthesis library.
[1654] requests: A library for sending HTTP requests.
[1655] Operational description of each means
[1656] 1. Input method:
[1657] This provides an interface for users to enter text data in emergencies. This interface will be implemented as a text input field on smartphones or robots. Users will enter emergency messages here.
[1658] 2. Natural language processing methods:
[1659] The system analyzes text data received through the input means and generates parameters for speech synthesis. This step is crucial as it involves understanding the content of the text data and converting it into speech data.
[1660] 3. Speech synthesis means:
[1661] Speech data is generated using parameters produced by natural language processing. During this process, pitch, speed, intonation, and other parameters are adjusted to produce natural-sounding speech that also conveys emotion.
[1662] 4. Emotion recognition means:
[1663] The system analyzes text data received from a natural language processing system to recognize the user's emotions. This recognized emotion information is then reflected in the parameters of the speech synthesis system. The text2emotion library is used for this emotion recognition.
[1664] 5. Transmission method:
[1665] This is a communication method for sending generated audio data to the user's device. This method uses the requests library to send audio data in HTTP request format.
[1666] 6. Regeneration means:
[1667] This is a device or software that plays back audio data transmitted to a user's terminal. It provides audio output on smartphones and robots.
[1668] 7. Means of notification:
[1669] In emergencies, the system analyzes entered text data, generates an emotionally charged voice message, and immediately notifies security personnel or emergency services. This notification is sent to a server by transmitting the generated voice data.
[1670] Specific example
[1671] User actions:
[1672] The user types "Help me, there's a fire" into the text input field of their smartphone or robot. Then they press the send button.
[1673] Server operation:
[1674] The server analyzes the received text data using natural language processing and recognizes the emotion of fear from the content of "Fire!". As a result, it generates audio data that includes the emotion of urgency and fear using speech synthesis.
[1675] Speech synthesis:
[1676] The speech synthesis means uses speech parameters that reflect the emotion of fear recognized by the emotion recognition means to generate an urgent voice message saying, "Help me, there's a fire."
[1677] Emergency notification:
[1678] The generated audio data is transmitted to security guards and emergency services via a transmission device to facilitate a rapid response.
[1679] Example of a prompt
[1680] Towards emotion recognition AI:
[1681] Text: "Help me, there's a fire."
[1682] Please return the emotion recognition results in JSON format (e.g., {"happy": 0.1, "sad": 0.3, "angry": 0.1, "fear": 0.5}).
[1683] Towards speech synthesis AI:
[1684] Text: "Help me, there's a fire."
[1685] Recognized emotion: fear
[1686] Please generate natural-sounding voice data that reflects emotions.
[1687] In this way, the present invention enables communication that includes the emotions of hearing-impaired individuals in emergencies, and facilitates a rapid and appropriate emergency response.
[1688] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[1689] Step 1:
[1690] The user enters an emergency message into a text input field on their smartphone or robot and presses the send button. The entered text data contains emergency information, such as "Help, there's a fire." This input is then sent to the server via the input device.
[1691] Input: User-entered text data "Help, there's a fire"
[1692] Output: Text data sent to the server
[1693] Step 2:
[1694] The server passes the received text data to a natural language processing system, which analyzes the text's content. The natural language processing system understands the meaning from the input text and generates basic parameters for speech synthesis.
[1695] Input: Received text data
[1696] Data Processing / Calculation: Analyze text content using natural language processing techniques and generate speech synthesis parameters.
[1697] Output: Generated speech synthesis parameters
[1698] Step 3:
[1699] The natural language processing system passes text data to the emotion recognition system to recognize the user's emotions. Here, the text2emotion library is used to analyze the emotions implied in the text and identify the main emotions.
[1700] Input: Text data passed from a natural language processing system.
[1701] Data processing / calculation: Recognizing emotions using the text2emotion library.
[1702] Output: Recognized emotion information (Example: {"fear": 0.5, "sad": 0.3, "happy": 0.1, "angry": 0.1})
[1703] Step 4:
[1704] Based on the emotion information recognized by the emotion recognition means, the speech synthesis means generates speech data. Using the gTTS library, speech data with speech parameters that reflect the emotion is created.
[1705] Input: Sentimental information and speech synthesis parameters
[1706] Data processing / calculation: Generate audio data using the gTTS library.
[1707] Output: Generated audio data file
[1708] Step 5:
[1709] The generated audio data file is sent via HTTP request to the security guard or emergency service system, which is the recipient of the emergency notification. The requests library is used.
[1710] Input: Generated audio data file
[1711] Data processing / calculation: Data transmission via HTTP requests
[1712] Output: Completion of voice data transmission to security guards and emergency services.
[1713] Step 6:
[1714] Security and emergency service systems play back transmitted audio data and initiate a rapid response. The playback mechanism uses an audio output device to play back emergency messages that reflect emotions.
[1715] Input: Received audio data file
[1716] Data processing / calculation: Audio data playback
[1717] Output: Initiate emergency response and play voice message.
[1718] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[1719] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1720] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[1721] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1722] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[1723] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[1724] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[1725] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[1726] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[1727] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[1728] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[1729] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[1730] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[1731] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1732] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[1733] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[1734] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[1735] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[1736] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[1737] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[1738] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted as being incorporated by reference.
[1739] The following is further disclosed regarding the embodiments described above.
[1740] (Claim 1)
[1741] An input means for receiving user input,
[1742] A natural language processing means that analyzes the text data received by the input means and generates parameters for speech synthesis,
[1743] A speech synthesis means that generates speech data using parameters generated by the natural language processing means,
[1744] A transmission means for transmitting the voice data generated by the voice synthesis means to the user's terminal,
[1745] A playback means for playing the aforementioned audio data on a user terminal,
[1746] A system that includes this.
[1747] (Claim 2)
[1748] The system according to claim 1, further comprising adjustment means for adjusting parameters such as language, pitch, and speed when generating the aforementioned audio data.
[1749] (Claim 3)
[1750] The system according to claim 1, wherein the input means is a terminal that provides a user with a text input field and sends the entered text to the server in POST request format.
[1751] "Example 1"
[1752] (Claim 1)
[1753] An input means for receiving user input,
[1754] A natural language processing means that analyzes the text data received by the input means and generates parameters for speech synthesis,
[1755] A speech synthesis means that generates speech data using parameters generated by the natural language processing means,
[1756] A transmission means for transmitting the voice data generated by the voice synthesis means to the user's terminal,
[1757] A playback means for playing the aforementioned audio data on a user terminal,
[1758] A system that includes this.
[1759] (Claim 2)
[1760] The system according to claim 1, further comprising adjustment means for adjusting parameters such as language, pitch, speed, and emotion when generating the aforementioned audio data.
[1761] (Claim 3)
[1762] The system according to claim 1, wherein the input means is a terminal that provides a text input field to a user and sends the entered text to a server in the form of an HTTP POST request.
[1763] "Application Example 1"
[1764] (Claim 1)
[1765] An input means for receiving user input,
[1766] A natural language processing means that analyzes the text data received by the input means and generates parameters for speech synthesis,
[1767] A speech synthesis means that generates speech data using parameters generated by the natural language processing means,
[1768] A transmission means for transmitting the voice data generated by the voice synthesis means to the user's terminal,
[1769] A playback means for playing the aforementioned audio data on a user terminal,
[1770] To support communication between people with hearing impairments and other people in physical stores, a means of generating audio data and transmitting it to other people in the store is needed.
[1771] A system that includes this.
[1772] (Claim 2)
[1773] The system according to claim 1, further comprising adjustment means for adjusting parameters such as language, pitch, and speed when generating the aforementioned audio data.
[1774] (Claim 3)
[1775] The system according to claim 1, wherein the input means is a terminal that provides a user with a text input field and sends the entered text to the server in POST request format.
[1776] "Example 2 of combining an emotion engine"
[1777] (Claim 1)
[1778] An input means for receiving user input,
[1779] A natural language processing means that analyzes the text data received by the input means, recognizes emotions, and generates parameters for speech synthesis,
[1780] A speech synthesis means that generates speech data using parameters based on emotions recognized by the natural language processing means,
[1781] A transmission means for transmitting the voice data generated by the voice synthesis means to the user's terminal,
[1782] A playback means for playing the aforementioned audio data on a user terminal,
[1783] A system that includes this.
[1784] (Claim 2)
[1785] The system according to claim 1, further comprising adjustment means for adjusting parameters such as language, pitch, speed, and intonation based on emotion when generating the aforementioned audio data.
[1786] (Claim 3)
[1787] The system according to claim 1, wherein the input means is a terminal that provides a text input field to a user and sends the entered text to a server in HTTP request format.
[1788] "Application example 2 when combining with an emotional engine"
[1789] (Claim 1)
[1790] An input means for receiving user input,
[1791] A natural language processing means that analyzes the text data received by the input means and generates parameters for speech synthesis,
[1792] A speech synthesis means that generates speech data using parameters generated by the natural language processing means,
[1793] The aforementioned natural language processing means recognizes emotions from text data and reflects that emotional information in speech synthesis parameters.
[1794] A transmission means for transmitting the voice data generated by the voice synthesis means to the user's terminal,
[1795] A playback means for playing the aforementioned audio data on a user terminal,
[1796] A notification system that analyzes text data entered in an emergency, generates an audio message including emotions, and notifies security guards or emergency services.
[1797] A system that includes this.
[1798] (Claim 2)
[1799] The system according to claim 1, further comprising adjustment means for adjusting calculation parameters such as language, pitch, and speed when generating the aforementioned audio data.
[1800] (Claim 3)
[1801] The system according to claim 1, wherein the input means is a terminal that provides a text input field to a user and sends the entered text data to a server in POST request format. [Explanation of Symbols]
[1802] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / audio> < / audio> < / audio> < / url:> < / audio> < / audio> < / audio> < / url:> < / audio> < / audio> < / audio> < / url:> < / audio> < / audio> < / audio>
Claims
1. An input means for receiving user input, A natural language processing means that analyzes the text data received by the input means and generates parameters for speech synthesis, A speech synthesis means that generates speech data using parameters generated by the natural language processing means, A transmission means for transmitting the voice data generated by the voice synthesis means to the user's terminal, A playback means for playing the aforementioned audio data on a user terminal, A system that includes this.
2. The system according to claim 1, further comprising adjustment means for adjusting parameters such as language, pitch, and speed when generating the aforementioned audio data.
3. The system according to claim 1, wherein the input means is a terminal that provides a text input field to a user and sends the entered text to a server in POST request format.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A