System
The system addresses privacy and noise issues in voice communication by converting silent speech into voice, ensuring privacy and quiet environments through high-resolution capture and AI-driven text-to-speech conversion.
Patent Information
- Application Number
- JP2024120602
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-25
- Publication Date
- 2026-02-05
AI Technical Summary
Voice communication in remote work environments and public places poses challenges in protecting privacy and maintaining quiet surroundings, making it difficult to exchange confidential information and maintain personal privacy.
A system that captures silent speech at high resolution, converts it into text, generates multiple utterance candidates, determines the most appropriate text, and converts it into voice data for transmission, using image processing and AI models to resemble the user's voice quality.
Enables users to communicate silently, protecting privacy and maintaining quiet surroundings while supporting smooth communication.
Smart Images

Figure 2026019193000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] In today's remote work environments and public places, voice communication poses many challenges in terms of protecting privacy and maintaining quiet surroundings. Furthermore, there are many situations where voice communication is difficult when away from home or in open spaces. This can make it difficult to exchange confidential information outside the workplace and protect personal privacy. The present invention aims to provide a system that allows users to speak silently, thereby realizing an environment where smooth communication can be carried out while protecting privacy and maintaining quiet surroundings. [Means for solving the problem]
[0005] The present invention provides a system including: a means for capturing a user's silent speech at high resolution; a means for transmitting the captured data to a server; a means for analyzing the transmitted data and converting it into text; a means for generating a plurality of utterance candidates based on the converted text; a means for comparing the generated utterance candidates to determine the most appropriate text; a means for converting the determined text into voice data; and a means for transmitting the voice data to the other party. The system may also include a means for image processing lip and tongue movement data of the user's silent speech, and a means for generating voice in real time that closely resembles the user's voice quality. This system enables users to make calls anywhere without speaking, protecting their privacy and maintaining quiet surroundings while supporting smooth communication.
[0006] "Silent speech" is a method in which a user speaks without making a sound, using only the movements of their lips and tongue.
[0007] "High-resolution capture means" refers to an image collection device or camera that records the movements of the user's lips and tongue in detail.
[0008] "Captured data" refers to image data of lip and tongue movements obtained from a user's silent speech.
[0009] "Means for transmitting to a server" refers to a device or software with a communication function for transmitting captured data to a server via a network.
[0010] "Means for parsing and converting to text" refers to image processing techniques and algorithms used to analyze the captured data and convert it to text data.
[0011] "Means for generating multiple utterance candidates" refers to AI models or software that generate multiple utterance candidates based on initial text data obtained through image processing, taking into account a variety of contexts.
[0012] "Means for determining the optimal text" refers to the algorithms and methods for selecting the most appropriate text from the multiple utterance candidates generated.
[0013] "Means for converting into voice data" refers to voice synthesis technology (e.g., text-to-speech (TTS) technology) for converting the determined text into voice data.
[0014] "Means for transmitting to the other party" refers to equipment or software with communication capabilities for transmitting the generated voice data to the other party's device.
[0015] "Means for image processing of lip and tongue movement data" refers to image processing technologies and algorithms for capturing images of the user's lip and tongue movements, analyzing them, and converting them into text data.
[0016] "Means for generating voice that closely resembles the user's voice quality in real time" refers to voice synthesis technology that generates voice that closely resembles the user's actual voice quality based on text data. [Brief explanation of the drawings]
[0017] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7]FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0018] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0019] First, the terms used in the following description will be explained.
[0020] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0021] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0022] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0023] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0024] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0025] [First embodiment]
[0026] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0027] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0028] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0029] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0030] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0031] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0032] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0033] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0034] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0035] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0036] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0037] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0038] The present invention relates to a system that converts a user's silent speech into text in real time, and generates and transmits voice data. This system exchanges data between the user, the terminal, and a server to complete the overall process.
[0039] Program processing explanation
[0040] When a user makes a silent speech, the system starts by capturing the movements of the user's lips and tongue. For example, consider a situation where a user wants to have a confidential conversation with a business partner at a cafe. In this case, the user speaks using only the movements of their mouth.
[0041] User behavior
[0042] The user puts on a speakerphone or headset and begins to speak silently, and the camera responds, capturing high-resolution video of the user's lip and tongue movements.
[0043] Device behavior
[0044] The device captures the user's speech in real time and transmits the data to a server, which encodes the captured data into an optimized format and sends it to the server over a secure communication channel.
[0045] Server Operation
[0046] The server decodes the received data and uses an image processing AI model to analyze the user's lip and tongue movements. This analyzed data is converted into text data. A generative AI model is then used to generate multiple utterance candidates based on the context of the conversation and the content before and after. These candidates are compared to determine the most suitable text. The determined text data is passed to a speech generation AI (TTS) model, which generates speech in real time that closely resembles the user's voice quality. The generated voice data is then encoded and sent to the device.
[0047] Specific examples
[0048] For example, suppose a user makes a silent utterance such as "How is the project progressing?" The user's lip and tongue movements are captured on the device and sent to the server. The server's image processing AI analyzes this movement and generates the text candidate "How is the project progressing?" The generative AI model then generates utterance candidates that capture the context of the conversation and determines the same text as the best candidate. Based on this text, the speech generation AI generates speech that closely resembles the user's voice quality and sends it to the other party.
[0049] The person on the other end of the line will feel as if the user is silently asking, "How's the project going?" This system allows users to maintain privacy and communicate smoothly while maintaining the surrounding silence.
[0050] The processing flow will be explained below.
[0051] Step 1:
[0052] The user initiates silent speech: The user silently initiates speech using lip and tongue movements.
[0053] Step 2:
[0054] The device activates its built-in camera to capture high-resolution video of the user's lip and tongue movements, with the camera set to a high frame rate to ensure smooth recording.
[0055] Step 3:
[0056] The device encodes the captured video data, optimizing the amount of data and converting it into a suitable format for transmission.
[0057] Step 4:
[0058] The device then transmits the encoded data to the server over a secure communication channel, where the communication is encrypted and privacy-protected.
[0059] Step 5:
[0060] The server decodes the received data, returning the encoded video data to a format that can be analyzed.
[0061] Step 6:
[0062] The server passes the decoded data to an image-processing AI model, which analyzes the user's lip and tongue movements and converts them into text suggestions.
[0063] Step 7:
[0064] The server receives text candidates generated by the image processing AI model, and uses the generative AI model to generate multiple utterance candidates based on the conversation context and surrounding content.
[0065] Step 8:
[0066] The server compares the generated utterance candidates and determines the most appropriate text, taking into account the context and past conversation history to select the most natural utterance.
[0067] Step 9:
[0068] The server then passes the determined text to a text-to-speech (TTS) model, which generates a voice in real time that closely resembles the user's voice quality.
[0069] Step 10:
[0070] The server encodes the generated audio data and sends it to the device, where it is converted into a format suitable for playback.
[0071] Step 11:
[0072] The device decodes the received audio data, which is then in a format that can be played back.
[0073] Step 12:
[0074] The terminal plays back the decoded voice data and outputs it as voice to the other party, who can then recognize the content of the user's silent utterance as voice.
[0075] In this way, the user's unvoiced speech can be transmitted to the other party in real time.
[0076] Example 1
[0077] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0078] Conventional voice communication systems require users to speak, resulting in issues of ambient noise and privacy. Furthermore, when silent speech is used, there is a lack of technology to accurately convert the content into text or speech, making smooth communication difficult. Therefore, there is a need for technology that can accurately convert silent speech into text and speech for communication.
[0079] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0080] In this invention, the server includes means for capturing a user's silent utterances at high resolution, means for encoding the captured data and transmitting it to the server, means for decoding the transmitted data, means for analyzing the user's lip and tongue movements using image processing on the decoded data, means for converting the analyzed data into text, means for generating a plurality of utterance candidates based on the converted text using a generative AI model, means for comparing the generated utterance candidates to determine the most appropriate text, means for converting the determined text into voice data using a voice generation AI model, and means for transmitting the voice data to the other party. This allows the user to accurately convert silent utterances into text and voice, enabling smooth communication.
[0081] "Silent speech" is a method in which a user speaks without making a sound, by moving their lips and tongue.
[0082] "High-resolution capture means" refers to high-precision cameras and sensors used to capture the movements of a user's lips and tongue in detail.
[0083] "Encoding" refers to the process of converting captured data into an optimized format, which preserves high quality information while minimizing data volume.
[0084] "Decoding" refers to the process of returning encoded data to its original form, making it possible to analyze the data.
[0085] "Image processing" refers to a series of calculations and algorithms used to extract and analyze useful information from captured video data, specifically analyzing the movements of the user's lips and tongue.
[0086] A "generative AI model" refers to an artificial intelligence model that generates text, speech, etc. based on given input data.
[0087] "Speech generation AI model" refers to an artificial intelligence model for generating natural-sounding speech based on text data.
[0088] The "means for converting into text" refers to a process for converting the analyzed motion data into a character string.
[0089] "Means for generating utterance candidates" refers to the process of using a generative AI model to create multiple utterance candidates based on input text data.
[0090] "Means for converting into audio data" refers to the process of generating audio based on text data and converting it into audio data format.
[0091] "Means for transmitting to the other party" refers to the communication means for delivering the final generated voice data to the other party via the terminal.
[0092] The present invention relates to a system that converts a user's silent speech into text in real time, and generates and transmits voice data. This system exchanges data between the user, the terminal, and a server to complete the overall process.
[0093] First, when a user makes a silent speech, the user uses a device equipped with a high-resolution camera to capture the speech. When the user puts on a speakerphone or headset and starts speaking silently, the camera captures this movement as high-resolution video. For example, imagine a situation where a user wants to have a confidential conversation with a business partner in a cafe. In this case, the user speaks only with the movement of their lips and tongue.
[0094] The device encodes the captured video data in real time and converts it into an optimized format (e.g., H.264). This encoding process uses a video processing library (e.g., FFmpeg). The encoded data is then sent to the server via a secure communication protocol (e.g., HTTPS).
[0095] The server decodes the received video data and uses image processing to analyze the user's lip and tongue movements. This analysis uses an image processing AI model (e.g., TensorFlow model). The analyzed movement data is converted into text data. For example, if a specific lip movement is determined to be a "P," the server adds a "P" to the text data.
[0096] Next, the server generates utterance candidates using a generative AI model (e.g., generative AI) based on the converted text data. The prompt text is entered as follows: "Generate a dialogue based on the content uttered silently by the user. The user is trying to say, 'How is the project progressing?'"
[0097] The generative AI model generates multiple utterance candidates and selects the most appropriate one. For example, it selects the most appropriate utterance based on the user's past conversation patterns and context. The selected text data is passed to a speech generation AI model (e.g., speech generation AI), which generates speech that closely resembles the user's voice quality.
[0098] The generated voice data is encoded (e.g., MP3 format) and sent to the device again via a secure communication protocol. The device decodes the received voice data and converts it into a playable format (e.g., WAV format). The decoded voice is played through the speaker and transmitted to the other party. For example, if a user utters a silent utterance such as "How is the project progressing?", the other party will feel as if the user is asking the question without making any sound.
[0099] According to the present invention, even if a user speaks silently, the content can be accurately converted into text and voice, enabling smooth communication.
[0100] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0101] Step 1:
[0102] The user initiates a silent utterance. The user puts on a speakerphone or headset and makes lip and tongue movements in front of a high-resolution camera. For example, the user might say, "How's the project going?" The input is lip and tongue movements, and the output is video data.
[0103] Step 2:
[0104] The device captures the user's lip and tongue movements as high-resolution video. The device's high-resolution camera captures the user's movements in real time and stores the captured video in a specific memory area. The input is the lip and tongue movements, and the output is high-resolution video data.
[0105] Step 3:
[0106] The device encodes the captured video data and sends it to the server via a secure communication protocol. The device uses a video processing library such as FFmpeg to encode the video data into H.264 format. The input is high-resolution video data, and the output is the encoded video data.
[0107] Step 4:
[0108] The server receives the encoded video data and decodes it. The server then converts the video data back to its original format using libraries such as FFmpeg, and passes it to the image processing AI model. The input is the encoded video data, and the output is the decoded video data.
[0109] Step 5:
[0110] The server analyzes the decoded video data and extracts the user's lip and tongue movements. The server uses an image processing AI model (e.g., TensorFlow model) to analyze the movements of each video frame. For example, a specific lip movement is determined to be "P." The input is the decoded video data, and the output is lip and tongue movement data.
[0111] Step 6:
[0112] The server converts lip and tongue movement data into text data. The server then maps the analyzed movement data into a string of characters and generates text data. For example, if it identifies the movement of the letter "P," it outputs the letter "P." The input is lip and tongue movement data, and the output is text data.
[0113] Step 7:
[0114] The server uses a generative AI model to generate utterance candidates based on text data. The prompt text is input as "Generate a dialogue based on the content uttered silently by the user. The user is trying to say, 'How is the project progressing?'" The generative AI model (e.g., generative AI) generates multiple utterance candidates based on the text data. The input is text data, and the output is a list of utterance candidates.
[0115] Step 8:
[0116] The server compares the generated utterance candidates and determines the optimal text. The server evaluates multiple candidates obtained from the generative AI model and selects the most appropriate utterance based on the user's past conversation patterns and context. The input is a list of utterance candidates, and the output is the optimal text.
[0117] Step 9:
[0118] The server uses a speech generation AI model to convert the optimal text into speech data. The server then passes the determined text to a speech generation AI model (e.g., speech generation AI) to generate speech that closely resembles the user's voice quality. The input is the optimal text, and the output is the generated speech data.
[0119] Step 10:
[0120] The server encodes the generated audio data and sends it to the device. The server encodes the generated audio data into MP3 format and sends it to the device via a secure communication protocol. The input is the generated audio data and the output is the encoded audio data.
[0121] Step 11:
[0122] The device decodes the audio data and plays it through the speaker. The device decodes the received audio data and converts it into a playable format (for example, WAV format). The decoded audio is played through the speaker and transmitted to the other party, making it sound as if the user is saying, "How's the project going?" The input is encoded audio data, and the output is the played audio.
[0123] (Application example 1)
[0124] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0125] In today's virtual stores, staff need to have strong communication skills to respond to customers smoothly and quickly. However, direct verbal communication can be difficult due to environmental noise and privacy issues. Furthermore, a method for silent communication using silent speech has yet to be established. This presents a challenge in providing effective customer service.
[0126] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0127] In this invention, the server includes means for capturing a user's silent utterances at high resolution, means for transmitting the captured data to the server, means for analyzing the transmitted data and converting it into text, means for generating a plurality of utterance candidates based on the converted text, means for comparing the generated utterance candidates and determining the most appropriate text, means for converting the determined text into voice data, means for transmitting the voice data to the other party, and means for the user to interact with customers silently in a virtual store to provide customer service support. This enables store staff to maintain a quiet environment while quickly and accurately interacting with customers through voiceless utterances.
[0128] "Silent speech" is speech that is made without producing any sound using lip or tongue movements.
[0129] "High-resolution capture means" refers to camera and sensor technology that accurately captures the subtle movements of the user's lips and tongue.
[0130] The "means for transmitting to the server" is a technology for transmitting the captured data to the server using a secure data communication protocol (e.g., gRPC or HTTP).
[0131] The "means of analyzing and converting into text" is a process of converting the received lip and tongue movements into text in natural language using image processing technology and natural language processing technology.
[0132] The "means for generating multiple candidate utterances" is a technology that uses a generative AI model to output multiple candidate utterances based on the context from text data.
[0133] The "means for determining the most appropriate text" is a process for selecting the most appropriate text for the context and situation from the multiple utterance candidates generated.
[0134] The "means of converting into voice data" refers to the process of generating voice that closely resembles the user's voice quality using voice synthesis technology based on the determined text data.
[0135] The "means for transmitting to the other party" is a communication technology for transferring the generated voice data to the other party in real time.
[0136] A "virtual store" is a store environment that uses virtual reality and augmented reality technology to offer products and services online.
[0137] The "means for users to provide customer service without vocal utterances to support customer service" is a technique in which staff members provide guidance and assistance to customers using vocal utterances.
[0138] The system embodying this invention converts a user's silent speech into text in real time, then converts it into voice data, and responds to customers in a virtual store. The system is composed of a user, a terminal, and a server.
[0139] The user puts on the smart glasses and begins silent speech. The smart glasses' camera captures the user's lip and tongue movements in high resolution. This data is encoded by the device in real time and transmitted to the server over a secure communication channel.
[0140] The server decodes the received data and uses an image processing AI model to analyze the user's lip and tongue movements. This analyzed data is converted into text data using a natural language processing library (e.g., spaCy or NLTK). A generative AI model is then used to generate multiple utterance candidates based on the context of the conversation and the surrounding content. These candidates are compared to determine the most suitable text. Next, a voice generation AI (Text-to-Speech, e.g., Google Text-to-Speech or Amazon Polly) uses the determined text to generate audio that closely resembles the user's voice quality. This audio data is encoded and sent back to the device via a secure communication channel, where it is played back to the customer.
[0141] As a specific example, consider the case where a customer asks about a new product in a virtual store. When the customer asks, "Do you have this product in other colors?", the staff user silently speaks into the smart glasses, "Do you have this product in other colors?" This speech is captured by a camera, and the encoded data is sent to the server. The server analyzes it and generates a similar text, "Do you have this product in other colors?" The generative AI model then determines that this text is appropriate for the context, and the speech generation AI generates a voice that is close to the staff member's voice quality. This voice data is sent to the device and played back to the customer in a natural way.
[0142] An example prompt for a generative AI model might look something like this:
[0143] "Convert the silent speech into text and generate voice data based on the following requirements: 'Speech content', Task: 'Guidance to customer', Voice quality: 'User's voice quality'"
[0144] This system allows store staff to provide quick and accurate customer service while maintaining a quiet and private environment.
[0145] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0146] Step 1:
[0147] The user puts on the smart glasses and begins silent speech, while the smart glasses' camera captures the user's lip and tongue movements in high resolution.
[0148] Input: User's lip and tongue movements
[0149] Output: High-resolution image data
[0150] How it works: The camera in the smart glasses captures the movements of the user's lips and tongue.
[0151] Step 2:
[0152] The device encodes the captured data in real time and transmits it to the server over a secure communication channel.
[0153] Input: High-resolution image data
[0154] Output: Encoded data
[0155] What it does: The data encoding module converts the image data into a compact format and sends it over the Internet.
[0156] Step 3:
[0157] The server decodes the received data and uses an image processing AI model to analyze the movements of the user's lips and tongue.
[0158] Input: Encoded data
[0159] Output: Analysis results (lip and tongue position information)
[0160] Specific operation: The data decoding module converts the encoded data back into the original image data, which is then analyzed by the image processing AI.
[0161] Step 4:
[0162] The server converts the analyzed data into text data using a natural language processing library.
[0163] Input: Analysis results (lip and tongue position information)
[0164] Output: Unedited text data
[0165] What it does: A natural language processing (NLP) library converts the parsed data into text.
[0166] Step 5:
[0167] The server uses a generative AI model to generate multiple utterance candidates based on the context of the conversation and the content before and after.
[0168] Input: Unedited text data
[0169] Output: Multiple utterance candidates
[0170] How it works: The generative AI model analyzes raw text and outputs multiple utterance candidates.
[0171] Step 6:
[0172] The server selects the most suitable text from the multiple generated utterance candidates.
[0173] Input: Multiple utterance candidates
[0174] Output: Optimal text
[0175] What it does: A contextual analysis algorithm selects the most appropriate text based on the context.
[0176] Step 7:
[0177] The server uses speech generation AI to convert the optimal text into a voice that closely resembles the user's voice quality.
[0178] Input: Optimal text
[0179] Output: Audio data
[0180] What it does: A text-to-speech engine converts text into speech.
[0181] Step 8:
[0182] The terminal receives the audio data from the server and plays it back to the customer.
[0183] Input: Audio data
[0184] Output: The audio to be played
[0185] Specific operation: The speaker of the smart glasses plays the audio data.
[0186] This process allows store staff to provide quick and accurate customer service while maintaining a quiet and private environment.
[0187] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0188] This invention relates to a system that converts a user's silent speech into text in real time and generates and transmits voice data. Furthermore, by combining it with an emotion engine, it is possible to recognize the emotion of the speech and reflect it in the text and voice. This system exchanges data between the user, terminal, and server to complete the overall processing.
[0189] Program processing explanation
[0190] When a user makes a silent speech, the system starts by capturing the movements of the user's lips and tongue. For example, consider a user who wants to express their opinion during a meeting. In this case, the user speaks only with the movements of their mouth.
[0191] User behavior
[0192] The user puts on a speakerphone or headset and begins to speak silently, and the camera responds, capturing high-resolution video of the user's lip and tongue movements.
[0193] Device behavior
[0194] The device captures the user's speech in real time and transmits the data to a server, which encodes the captured data into an optimized format and sends it to the server over a secure communication channel.
[0195] Server Operation
[0196] The server decodes the received data and uses an image processing AI model to analyze the user's lip and tongue movements. This analyzed data is converted into text. A generative AI model is then used to generate multiple utterance candidates based on the context of the conversation and the surrounding content. These candidates are then collated to determine the most suitable text.
[0197] Emotion Engine Operation
[0198] The server's emotion engine recognizes the user's emotions from silent speech and captured data. Based on this emotion information, the generated text and voice data are adjusted. For example, if the user is excited or depressed, the generated text and voice tone are adjusted to match that emotion.
[0199] Voice generation
[0200] The server then passes the determined text and emotional information to a text-to-speech (TTS) model, which then generates a voice that closely resembles the user's voice in real time. This voice data also reflects the recognized emotions.
[0201] Specific examples
[0202] For example, suppose a user makes a silent utterance such as, "How is the project progressing?". At this time, the user is slightly nervous. The movements of the user's lips and tongue are captured by the device and sent to the server. The server's image processing AI analyzes this movement and generates the text candidate, "How is the project progressing?". The generative AI model then generates utterance candidates that capture the context of the conversation and determines the most suitable text. The emotion engine recognizes the user's nervousness and adds a voice tone that reflects this. Based on this text and emotional information, the voice generation AI generates a voice that closely resembles the user's voice quality and sends it to the other party.
[0203] The person on the other end of the line can recognize that the user is silently asking, "How's the project going?" and sense the tension in their voice. This system allows users to maintain privacy and smooth communication that conveys emotion while maintaining the silence of the surrounding area.
[0204] The processing flow will be explained below.
[0205] Step 1:
[0206] The user initiates silent speech: the user begins speaking without making any sound using lip or tongue movements.
[0207] Step 2:
[0208] The device activates its built-in camera to capture high-resolution video of the user's lip and tongue movements, with the camera set to a high frame rate to ensure smooth recording of the mouth movements.
[0209] Step 3:
[0210] The device encodes the captured video data in real time, optimizing the data volume and converting it into a suitable format for transmission.
[0211] Step 4:
[0212] The device then transmits the encoded data to the server over a secure communication channel, where the communication is encrypted and privacy-protected.
[0213] Step 5:
[0214] The server decodes the received data, returning the encoded video data to a format that can be analyzed.
[0215] Step 6:
[0216] The server's image processing AI model analyzes the decoded video data and converts lip and tongue movements into text data.
[0217] Step 7:
[0218] The server's emotion engine recognizes emotions from video data and user speech, and this emotion information is used for subsequent processing.
[0219] Step 8:
[0220] The server receives text candidates generated by the image processing AI model, and uses the generative AI model to generate multiple utterance candidates based on the conversation context and surrounding content.
[0221] Step 9:
[0222] The server compares the generated utterance candidates and determines the most appropriate text, taking into account the context and past conversation history to select the most natural utterance.
[0223] Step 10:
[0224] The server adjusts the optimized text based on the emotional information from the emotion engine, adjusting the content and tone of the text according to the emotional information.
[0225] Step 11:
[0226] The server then passes the determined text and emotional information to a text-to-speech (TTS) model, which uses this information to generate a voice that closely resembles the user's voice in real time.
[0227] Step 12:
[0228] The server encodes the generated audio data and sends it to the device, where it is converted into a format suitable for playback.
[0229] Step 13:
[0230] The device decodes the received audio data, which is then in a format that can be played back.
[0231] Step 14:
[0232] The terminal plays back the decoded voice data and outputs it as voice to the other party, who can then recognize the content of the user's silent utterance as voice.
[0233] In this way, not only can the user's unvoiced utterances be conveyed to the other party in real time, but the user's emotions can also be reflected.
[0234] Example 2
[0235] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0236] Conventional communication systems make it difficult to ensure privacy through speech, making smooth communication difficult in public places or environments where silence is desired. Furthermore, when silent speech is used, there are no systems that can accurately convert speech from lip and tongue movements alone, and generate speech that reflects the user's emotions. Therefore, a new communication method that combines real-time conversion of silent speech into text and emotion recognition is needed.
[0237] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for capturing a user's silent utterance at high resolution, means for encoding the captured data, means for transmitting the encoded data to the server, means for decoding and analyzing the transmitted data, means for converting the analyzed data into text, means for generating a plurality of utterance candidates using a generative AI model, means for collating the generated utterance candidates to determine the optimal text, means for recognizing the user's emotion using an emotion engine, means for adjusting the text and voice data based on the recognized emotion information, means for converting the determined text into voice data, and means for transmitting the voice data to the other party. This enables smooth communication using silent utterances, converting them into text and voice in real time, and further reflecting the user's emotions in the voice.
[0238] "Capture" means recording the movements of the user's lips and tongue as high-resolution video.
[0239] "Encoding" refers to converting captured data into an optimized format.
[0240] "Sending to the server" means sending the encoded data to the server while taking security into consideration.
[0241] "Decoding" means restoring data received by the server to its original form.
[0242] "Analysis" refers to analyzing the movements of the user's lips and tongue based on the decoded data.
[0243] "Convert to text" means generating appropriate sentences from the analyzed data.
[0244] A "generative AI model" is an artificial intelligence that generates multiple utterance candidates based on textual data.
[0245] "Matching utterance candidates" refers to comparing the generated utterance candidates and selecting the most suitable text.
[0246] The "emotion engine" is a system that recognizes emotions from the user's silent utterances.
[0247] "Adjusting based on emotional information" means adjusting the tone and nuance of the generated text or audio data based on the emotions recognized by the emotion engine.
[0248] "Conversion to audio data" refers to converting the determined text into audio format.
[0249] "Sending voice data to the other party" means delivering voice data to the other party in real time.
[0250] This invention relates to a system that converts a user's silent speech into text in real time and generates and transmits voice data. Furthermore, by combining it with an emotion engine, it is possible to recognize the emotion of the speech and reflect it in the text and voice. This system exchanges data between the user, terminal, and server to complete the overall processing.
[0251] The user puts on a speakerphone or headset and begins to speak silently. The camera captures the user's lip and tongue movements as high-resolution video. For example, imagine a user silently says, "How's the project going?" during a meeting. The user's mouth movements are recorded in real time.
[0252] The device encodes the captured video data into an optimized format using video encoding software (e.g., FFmpeg), and then transmits the encoded data to the server over a secure communication channel (e.g., HTTPS or WebSocket).
[0253] The server first decodes the received data. Next, it uses an image processing AI model (e.g., OpenCV or TensorFlow) to analyze the user's lip and tongue movements. Text data is generated based on this analysis data. Next, it uses a generative AI model (e.g., GPT-3) to generate multiple utterance candidates based on the context of the conversation and the content before and after. These candidates are compared to determine the most appropriate text.
[0254] Next, the server's emotion engine (e.g., IBM Watson Emotion Analysis) recognizes the user's emotions from their silent speech and captured data. Based on this emotion information, the generated text and voice data are adjusted. For example, if the user is nervous, the generated text and voice tone are adjusted to match that emotion.
[0255] Finally, the server passes the determined text and emotional information to a speech generation AI (TTS) model (e.g., Google Text-to-Speech API), which generates a voice that closely resembles the user's voice in real time. This voice data is then transmitted to the other party via a secure communication channel.
[0256] For example, if a user silently says, "How's the project going?" and sounds a little nervous, the camera captures this movement and sends it from the device to the server. The server analyzes and converts the text, and a generative AI model generates utterance candidates that capture the context of the conversation and determines the most appropriate text. The emotion engine recognizes the user's nervousness and adds a voice tone that reflects this. Based on this text and emotional information, the voice generation AI generates a voice that closely resembles the user's voice quality and sends it to the other party.
[0257] An example of a prompt sentence might be:
[0258] "The user silently utters, 'How's the project going?' Please convert this into text. Also, the user seems a little nervous, so please generate a voice that reflects that emotion."
[0259] This allows users to maintain privacy and quiet surroundings while still being able to communicate smoothly and express their emotions.
[0260] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0261] Step 1:
[0262] The user puts on a speakerphone or headset and begins to speak silently. The camera captures the user's lip and tongue movements as high-resolution video. The input is the user's lip and tongue movements, and the output is the captured high-resolution video data. For example, if the user silently speaks, "How is the project going?", the camera records this movement in real time.
[0263] Step 2:
[0264] The device encodes the captured video data into an optimized format. The input is high-resolution video data, and the output is the encoded data. This process uses video encoding software (e.g., FFmpeg). Encoding reduces the data size and improves transmission efficiency.
[0265] Step 3:
[0266] The device sends the encoded data to the server via a secure communication channel (e.g., HTTPS or WebSocket). The input is the encoded data, and the output is the data sent to the server. The device encrypts the data when sending it to ensure its security. For example, the device sends the data using the SSL / TLS protocol.
[0267] Step 4:
[0268] The server decodes the received data. The input is the data sent from the device, and the output is the decoded data. The server then uses an image processing AI model (e.g., OpenCV or TensorFlow) to analyze the user's lip and tongue movements. This analysis generates the basic data for identifying what the user is saying.
[0269] Step 5:
[0270] The server converts the analyzed data into text. The input is data obtained by analyzing lip and tongue movements, and the output is text data. Next, a generative AI model (e.g., GPT-3) is used to generate multiple utterance candidates. The text is generated taking into account the context of the conversation and the content before and after.
[0271] Step 6:
[0272] The server compares the generated utterance candidates and determines the optimal text. The input is multiple utterance candidates, and the output is the optimal text. The generative AI model performs context analysis and selects the most appropriate candidate.
[0273] Step 7:
[0274] The server's emotion engine (e.g., IBM Watson Emotion Analysis) recognizes the user's emotions from the captured data of the user's silent speech. The input is real-time speech data and candidate utterance text, and the output is emotional information. Based on this emotional information, the generated text and voice data are adjusted.
[0275] Step 8:
[0276] The server passes the adjusted text and emotional information to a speech generation AI (TTS) model (e.g., Google Text-to-Speech API). The input is the adjusted text and emotional information, and the output is audio data. Based on this, the speech generation AI generates a voice that closely resembles the user's voice quality in real time.
[0277] Step 9:
[0278] The generated voice data is sent to the other party via a secure communication channel. The input is the voice data, and the output is the voice delivered to the other party. By listening to this voice, the other party can understand the silent message and its emotion that the user has made.
[0279] This allows users to communicate smoothly and convey their emotions without speaking.
[0280] (Application example 2)
[0281] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0282] When users communicate efficiently using silent speech in autonomous vehicles, they need a way to fully convey their emotions without being distracted by surrounding noise. However, existing technologies have had challenges, such as difficulty accurately grasping the user's intentions and emotions and generating unnatural speech. Furthermore, there has been a lack of means for quiet communication while driving or in noisy environments.
[0283] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0284] In this invention, the server includes means for capturing the user's silent utterances at high resolution, means for transmitting the captured data to the server, means for analyzing the transmitted data and converting it into text, means for generating a plurality of utterance candidates based on the converted text, means for comparing the generated utterance candidates to determine the most appropriate text, means for converting the determined text into voice data, means for transmitting the voice data to an in-vehicle audio device, and means for analyzing emotional information and reflecting it in the generated voice data. This enables the user to communicate effectively in an autonomous vehicle even without vocal utterances, and by generating natural voice that reflects emotions, smooth conversations can be held while maintaining quietness during the ride.
[0285] "User" means an end user of the System.
[0286] "Silent speech" is the act of communicating without making any sounds, but through the movement of the lips and tongue.
[0287] "High-resolution capture means" refers to equipment or methods for capturing the minute movements of a user's lips and tongue as high-resolution video data.
[0288] A "server" is a computer system for processing, storing, and analyzing data.
[0289] The "analysis means" refers to the algorithms or techniques used to analyze the received data and convert it into a particular format.
[0290] "Means for converting into text" refers to a technique or device that converts the analyzed data into text data in sentence format.
[0291] The "means for generating utterance candidates" refers to a technique or method for generating multiple candidate texts intended by the user based on the converted text data.
[0292] The "means for determining the optimal text" refers to an algorithm or process for selecting the optimal utterance from among the multiple generated candidate utterances.
[0293] The "means for converting into voice data" refers to a technique or method for generating voice data from the determined text.
[0294] "In-vehicle sound equipment" refers to speakers and sound systems installed inside autonomous vehicles.
[0295] "Means for analyzing emotional information" refers to technologies and algorithms that extract and analyze emotions from users' silent speech and captured data.
[0296] "Means for reflecting in voice data" refers to techniques or methods for reflecting analyzed emotional information in the tone and intonation of voice data.
[0297] The present invention is a system for enabling a user to communicate effectively using silent speech within an autonomous vehicle, and can be implemented using the following specific procedures and configurations.
[0298] System Configuration
[0299] User behavior
[0300] The user speaks silently. For example, if they want to give silent instructions while driving, they speak only with their mouths. The movements of the user's lips and tongue are captured by a high-resolution camera installed inside the car.
[0301] Device behavior
[0302] A system inside the autonomous vehicle captures the user's lip and tongue movements in real time and transmits the captured data to a server, where it is encoded in an optimized format and transmitted over a secure communication channel.
[0303] Server Operation
[0304] The server decodes the received data and uses an image processing AI model to analyze the user's lip and tongue movements. This analysis data is used to generate text data. A generative AI model is then used to generate multiple utterance candidates based on the conversation context and surrounding content. These candidates are then collated to determine the most suitable text.
[0305] Emotion Engine Operation
[0306] The server's emotion engine recognizes the user's emotions from silent speech and captured data. Based on this emotion information, the generated text and voice data are adjusted. For example, if the user is nervous or happy, the tone of the generated voice will be adjusted to match that emotion.
[0307] Audio Generation and Output
[0308] The server then passes the determined text and emotional information to a text-to-speech (TTS) model, which then generates a voice in real time that closely resembles the user's voice. This voice data is then output through the car's audio system.
[0309] Hardware and Software
[0310] This system uses the following hardware and software:
[0311] High-resolution camera: Used to capture the user's lip and tongue movements.
[0312] Server: Handles all of the analysis of received data, text generation, emotion recognition, and speech generation.
[0313] Image processing AI model: A model for analyzing the movements of the user's lips and tongue.
[0314] Generative AI model: A model for generating utterance candidates that capture the context of the conversation.
[0315] Emotion engine: A system for recognizing user emotions and reflecting them in voice data.
[0316] Speech generation AI (TTS) model: Generates speech based on generated text and emotional information.
[0317] Specific examples
[0318] For example, consider a case where a user silently utters, "Please tell me the distance to the next rest area." The user's lip and tongue movements are captured by a high-resolution camera inside the vehicle, and the data is sent to the server. The server's image processing AI model analyzes this movement and generates text data, "Please tell me the distance to the next rest area." Next, the generative AI model generates utterance candidates that capture the context, and the optimal text is determined. The emotion engine also recognizes the user's relaxed emotion and reflects that emotion in the generated voice data. Finally, the voice generation AI generates speech based on this text and emotional information, which is output through the vehicle's sound system.
[0319] Prompt Sentence Examples
[0320] "Please tell me how far it is to the next rest area."
[0321] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0322] Processing flow
[0323] Step 1:
[0324] The user makes silent speech. The user communicates without vocalization, using only lip and tongue movements. The input to this process is the user's lip and tongue movements, and the output is the captured video data.
[0325] Step 2:
[0326] The device uses a high-resolution camera to capture the user's lip and tongue movements. The input is the user's movements, and the output is high-resolution video data. The captured data is processed in real time.
[0327] Step 3:
[0328] The device sends the captured video data to the server. The input is the captured video data, and the output is the encoded data sent to the server over a secure communication channel.
[0329] Step 4:
[0330] The server decodes the received data and uses an image processing AI model to analyze the user's lip and tongue movements. The input is encoded video data, and the output is analyzed movement data. Text data is generated based on this analyzed data.
[0331] Step 5:
[0332] The server uses a generative AI model to generate multiple utterance candidates based on the converted text data. The input is text from the analyzed data, and the output is multiple generated utterance candidates. The utterance candidates are generated based on the context and surrounding content.
[0333] Step 6:
[0334] The server matches the generated candidate utterances and determines the best text. The input is multiple candidate utterances, and the output is the text that is judged to be the best. The matching process includes contextual analysis.
[0335] Step 7:
[0336] The emotion engine on the server recognizes the user's emotion from the user's silent utterances and captured data. The input is the analyzed data and text, and the output is the recognized emotion information.
[0337] Step 8:
[0338] The server passes the determined text and emotional information to a speech generation AI (TTS) model to generate speech data. The input is the optimal text and emotional information, and the output is the generated speech data. The speech data reflects the recognized emotion.
[0339] Step 9:
[0340] The server sends the generated voice data to the in-car audio system, which then outputs the voice from the car's speakers. The input is the generated voice data, and the output is the voice played in the car. This voice is natural and reflects the user's emotions.
[0341] Specific actions
[0342] As a specific example, when a user makes a silent utterance such as "Please tell me the distance to the next rest area," the process proceeds as follows:
[0343] 1. The user makes lip and tongue movements without speaking.
[0344] 2. The device captures the movement with a high-resolution camera.
[0345] 3. The device encodes the captured video data and sends it to the server.
[0346] 4. The server decodes and parses the received data.
[0347] 5. The server generates text data based on the analysis data.
[0348] 6. The generative AI model generates multiple utterance candidates.
[0349] 7. The server determines the best text.
[0350] 8. The emotion engine recognizes the user's emotions.
[0351] 9. Voice generation AI generates voice data based on optimal text and emotional information.
[0352] 10. The server sends the audio data to the car's audio device, which outputs it as audio from the car's speakers.
[0353] Prompt Sentence Examples
[0354] "Please tell me how far it is to the next rest area."
[0355] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0356] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0357] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0358] [Second embodiment]
[0359] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0360] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0361] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0362] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0363] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0364] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0365] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0366] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0367] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0368] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0369] In the smart glasses 214, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0370] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0371] The present invention relates to a system that converts a user's silent speech into text in real time, and generates and transmits voice data. This system exchanges data between the user, the terminal, and a server to complete the overall process.
[0372] Program processing explanation
[0373] When a user makes a silent speech, the system starts by capturing the movements of the user's lips and tongue. For example, consider a situation where a user wants to have a confidential conversation with a business partner at a cafe. In this case, the user speaks using only the movements of their mouth.
[0374] User behavior
[0375] The user puts on a speakerphone or headset and begins to speak silently, and the camera responds, capturing high-resolution video of the user's lip and tongue movements.
[0376] Device behavior
[0377] The device captures the user's speech in real time and transmits the data to a server, which encodes the captured data into an optimized format and sends it to the server over a secure communication channel.
[0378] Server Operation
[0379] The server decodes the received data and uses an image processing AI model to analyze the user's lip and tongue movements. This analyzed data is converted into text data. A generative AI model is then used to generate multiple utterance candidates based on the context of the conversation and the content before and after. These candidates are compared to determine the most suitable text. The determined text data is passed to a speech generation AI (TTS) model, which generates speech in real time that closely resembles the user's voice quality. The generated voice data is then encoded and sent to the device.
[0380] Specific examples
[0381] For example, suppose a user makes a silent utterance such as "How is the project progressing?" The user's lip and tongue movements are captured on the device and sent to the server. The server's image processing AI analyzes this movement and generates the text candidate "How is the project progressing?" The generative AI model then generates utterance candidates that capture the context of the conversation and determines the same text as the best candidate. Based on this text, the speech generation AI generates speech that closely resembles the user's voice quality and sends it to the other party.
[0382] The person on the other end of the line will feel as if the user is silently asking, "How's the project going?" This system allows users to maintain privacy and communicate smoothly while maintaining the surrounding silence.
[0383] The processing flow will be explained below.
[0384] Step 1:
[0385] The user initiates silent speech: The user silently initiates speech using lip and tongue movements.
[0386] Step 2:
[0387] The device activates its built-in camera to capture high-resolution video of the user's lip and tongue movements, with the camera set to a high frame rate to ensure smooth recording.
[0388] Step 3:
[0389] The device encodes the captured video data, optimizing the amount of data and converting it into a suitable format for transmission.
[0390] Step 4:
[0391] The device then transmits the encoded data to the server over a secure communication channel, where the communication is encrypted and privacy-protected.
[0392] Step 5:
[0393] The server decodes the received data, returning the encoded video data to a format that can be analyzed.
[0394] Step 6:
[0395] The server passes the decoded data to an image-processing AI model, which analyzes the user's lip and tongue movements and converts them into text suggestions.
[0396] Step 7:
[0397] The server receives text candidates generated by the image processing AI model, and uses the generative AI model to generate multiple utterance candidates based on the conversation context and surrounding content.
[0398] Step 8:
[0399] The server compares the generated utterance candidates and determines the most appropriate text, taking into account the context and past conversation history to select the most natural utterance.
[0400] Step 9:
[0401] The server then passes the determined text to a text-to-speech (TTS) model, which generates a voice in real time that closely resembles the user's voice quality.
[0402] Step 10:
[0403] The server encodes the generated audio data and sends it to the device, where it is converted into a format suitable for playback.
[0404] Step 11:
[0405] The device decodes the received audio data, which is then in a format that can be played back.
[0406] Step 12:
[0407] The terminal plays back the decoded voice data and outputs it as voice to the other party, who can then recognize the content of the user's silent utterance as voice.
[0408] In this way, the user's unvoiced speech can be transmitted to the other party in real time.
[0409] Example 1
[0410] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0411] Conventional voice communication systems require users to speak, resulting in issues of ambient noise and privacy. Furthermore, when silent speech is used, there is a lack of technology to accurately convert the content into text or speech, making smooth communication difficult. Therefore, there is a need for technology that can accurately convert silent speech into text and speech for communication.
[0412] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0413] In this invention, the server includes means for capturing a user's silent utterances at high resolution, means for encoding the captured data and transmitting it to the server, means for decoding the transmitted data, means for analyzing the user's lip and tongue movements using image processing on the decoded data, means for converting the analyzed data into text, means for generating a plurality of utterance candidates based on the converted text using a generative AI model, means for comparing the generated utterance candidates to determine the most appropriate text, means for converting the determined text into voice data using a voice generation AI model, and means for transmitting the voice data to the other party. This allows the user to accurately convert silent utterances into text and voice, enabling smooth communication.
[0414] "Silent speech" is a method in which a user speaks without making a sound, by moving their lips and tongue.
[0415] "High-resolution capture means" refers to high-precision cameras and sensors used to capture the movements of a user's lips and tongue in detail.
[0416] "Encoding" refers to the process of converting captured data into an optimized format, which preserves high quality information while minimizing data volume.
[0417] "Decoding" refers to the process of returning encoded data to its original form, making it possible to analyze the data.
[0418] "Image processing" refers to a series of calculations and algorithms used to extract and analyze useful information from captured video data, specifically analyzing the movements of the user's lips and tongue.
[0419] A "generative AI model" refers to an artificial intelligence model that generates text, speech, etc. based on given input data.
[0420] "Speech generation AI model" refers to an artificial intelligence model for generating natural-sounding speech based on text data.
[0421] The "means for converting into text" refers to a process for converting the analyzed motion data into a character string.
[0422] "Means for generating utterance candidates" refers to the process of using a generative AI model to create multiple utterance candidates based on input text data.
[0423] "Means for converting into audio data" refers to the process of generating audio based on text data and converting it into audio data format.
[0424] "Means for transmitting to the other party" refers to the communication means for delivering the final generated voice data to the other party via the terminal.
[0425] The present invention relates to a system that converts a user's silent speech into text in real time, and generates and transmits voice data. This system exchanges data between the user, the terminal, and a server to complete the overall process.
[0426] First, when a user makes a silent speech, the user uses a device equipped with a high-resolution camera to capture the speech. When the user puts on a speakerphone or headset and starts speaking silently, the camera captures this movement as high-resolution video. For example, imagine a situation where a user wants to have a confidential conversation with a business partner in a cafe. In this case, the user speaks only with the movement of their lips and tongue.
[0427] The device encodes the captured video data in real time and converts it into an optimized format (e.g., H.264). This encoding process uses a video processing library (e.g., FFmpeg). The encoded data is then sent to the server via a secure communication protocol (e.g., HTTPS).
[0428] The server decodes the received video data and uses image processing to analyze the user's lip and tongue movements. This analysis uses an image processing AI model (e.g., TensorFlow model). The analyzed movement data is converted into text data. For example, if a specific lip movement is determined to be a "P," the server adds a "P" to the text data.
[0429] Next, the server generates utterance candidates using a generative AI model (e.g., generative AI) based on the converted text data. The prompt text is entered as follows: "Generate a dialogue based on the content uttered silently by the user. The user is trying to say, 'How is the project progressing?'"
[0430] The generative AI model generates multiple utterance candidates and selects the most appropriate one. For example, it selects the most appropriate utterance based on the user's past conversation patterns and context. The selected text data is passed to a speech generation AI model (e.g., speech generation AI), which generates speech that closely resembles the user's voice quality.
[0431] The generated voice data is encoded (e.g., MP3 format) and sent to the device again via a secure communication protocol. The device decodes the received voice data and converts it into a playable format (e.g., WAV format). The decoded voice is played through the speaker and transmitted to the other party. For example, if a user utters a silent utterance such as "How is the project progressing?", the other party will feel as if the user is asking the question without making any sound.
[0432] According to the present invention, even if a user speaks silently, the content can be accurately converted into text and voice, enabling smooth communication.
[0433] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0434] Step 1:
[0435] The user initiates a silent utterance. The user puts on a speakerphone or headset and makes lip and tongue movements in front of a high-resolution camera. For example, the user might say, "How's the project going?" The input is lip and tongue movements, and the output is video data.
[0436] Step 2:
[0437] The device captures the user's lip and tongue movements as high-resolution video. The device's high-resolution camera captures the user's movements in real time and stores the captured video in a specific memory area. The input is the lip and tongue movements, and the output is high-resolution video data.
[0438] Step 3:
[0439] The device encodes the captured video data and sends it to the server via a secure communication protocol. The device uses a video processing library such as FFmpeg to encode the video data into H.264 format. The input is high-resolution video data, and the output is the encoded video data.
[0440] Step 4:
[0441] The server receives the encoded video data and decodes it. The server then converts the video data back to its original format using libraries such as FFmpeg, and passes it to the image processing AI model. The input is the encoded video data, and the output is the decoded video data.
[0442] Step 5:
[0443] The server analyzes the decoded video data and extracts the user's lip and tongue movements. The server uses an image processing AI model (e.g., TensorFlow model) to analyze the movements of each video frame. For example, a specific lip movement is determined to be "P." The input is the decoded video data, and the output is lip and tongue movement data.
[0444] Step 6:
[0445] The server converts lip and tongue movement data into text data. The server then maps the analyzed movement data into a string of characters and generates text data. For example, if it identifies the movement of the letter "P," it outputs the letter "P." The input is lip and tongue movement data, and the output is text data.
[0446] Step 7:
[0447] The server uses a generative AI model to generate utterance candidates based on text data. The prompt text is input as "Generate a dialogue based on the content uttered silently by the user. The user is trying to say, 'How is the project progressing?'" The generative AI model (e.g., generative AI) generates multiple utterance candidates based on the text data. The input is text data, and the output is a list of utterance candidates.
[0448] Step 8:
[0449] The server compares the generated utterance candidates and determines the optimal text. The server evaluates multiple candidates obtained from the generative AI model and selects the most appropriate utterance based on the user's past conversation patterns and context. The input is a list of utterance candidates, and the output is the optimal text.
[0450] Step 9:
[0451] The server uses a speech generation AI model to convert the optimal text into speech data. The server then passes the determined text to a speech generation AI model (e.g., speech generation AI) to generate speech that closely resembles the user's voice quality. The input is the optimal text, and the output is the generated speech data.
[0452] Step 10:
[0453] The server encodes the generated audio data and sends it to the device. The server encodes the generated audio data into MP3 format and sends it to the device via a secure communication protocol. The input is the generated audio data and the output is the encoded audio data.
[0454] Step 11:
[0455] The device decodes the audio data and plays it through the speaker. The device decodes the received audio data and converts it into a playable format (for example, WAV format). The decoded audio is played through the speaker and transmitted to the other party, making it sound as if the user is saying, "How's the project going?" The input is encoded audio data, and the output is the played audio.
[0456] (Application example 1)
[0457] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0458] In today's virtual stores, staff need to have strong communication skills to respond to customers smoothly and quickly. However, direct verbal communication can be difficult due to environmental noise and privacy issues. Furthermore, a method for silent communication using silent speech has yet to be established. This presents a challenge in providing effective customer service.
[0459] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0460] In this invention, the server includes means for capturing a user's silent utterances at high resolution, means for transmitting the captured data to the server, means for analyzing the transmitted data and converting it into text, means for generating a plurality of utterance candidates based on the converted text, means for comparing the generated utterance candidates and determining the most appropriate text, means for converting the determined text into voice data, means for transmitting the voice data to the other party, and means for the user to interact with customers silently in a virtual store to provide customer service support. This enables store staff to maintain a quiet environment while quickly and accurately interacting with customers through voiceless utterances.
[0461] "Silent speech" is speech that is made without producing any sound using lip or tongue movements.
[0462] "High-resolution capture means" refers to camera and sensor technology that accurately captures the subtle movements of the user's lips and tongue.
[0463] The "means for transmitting to the server" is a technology for transmitting the captured data to the server using a secure data communication protocol (e.g., gRPC or HTTP).
[0464] The "means of analyzing and converting into text" is a process of converting the received lip and tongue movements into text in natural language using image processing technology and natural language processing technology.
[0465] The "means for generating multiple candidate utterances" is a technology that uses a generative AI model to output multiple candidate utterances based on the context from text data.
[0466] The "means for determining the most appropriate text" is a process for selecting the most appropriate text for the context and situation from the multiple utterance candidates generated.
[0467] The "means of converting into voice data" refers to the process of generating voice that closely resembles the user's voice quality using voice synthesis technology based on the determined text data.
[0468] The "means for transmitting to the other party" is a communication technology for transferring the generated voice data to the other party in real time.
[0469] A "virtual store" is a store environment that uses virtual reality and augmented reality technology to offer products and services online.
[0470] The "means for users to provide customer service without vocal utterances to support customer service" is a technique in which staff members provide guidance and assistance to customers using vocal utterances.
[0471] The system embodying this invention converts a user's silent speech into text in real time, then converts it into voice data, and responds to customers in a virtual store. The system is composed of a user, a terminal, and a server.
[0472] The user puts on the smart glasses and begins silent speech. The smart glasses' camera captures the user's lip and tongue movements in high resolution. This data is encoded by the device in real time and transmitted to the server over a secure communication channel.
[0473] The server decodes the received data and uses an image processing AI model to analyze the user's lip and tongue movements. This analyzed data is converted into text data using a natural language processing library (e.g., spaCy or NLTK). A generative AI model is then used to generate multiple utterance candidates based on the context of the conversation and the surrounding content. These candidates are compared to determine the most suitable text. Next, a voice generation AI (Text-to-Speech, e.g., Google Text-to-Speech or Amazon Polly) uses the determined text to generate audio that closely resembles the user's voice quality. This audio data is encoded and sent back to the device via a secure communication channel, where it is played back to the customer.
[0474] As a specific example, consider the case where a customer asks about a new product in a virtual store. When the customer asks, "Do you have this product in other colors?", the staff user silently speaks into the smart glasses, "Do you have this product in other colors?" This speech is captured by a camera, and the encoded data is sent to the server. The server analyzes it and generates a similar text, "Do you have this product in other colors?" The generative AI model then determines that this text is appropriate for the context, and the speech generation AI generates a voice that is close to the staff member's voice quality. This voice data is sent to the device and played back to the customer in a natural way.
[0475] An example prompt for a generative AI model might look something like this:
[0476] "Convert the silent speech into text and generate voice data based on the following requirements: 'Speech content', Task: 'Guidance to customer', Voice quality: 'User's voice quality'"
[0477] This system allows store staff to provide quick and accurate customer service while maintaining a quiet and private environment.
[0478] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0479] Step 1:
[0480] The user puts on the smart glasses and begins silent speech, while the smart glasses' camera captures the user's lip and tongue movements in high resolution.
[0481] Input: User's lip and tongue movements
[0482] Output: High-resolution image data
[0483] How it works: The camera in the smart glasses captures the movements of the user's lips and tongue.
[0484] Step 2:
[0485] The device encodes the captured data in real time and transmits it to the server over a secure communication channel.
[0486] Input: High-resolution image data
[0487] Output: Encoded data
[0488] What it does: The data encoding module converts the image data into a compact format and sends it over the Internet.
[0489] Step 3:
[0490] The server decodes the received data and uses an image processing AI model to analyze the movements of the user's lips and tongue.
[0491] Input: Encoded data
[0492] Output: Analysis results (lip and tongue position information)
[0493] Specific operation: The data decoding module converts the encoded data back into the original image data, which is then analyzed by the image processing AI.
[0494] Step 4:
[0495] The server converts the analyzed data into text data using a natural language processing library.
[0496] Input: Analysis results (lip and tongue position information)
[0497] Output: Unedited text data
[0498] What it does: A natural language processing (NLP) library converts the parsed data into text.
[0499] Step 5:
[0500] The server uses a generative AI model to generate multiple utterance candidates based on the context of the conversation and the content before and after.
[0501] Input: Unedited text data
[0502] Output: Multiple utterance candidates
[0503] How it works: The generative AI model analyzes raw text and outputs multiple utterance candidates.
[0504] Step 6:
[0505] The server selects the most suitable text from the multiple generated utterance candidates.
[0506] Input: Multiple utterance candidates
[0507] Output: Optimal text
[0508] What it does: A contextual analysis algorithm selects the most appropriate text based on the context.
[0509] Step 7:
[0510] The server uses speech generation AI to convert the optimal text into a voice that closely resembles the user's voice quality.
[0511] Input: Optimal text
[0512] Output: Audio data
[0513] What it does: A text-to-speech engine converts text into speech.
[0514] Step 8:
[0515] The terminal receives the audio data from the server and plays it back to the customer.
[0516] Input: Audio data
[0517] Output: The audio to be played
[0518] Specific operation: The speaker of the smart glasses plays the audio data.
[0519] This process allows store staff to provide quick and accurate customer service while maintaining a quiet and private environment.
[0520] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0521] This invention relates to a system that converts a user's silent speech into text in real time and generates and transmits voice data. Furthermore, by combining it with an emotion engine, it is possible to recognize the emotion of the speech and reflect it in the text and voice. This system exchanges data between the user, terminal, and server to complete the overall processing.
[0522] Program processing explanation
[0523] When a user makes a silent speech, the system starts by capturing the movements of the user's lips and tongue. For example, consider a user who wants to express their opinion during a meeting. In this case, the user speaks only with the movements of their mouth.
[0524] User behavior
[0525] The user puts on a speakerphone or headset and begins to speak silently, and the camera responds, capturing high-resolution video of the user's lip and tongue movements.
[0526] Device behavior
[0527] The device captures the user's speech in real time and transmits the data to a server, which encodes the captured data into an optimized format and sends it to the server over a secure communication channel.
[0528] Server Operation
[0529] The server decodes the received data and uses an image processing AI model to analyze the user's lip and tongue movements. This analyzed data is converted into text. A generative AI model is then used to generate multiple utterance candidates based on the context of the conversation and the surrounding content. These candidates are then collated to determine the most suitable text.
[0530] Emotion Engine Operation
[0531] The server's emotion engine recognizes the user's emotions from silent speech and captured data. Based on this emotion information, the generated text and voice data are adjusted. For example, if the user is excited or depressed, the generated text and voice tone are adjusted to match that emotion.
[0532] Voice generation
[0533] The server then passes the determined text and emotional information to a text-to-speech (TTS) model, which then generates a voice that closely resembles the user's voice in real time. This voice data also reflects the recognized emotions.
[0534] Specific examples
[0535] For example, suppose a user makes a silent utterance such as, "How is the project progressing?". At this time, the user is slightly nervous. The movements of the user's lips and tongue are captured by the device and sent to the server. The server's image processing AI analyzes this movement and generates the text candidate, "How is the project progressing?". The generative AI model then generates utterance candidates that capture the context of the conversation and determines the most suitable text. The emotion engine recognizes the user's nervousness and adds a voice tone that reflects this. Based on this text and emotional information, the voice generation AI generates a voice that closely resembles the user's voice quality and sends it to the other party.
[0536] The person on the other end of the line can recognize that the user is silently asking, "How's the project going?" and sense the tension in their voice. This system allows users to maintain privacy and smooth communication that conveys emotion while maintaining the silence of the surrounding area.
[0537] The processing flow will be explained below.
[0538] Step 1:
[0539] The user initiates silent speech: the user begins speaking without making any sound using lip or tongue movements.
[0540] Step 2:
[0541] The device activates its built-in camera to capture high-resolution video of the user's lip and tongue movements, with the camera set to a high frame rate to ensure smooth recording of the mouth movements.
[0542] Step 3:
[0543] The device encodes the captured video data in real time, optimizing the data volume and converting it into a suitable format for transmission.
[0544] Step 4:
[0545] The device then transmits the encoded data to the server over a secure communication channel, where the communication is encrypted and privacy-protected.
[0546] Step 5:
[0547] The server decodes the received data, returning the encoded video data to a format that can be analyzed.
[0548] Step 6:
[0549] The server's image processing AI model analyzes the decoded video data and converts lip and tongue movements into text data.
[0550] Step 7:
[0551] The server's emotion engine recognizes emotions from video data and user speech, and this emotion information is used for subsequent processing.
[0552] Step 8:
[0553] The server receives text candidates generated by the image processing AI model, and uses the generative AI model to generate multiple utterance candidates based on the conversation context and surrounding content.
[0554] Step 9:
[0555] The server compares the generated utterance candidates and determines the most appropriate text, taking into account the context and past conversation history to select the most natural utterance.
[0556] Step 10:
[0557] The server adjusts the optimized text based on the emotional information from the emotion engine, adjusting the content and tone of the text according to the emotional information.
[0558] Step 11:
[0559] The server then passes the determined text and emotional information to a text-to-speech (TTS) model, which uses this information to generate a voice that closely resembles the user's voice in real time.
[0560] Step 12:
[0561] The server encodes the generated audio data and sends it to the device, where it is converted into a format suitable for playback.
[0562] Step 13:
[0563] The device decodes the received audio data, which is then in a format that can be played back.
[0564] Step 14:
[0565] The terminal plays back the decoded voice data and outputs it as voice to the other party, who can then recognize the content of the user's silent utterance as voice.
[0566] In this way, not only can the user's unvoiced utterances be conveyed to the other party in real time, but the user's emotions can also be reflected.
[0567] Example 2
[0568] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0569] Conventional communication systems make it difficult to ensure privacy through speech, making smooth communication difficult in public places or environments where silence is desired. Furthermore, when silent speech is used, there are no systems that can accurately convert speech from lip and tongue movements alone, and generate speech that reflects the user's emotions. Therefore, a new communication method that combines real-time conversion of silent speech into text and emotion recognition is needed.
[0570] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for capturing a user's silent utterance at high resolution, means for encoding the captured data, means for transmitting the encoded data to the server, means for decoding and analyzing the transmitted data, means for converting the analyzed data into text, means for generating a plurality of utterance candidates using a generative AI model, means for collating the generated utterance candidates to determine the optimal text, means for recognizing the user's emotion using an emotion engine, means for adjusting the text and voice data based on the recognized emotion information, means for converting the determined text into voice data, and means for transmitting the voice data to the other party. This enables smooth communication using silent utterances, converting them into text and voice in real time, and further reflecting the user's emotions in the voice.
[0571] "Capture" means recording the movements of the user's lips and tongue as high-resolution video.
[0572] "Encoding" refers to converting captured data into an optimized format.
[0573] "Sending to the server" means sending the encoded data to the server while taking security into consideration.
[0574] "Decoding" means restoring data received by the server to its original form.
[0575] "Analysis" refers to analyzing the movements of the user's lips and tongue based on the decoded data.
[0576] "Convert to text" means generating appropriate sentences from the analyzed data.
[0577] A "generative AI model" is an artificial intelligence that generates multiple utterance candidates based on textual data.
[0578] "Matching utterance candidates" refers to comparing the generated utterance candidates and selecting the most suitable text.
[0579] The "emotion engine" is a system that recognizes emotions from the user's silent utterances.
[0580] "Adjusting based on emotional information" means adjusting the tone and nuance of the generated text or audio data based on the emotions recognized by the emotion engine.
[0581] "Conversion to audio data" refers to converting the determined text into audio format.
[0582] "Sending voice data to the other party" means delivering voice data to the other party in real time.
[0583] This invention relates to a system that converts a user's silent speech into text in real time and generates and transmits voice data. Furthermore, by combining it with an emotion engine, it is possible to recognize the emotion of the speech and reflect it in the text and voice. This system exchanges data between the user, terminal, and server to complete the overall processing.
[0584] The user puts on a speakerphone or headset and begins to speak silently. The camera captures the user's lip and tongue movements as high-resolution video. For example, imagine a user silently says, "How's the project going?" during a meeting. The user's mouth movements are recorded in real time.
[0585] The device encodes the captured video data into an optimized format using video encoding software (e.g., FFmpeg), and then transmits the encoded data to the server over a secure communication channel (e.g., HTTPS or WebSocket).
[0586] The server first decodes the received data. Next, it uses an image processing AI model (e.g., OpenCV or TensorFlow) to analyze the user's lip and tongue movements. Text data is generated based on this analysis data. Next, it uses a generative AI model (e.g., GPT-3) to generate multiple utterance candidates based on the context of the conversation and the content before and after. These candidates are compared to determine the most appropriate text.
[0587] Next, the server's emotion engine (e.g., IBM Watson Emotion Analysis) recognizes the user's emotions from their silent speech and captured data. Based on this emotion information, the generated text and voice data are adjusted. For example, if the user is nervous, the generated text and voice tone are adjusted to match that emotion.
[0588] Finally, the server passes the determined text and emotional information to a speech generation AI (TTS) model (e.g., Google Text-to-Speech API), which generates a voice that closely resembles the user's voice in real time. This voice data is then transmitted to the other party via a secure communication channel.
[0589] For example, if a user silently says, "How's the project going?" and sounds a little nervous, the camera captures this movement and sends it from the device to the server. The server analyzes and converts the text, and a generative AI model generates utterance candidates that capture the context of the conversation and determines the most appropriate text. The emotion engine recognizes the user's nervousness and adds a voice tone that reflects this. Based on this text and emotional information, the voice generation AI generates a voice that closely resembles the user's voice quality and sends it to the other party.
[0590] An example of a prompt sentence might be:
[0591] "The user silently utters, 'How's the project going?' Please convert this into text. Also, the user seems a little nervous, so please generate a voice that reflects that emotion."
[0592] This allows users to maintain privacy and quiet surroundings while still being able to communicate smoothly and express their emotions.
[0593] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0594] Step 1:
[0595] The user puts on a speakerphone or headset and begins to speak silently. The camera captures the user's lip and tongue movements as high-resolution video. The input is the user's lip and tongue movements, and the output is the captured high-resolution video data. For example, if the user silently speaks, "How is the project going?", the camera records this movement in real time.
[0596] Step 2:
[0597] The device encodes the captured video data into an optimized format. The input is high-resolution video data, and the output is the encoded data. This process uses video encoding software (e.g., FFmpeg). Encoding reduces the data size and improves transmission efficiency.
[0598] Step 3:
[0599] The device sends the encoded data to the server via a secure communication channel (e.g., HTTPS or WebSocket). The input is the encoded data, and the output is the data sent to the server. The device encrypts the data when sending it to ensure its security. For example, the device sends the data using the SSL / TLS protocol.
[0600] Step 4:
[0601] The server decodes the received data. The input is the data sent from the device, and the output is the decoded data. The server then uses an image processing AI model (e.g., OpenCV or TensorFlow) to analyze the user's lip and tongue movements. This analysis generates the basic data for identifying what the user is saying.
[0602] Step 5:
[0603] The server converts the analyzed data into text. The input is data obtained by analyzing lip and tongue movements, and the output is text data. Next, a generative AI model (e.g., GPT-3) is used to generate multiple utterance candidates. The text is generated taking into account the context of the conversation and the content before and after.
[0604] Step 6:
[0605] The server compares the generated utterance candidates and determines the optimal text. The input is multiple utterance candidates, and the output is the optimal text. The generative AI model performs context analysis and selects the most appropriate candidate.
[0606] Step 7:
[0607] The server's emotion engine (e.g., IBM Watson Emotion Analysis) recognizes the user's emotions from the captured data of the user's silent speech. The input is real-time speech data and candidate utterance text, and the output is emotional information. Based on this emotional information, the generated text and voice data are adjusted.
[0608] Step 8:
[0609] The server passes the adjusted text and emotional information to a speech generation AI (TTS) model (e.g., Google Text-to-Speech API). The input is the adjusted text and emotional information, and the output is audio data. Based on this, the speech generation AI generates a voice that closely resembles the user's voice quality in real time.
[0610] Step 9:
[0611] The generated voice data is sent to the other party via a secure communication channel. The input is the voice data, and the output is the voice delivered to the other party. By listening to this voice, the other party can understand the silent message and its emotion that the user has made.
[0612] This allows users to communicate smoothly and convey their emotions without speaking.
[0613] (Application example 2)
[0614] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0615] When users communicate efficiently using silent speech in autonomous vehicles, they need a way to fully convey their emotions without being distracted by surrounding noise. However, existing technologies have had challenges, such as difficulty accurately grasping the user's intentions and emotions and generating unnatural speech. Furthermore, there has been a lack of means for quiet communication while driving or in noisy environments.
[0616] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0617] In this invention, the server includes means for capturing the user's silent utterances at high resolution, means for transmitting the captured data to the server, means for analyzing the transmitted data and converting it into text, means for generating a plurality of utterance candidates based on the converted text, means for comparing the generated utterance candidates to determine the most appropriate text, means for converting the determined text into voice data, means for transmitting the voice data to an in-vehicle audio device, and means for analyzing emotional information and reflecting it in the generated voice data. This enables the user to communicate effectively in an autonomous vehicle even without vocal utterances, and by generating natural voice that reflects emotions, smooth conversations can be held while maintaining quietness during the ride.
[0618] "User" means an end user of the System.
[0619] "Silent speech" is the act of communicating without making any sounds, but through the movement of the lips and tongue.
[0620] "High-resolution capture means" refers to equipment or methods for capturing the minute movements of a user's lips and tongue as high-resolution video data.
[0621] A "server" is a computer system for processing, storing, and analyzing data.
[0622] The "analysis means" refers to the algorithms or techniques used to analyze the received data and convert it into a particular format.
[0623] "Means for converting into text" refers to a technique or device that converts the analyzed data into text data in sentence format.
[0624] The "means for generating utterance candidates" refers to a technique or method for generating multiple candidate texts intended by the user based on the converted text data.
[0625] The "means for determining the optimal text" refers to an algorithm or process for selecting the optimal utterance from among the multiple generated candidate utterances.
[0626] The "means for converting into voice data" refers to a technique or method for generating voice data from the determined text.
[0627] "In-vehicle sound equipment" refers to speakers and sound systems installed inside autonomous vehicles.
[0628] "Means for analyzing emotional information" refers to technologies and algorithms that extract and analyze emotions from users' silent speech and captured data.
[0629] "Means for reflecting in voice data" refers to techniques or methods for reflecting analyzed emotional information in the tone and intonation of voice data.
[0630] The present invention is a system for enabling a user to communicate effectively using silent speech within an autonomous vehicle, and can be implemented using the following specific procedures and configurations.
[0631] System Configuration
[0632] User behavior
[0633] The user speaks silently. For example, if they want to give silent instructions while driving, they speak only with their mouths. The movements of the user's lips and tongue are captured by a high-resolution camera installed inside the car.
[0634] Device behavior
[0635] A system inside the autonomous vehicle captures the user's lip and tongue movements in real time and transmits the captured data to a server, where it is encoded in an optimized format and transmitted over a secure communication channel.
[0636] Server Operation
[0637] The server decodes the received data and uses an image processing AI model to analyze the user's lip and tongue movements. This analysis data is used to generate text data. A generative AI model is then used to generate multiple utterance candidates based on the conversation context and surrounding content. These candidates are then collated to determine the most suitable text.
[0638] Emotion Engine Operation
[0639] The server's emotion engine recognizes the user's emotions from silent speech and captured data. Based on this emotion information, the generated text and voice data are adjusted. For example, if the user is nervous or happy, the tone of the generated voice will be adjusted to match that emotion.
[0640] Audio Generation and Output
[0641] The server then passes the determined text and emotional information to a text-to-speech (TTS) model, which then generates a voice in real time that closely resembles the user's voice. This voice data is then output through the car's audio system.
[0642] Hardware and Software
[0643] This system uses the following hardware and software:
[0644] High-resolution camera: Used to capture the user's lip and tongue movements.
[0645] Server: Handles all of the analysis of received data, text generation, emotion recognition, and speech generation.
[0646] Image processing AI model: A model for analyzing the movements of the user's lips and tongue.
[0647] Generative AI model: A model for generating utterance candidates that capture the context of the conversation.
[0648] Emotion engine: A system for recognizing user emotions and reflecting them in voice data.
[0649] Speech generation AI (TTS) model: Generates speech based on generated text and emotional information.
[0650] Specific examples
[0651] For example, consider a case where a user silently utters, "Please tell me the distance to the next rest area." The user's lip and tongue movements are captured by a high-resolution camera inside the vehicle, and the data is sent to the server. The server's image processing AI model analyzes this movement and generates text data, "Please tell me the distance to the next rest area." Next, the generative AI model generates utterance candidates that capture the context, and the optimal text is determined. The emotion engine also recognizes the user's relaxed emotion and reflects that emotion in the generated voice data. Finally, the voice generation AI generates speech based on this text and emotional information, which is output through the vehicle's sound system.
[0652] Prompt Sentence Examples
[0653] "Please tell me how far it is to the next rest area."
[0654] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0655] Processing flow
[0656] Step 1:
[0657] The user makes silent speech. The user communicates without vocalization, using only lip and tongue movements. The input to this process is the user's lip and tongue movements, and the output is the captured video data.
[0658] Step 2:
[0659] The device uses a high-resolution camera to capture the user's lip and tongue movements. The input is the user's movements, and the output is high-resolution video data. The captured data is processed in real time.
[0660] Step 3:
[0661] The device sends the captured video data to the server. The input is the captured video data, and the output is the encoded data sent to the server over a secure communication channel.
[0662] Step 4:
[0663] The server decodes the received data and uses an image processing AI model to analyze the user's lip and tongue movements. The input is encoded video data, and the output is analyzed movement data. Text data is generated based on this analyzed data.
[0664] Step 5:
[0665] The server uses a generative AI model to generate multiple utterance candidates based on the converted text data. The input is text from the analyzed data, and the output is multiple generated utterance candidates. The utterance candidates are generated based on the context and surrounding content.
[0666] Step 6:
[0667] The server matches the generated candidate utterances and determines the best text. The input is multiple candidate utterances, and the output is the text that is judged to be the best. The matching process includes contextual analysis.
[0668] Step 7:
[0669] The emotion engine on the server recognizes the user's emotion from the user's silent utterances and captured data. The input is the analyzed data and text, and the output is the recognized emotion information.
[0670] Step 8:
[0671] The server passes the determined text and emotional information to a speech generation AI (TTS) model to generate speech data. The input is the optimal text and emotional information, and the output is the generated speech data. The speech data reflects the recognized emotion.
[0672] Step 9:
[0673] The server sends the generated voice data to the in-car audio system, which then outputs the voice from the car's speakers. The input is the generated voice data, and the output is the voice played in the car. This voice is natural and reflects the user's emotions.
[0674] Specific actions
[0675] As a specific example, when a user makes a silent utterance such as "Please tell me the distance to the next rest area," the process proceeds as follows:
[0676] 1. The user makes lip and tongue movements without speaking.
[0677] 2. The device captures the movement with a high-resolution camera.
[0678] 3. The device encodes the captured video data and sends it to the server.
[0679] 4. The server decodes and parses the received data.
[0680] 5. The server generates text data based on the analysis data.
[0681] 6. The generative AI model generates multiple utterance candidates.
[0682] 7. The server determines the best text.
[0683] 8. The emotion engine recognizes the user's emotions.
[0684] 9. Voice generation AI generates voice data based on optimal text and emotional information.
[0685] 10. The server sends the audio data to the car's audio device, which outputs it as audio from the car's speakers.
[0686] Prompt Sentence Examples
[0687] "Please tell me how far it is to the next rest area."
[0688] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0689] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0690] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0691] [Third embodiment]
[0692] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0693] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0694] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0695] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0696] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0697] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0698] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0699] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0700] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0701] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0702] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0703] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0704] The present invention relates to a system that converts a user's silent speech into text in real time, and generates and transmits voice data. This system exchanges data between the user, the terminal, and a server to complete the overall process.
[0705] Program processing explanation
[0706] When a user makes a silent speech, the system starts by capturing the movements of the user's lips and tongue. For example, consider a situation where a user wants to have a confidential conversation with a business partner at a cafe. In this case, the user speaks using only the movements of their mouth.
[0707] User behavior
[0708] The user puts on a speakerphone or headset and begins to speak silently, and the camera responds, capturing high-resolution video of the user's lip and tongue movements.
[0709] Device behavior
[0710] The device captures the user's speech in real time and transmits the data to a server, which encodes the captured data into an optimized format and sends it to the server over a secure communication channel.
[0711] Server Operation
[0712] The server decodes the received data and uses an image processing AI model to analyze the user's lip and tongue movements. This analyzed data is converted into text data. A generative AI model is then used to generate multiple utterance candidates based on the context of the conversation and the content before and after. These candidates are compared to determine the most suitable text. The determined text data is passed to a speech generation AI (TTS) model, which generates speech in real time that closely resembles the user's voice quality. The generated voice data is then encoded and sent to the device.
[0713] Specific examples
[0714] For example, suppose a user makes a silent utterance such as "How is the project progressing?" The user's lip and tongue movements are captured on the device and sent to the server. The server's image processing AI analyzes this movement and generates the text candidate "How is the project progressing?" The generative AI model then generates utterance candidates that capture the context of the conversation and determines the same text as the best candidate. Based on this text, the speech generation AI generates speech that closely resembles the user's voice quality and sends it to the other party.
[0715] The person on the other end of the line will feel as if the user is silently asking, "How's the project going?" This system allows users to maintain privacy and communicate smoothly while maintaining the surrounding silence.
[0716] The processing flow will be explained below.
[0717] Step 1:
[0718] The user initiates silent speech: The user silently initiates speech using lip and tongue movements.
[0719] Step 2:
[0720] The device activates its built-in camera to capture high-resolution video of the user's lip and tongue movements, with the camera set to a high frame rate to ensure smooth recording.
[0721] Step 3:
[0722] The device encodes the captured video data, optimizing the amount of data and converting it into a suitable format for transmission.
[0723] Step 4:
[0724] The device then transmits the encoded data to the server over a secure communication channel, where the communication is encrypted and privacy-protected.
[0725] Step 5:
[0726] The server decodes the received data, returning the encoded video data to a format that can be analyzed.
[0727] Step 6:
[0728] The server passes the decoded data to an image-processing AI model, which analyzes the user's lip and tongue movements and converts them into text suggestions.
[0729] Step 7:
[0730] The server receives text candidates generated by the image processing AI model, and uses the generative AI model to generate multiple utterance candidates based on the conversation context and surrounding content.
[0731] Step 8:
[0732] The server compares the generated utterance candidates and determines the most appropriate text, taking into account the context and past conversation history to select the most natural utterance.
[0733] Step 9:
[0734] The server then passes the determined text to a text-to-speech (TTS) model, which generates a voice in real time that closely resembles the user's voice quality.
[0735] Step 10:
[0736] The server encodes the generated audio data and sends it to the device, where it is converted into a format suitable for playback.
[0737] Step 11:
[0738] The device decodes the received audio data, which is then in a format that can be played back.
[0739] Step 12:
[0740] The terminal plays back the decoded voice data and outputs it as voice to the other party, who can then recognize the content of the user's silent utterance as voice.
[0741] In this way, the user's unvoiced speech can be transmitted to the other party in real time.
[0742] Example 1
[0743] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0744] Conventional voice communication systems require users to speak, resulting in issues of ambient noise and privacy. Furthermore, when silent speech is used, there is a lack of technology to accurately convert the content into text or speech, making smooth communication difficult. Therefore, there is a need for technology that can accurately convert silent speech into text and speech for communication.
[0745] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0746] In this invention, the server includes means for capturing a user's silent utterances at high resolution, means for encoding the captured data and transmitting it to the server, means for decoding the transmitted data, means for analyzing the user's lip and tongue movements using image processing on the decoded data, means for converting the analyzed data into text, means for generating a plurality of utterance candidates based on the converted text using a generative AI model, means for comparing the generated utterance candidates to determine the most appropriate text, means for converting the determined text into voice data using a voice generation AI model, and means for transmitting the voice data to the other party. This allows the user to accurately convert silent utterances into text and voice, enabling smooth communication.
[0747] "Silent speech" is a method in which a user speaks without making a sound, by moving their lips and tongue.
[0748] "High-resolution capture means" refers to high-precision cameras and sensors used to capture the movements of a user's lips and tongue in detail.
[0749] "Encoding" refers to the process of converting captured data into an optimized format, which preserves high quality information while minimizing data volume.
[0750] "Decoding" refers to the process of returning encoded data to its original form, making it possible to analyze the data.
[0751] "Image processing" refers to a series of calculations and algorithms used to extract and analyze useful information from captured video data, specifically analyzing the movements of the user's lips and tongue.
[0752] A "generative AI model" refers to an artificial intelligence model that generates text, speech, etc. based on given input data.
[0753] "Speech generation AI model" refers to an artificial intelligence model for generating natural-sounding speech based on text data.
[0754] The "means for converting into text" refers to a process for converting the analyzed motion data into a character string.
[0755] "Means for generating utterance candidates" refers to the process of using a generative AI model to create multiple utterance candidates based on input text data.
[0756] "Means for converting into audio data" refers to the process of generating audio based on text data and converting it into audio data format.
[0757] "Means for transmitting to the other party" refers to the communication means for delivering the final generated voice data to the other party via the terminal.
[0758] The present invention relates to a system that converts a user's silent speech into text in real time, and generates and transmits voice data. This system exchanges data between the user, the terminal, and a server to complete the overall process.
[0759] First, when a user makes a silent speech, the user uses a device equipped with a high-resolution camera to capture the speech. When the user puts on a speakerphone or headset and starts speaking silently, the camera captures this movement as high-resolution video. For example, imagine a situation where a user wants to have a confidential conversation with a business partner in a cafe. In this case, the user speaks only with the movement of their lips and tongue.
[0760] The device encodes the captured video data in real time and converts it into an optimized format (e.g., H.264). This encoding process uses a video processing library (e.g., FFmpeg). The encoded data is then sent to the server via a secure communication protocol (e.g., HTTPS).
[0761] The server decodes the received video data and uses image processing to analyze the user's lip and tongue movements. This analysis uses an image processing AI model (e.g., TensorFlow model). The analyzed movement data is converted into text data. For example, if a specific lip movement is determined to be a "P," the server adds a "P" to the text data.
[0762] Next, the server generates utterance candidates using a generative AI model (e.g., generative AI) based on the converted text data. The prompt text is entered as follows: "Generate a dialogue based on the content uttered silently by the user. The user is trying to say, 'How is the project progressing?'"
[0763] The generative AI model generates multiple utterance candidates and selects the most appropriate one. For example, it selects the most appropriate utterance based on the user's past conversation patterns and context. The selected text data is passed to a speech generation AI model (e.g., speech generation AI), which generates speech that closely resembles the user's voice quality.
[0764] The generated voice data is encoded (e.g., MP3 format) and sent to the device again via a secure communication protocol. The device decodes the received voice data and converts it into a playable format (e.g., WAV format). The decoded voice is played through the speaker and transmitted to the other party. For example, if a user utters a silent utterance such as "How is the project progressing?", the other party will feel as if the user is asking the question without making any sound.
[0765] According to the present invention, even if a user speaks silently, the content can be accurately converted into text and voice, enabling smooth communication.
[0766] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0767] Step 1:
[0768] The user initiates a silent utterance. The user puts on a speakerphone or headset and makes lip and tongue movements in front of a high-resolution camera. For example, the user might say, "How's the project going?" The input is lip and tongue movements, and the output is video data.
[0769] Step 2:
[0770] The device captures the user's lip and tongue movements as high-resolution video. The device's high-resolution camera captures the user's movements in real time and stores the captured video in a specific memory area. The input is the lip and tongue movements, and the output is high-resolution video data.
[0771] Step 3:
[0772] The device encodes the captured video data and sends it to the server via a secure communication protocol. The device uses a video processing library such as FFmpeg to encode the video data into H.264 format. The input is high-resolution video data, and the output is the encoded video data.
[0773] Step 4:
[0774] The server receives the encoded video data and decodes it. The server then converts the video data back to its original format using libraries such as FFmpeg, and passes it to the image processing AI model. The input is the encoded video data, and the output is the decoded video data.
[0775] Step 5:
[0776] The server analyzes the decoded video data and extracts the user's lip and tongue movements. The server uses an image processing AI model (e.g., TensorFlow model) to analyze the movements of each video frame. For example, a specific lip movement is determined to be "P." The input is the decoded video data, and the output is lip and tongue movement data.
[0777] Step 6:
[0778] The server converts lip and tongue movement data into text data. The server then maps the analyzed movement data into a string of characters and generates text data. For example, if it identifies the movement of the letter "P," it outputs the letter "P." The input is lip and tongue movement data, and the output is text data.
[0779] Step 7:
[0780] The server uses a generative AI model to generate utterance candidates based on text data. The prompt text is input as "Generate a dialogue based on the content uttered silently by the user. The user is trying to say, 'How is the project progressing?'" The generative AI model (e.g., generative AI) generates multiple utterance candidates based on the text data. The input is text data, and the output is a list of utterance candidates.
[0781] Step 8:
[0782] The server compares the generated utterance candidates and determines the optimal text. The server evaluates multiple candidates obtained from the generative AI model and selects the most appropriate utterance based on the user's past conversation patterns and context. The input is a list of utterance candidates, and the output is the optimal text.
[0783] Step 9:
[0784] The server uses a speech generation AI model to convert the optimal text into speech data. The server then passes the determined text to a speech generation AI model (e.g., speech generation AI) to generate speech that closely resembles the user's voice quality. The input is the optimal text, and the output is the generated speech data.
[0785] Step 10:
[0786] The server encodes the generated audio data and sends it to the device. The server encodes the generated audio data into MP3 format and sends it to the device via a secure communication protocol. The input is the generated audio data and the output is the encoded audio data.
[0787] Step 11:
[0788] The device decodes the audio data and plays it through the speaker. The device decodes the received audio data and converts it into a playable format (for example, WAV format). The decoded audio is played through the speaker and transmitted to the other party, making it sound as if the user is saying, "How's the project going?" The input is encoded audio data, and the output is the played audio.
[0789] (Application example 1)
[0790] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0791] In today's virtual stores, staff need to have strong communication skills to respond to customers smoothly and quickly. However, direct verbal communication can be difficult due to environmental noise and privacy issues. Furthermore, a method for silent communication using silent speech has yet to be established. This presents a challenge in providing effective customer service.
[0792] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0793] In this invention, the server includes means for capturing a user's silent utterances at high resolution, means for transmitting the captured data to the server, means for analyzing the transmitted data and converting it into text, means for generating a plurality of utterance candidates based on the converted text, means for comparing the generated utterance candidates and determining the most appropriate text, means for converting the determined text into voice data, means for transmitting the voice data to the other party, and means for the user to interact with customers silently in a virtual store to provide customer service support. This enables store staff to maintain a quiet environment while quickly and accurately interacting with customers through voiceless utterances.
[0794] "Silent speech" is speech that is made without producing any sound using lip or tongue movements.
[0795] "High-resolution capture means" refers to camera and sensor technology that accurately captures the subtle movements of the user's lips and tongue.
[0796] The "means for transmitting to the server" is a technology for transmitting the captured data to the server using a secure data communication protocol (e.g., gRPC or HTTP).
[0797] The "means of analyzing and converting into text" is a process of converting the received lip and tongue movements into text in natural language using image processing technology and natural language processing technology.
[0798] The "means for generating multiple candidate utterances" is a technology that uses a generative AI model to output multiple candidate utterances based on the context from text data.
[0799] The "means for determining the most appropriate text" is a process for selecting the most appropriate text for the context and situation from the multiple utterance candidates generated.
[0800] The "means of converting into voice data" refers to the process of generating voice that closely resembles the user's voice quality using voice synthesis technology based on the determined text data.
[0801] The "means for transmitting to the other party" is a communication technology for transferring the generated voice data to the other party in real time.
[0802] A "virtual store" is a store environment that uses virtual reality and augmented reality technology to offer products and services online.
[0803] The "means for users to provide customer service without vocal utterances to support customer service" is a technique in which staff members provide guidance and assistance to customers using vocal utterances.
[0804] The system embodying this invention converts a user's silent speech into text in real time, then converts it into voice data, and responds to customers in a virtual store. The system is composed of a user, a terminal, and a server.
[0805] The user puts on the smart glasses and begins silent speech. The smart glasses' camera captures the user's lip and tongue movements in high resolution. This data is encoded by the device in real time and transmitted to the server over a secure communication channel.
[0806] The server decodes the received data and uses an image processing AI model to analyze the user's lip and tongue movements. This analyzed data is converted into text data using a natural language processing library (e.g., spaCy or NLTK). A generative AI model is then used to generate multiple utterance candidates based on the context of the conversation and the surrounding content. These candidates are compared to determine the most suitable text. Next, a voice generation AI (Text-to-Speech, e.g., Google Text-to-Speech or Amazon Polly) uses the determined text to generate audio that closely resembles the user's voice quality. This audio data is encoded and sent back to the device via a secure communication channel, where it is played back to the customer.
[0807] As a specific example, consider the case where a customer asks about a new product in a virtual store. When the customer asks, "Do you have this product in other colors?", the staff user silently speaks into the smart glasses, "Do you have this product in other colors?" This speech is captured by a camera, and the encoded data is sent to the server. The server analyzes it and generates a similar text, "Do you have this product in other colors?" The generative AI model then determines that this text is appropriate for the context, and the speech generation AI generates a voice that is close to the staff member's voice quality. This voice data is sent to the device and played back to the customer in a natural way.
[0808] An example prompt for a generative AI model might look something like this:
[0809] "Convert the silent speech into text and generate voice data based on the following requirements: 'Speech content', Task: 'Guidance to customer', Voice quality: 'User's voice quality'"
[0810] This system allows store staff to provide quick and accurate customer service while maintaining a quiet and private environment.
[0811] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0812] Step 1:
[0813] The user puts on the smart glasses and begins silent speech, while the smart glasses' camera captures the user's lip and tongue movements in high resolution.
[0814] Input: User's lip and tongue movements
[0815] Output: High-resolution image data
[0816] How it works: The camera in the smart glasses captures the movements of the user's lips and tongue.
[0817] Step 2:
[0818] The device encodes the captured data in real time and transmits it to the server over a secure communication channel.
[0819] Input: High-resolution image data
[0820] Output: Encoded data
[0821] What it does: The data encoding module converts the image data into a compact format and sends it over the Internet.
[0822] Step 3:
[0823] The server decodes the received data and uses an image processing AI model to analyze the movements of the user's lips and tongue.
[0824] Input: Encoded data
[0825] Output: Analysis results (lip and tongue position information)
[0826] Specific operation: The data decoding module converts the encoded data back into the original image data, which is then analyzed by the image processing AI.
[0827] Step 4:
[0828] The server converts the analyzed data into text data using a natural language processing library.
[0829] Input: Analysis results (lip and tongue position information)
[0830] Output: Unedited text data
[0831] What it does: A natural language processing (NLP) library converts the parsed data into text.
[0832] Step 5:
[0833] The server uses a generative AI model to generate multiple utterance candidates based on the context of the conversation and the content before and after.
[0834] Input: Unedited text data
[0835] Output: Multiple utterance candidates
[0836] How it works: The generative AI model analyzes raw text and outputs multiple utterance candidates.
[0837] Step 6:
[0838] The server selects the most suitable text from the multiple generated utterance candidates.
[0839] Input: Multiple utterance candidates
[0840] Output: Optimal text
[0841] What it does: A contextual analysis algorithm selects the most appropriate text based on the context.
[0842] Step 7:
[0843] The server uses speech generation AI to convert the optimal text into a voice that closely resembles the user's voice quality.
[0844] Input: Optimal text
[0845] Output: Audio data
[0846] What it does: A text-to-speech engine converts text into speech.
[0847] Step 8:
[0848] The terminal receives the audio data from the server and plays it back to the customer.
[0849] Input: Audio data
[0850] Output: The audio to be played
[0851] Specific operation: The speaker of the smart glasses plays the audio data.
[0852] This process allows store staff to provide quick and accurate customer service while maintaining a quiet and private environment.
[0853] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0854] This invention relates to a system that converts a user's silent speech into text in real time and generates and transmits voice data. Furthermore, by combining it with an emotion engine, it is possible to recognize the emotion of the speech and reflect it in the text and voice. This system exchanges data between the user, terminal, and server to complete the overall processing.
[0855] Program processing explanation
[0856] When a user makes a silent speech, the system starts by capturing the movements of the user's lips and tongue. For example, consider a user who wants to express their opinion during a meeting. In this case, the user speaks only with the movements of their mouth.
[0857] User behavior
[0858] The user puts on a speakerphone or headset and begins to speak silently, and the camera responds, capturing high-resolution video of the user's lip and tongue movements.
[0859] Device behavior
[0860] The device captures the user's speech in real time and transmits the data to a server, which encodes the captured data into an optimized format and sends it to the server over a secure communication channel.
[0861] Server Operation
[0862] The server decodes the received data and uses an image processing AI model to analyze the user's lip and tongue movements. This analyzed data is converted into text. A generative AI model is then used to generate multiple utterance candidates based on the context of the conversation and the surrounding content. These candidates are then collated to determine the most suitable text.
[0863] Emotion Engine Operation
[0864] The server's emotion engine recognizes the user's emotions from silent speech and captured data. Based on this emotion information, the generated text and voice data are adjusted. For example, if the user is excited or depressed, the generated text and voice tone are adjusted to match that emotion.
[0865] Voice generation
[0866] The server then passes the determined text and emotional information to a text-to-speech (TTS) model, which then generates a voice that closely resembles the user's voice in real time. This voice data also reflects the recognized emotions.
[0867] Specific examples
[0868] For example, suppose a user makes a silent utterance such as, "How is the project progressing?". At this time, the user is slightly nervous. The movements of the user's lips and tongue are captured by the device and sent to the server. The server's image processing AI analyzes this movement and generates the text candidate, "How is the project progressing?". The generative AI model then generates utterance candidates that capture the context of the conversation and determines the most suitable text. The emotion engine recognizes the user's nervousness and adds a voice tone that reflects this. Based on this text and emotional information, the voice generation AI generates a voice that closely resembles the user's voice quality and sends it to the other party.
[0869] The person on the other end of the line can recognize that the user is silently asking, "How's the project going?" and sense the tension in their voice. This system allows users to maintain privacy and smooth communication that conveys emotion while maintaining the silence of the surrounding area.
[0870] The processing flow will be explained below.
[0871] Step 1:
[0872] The user initiates silent speech: the user begins speaking without making any sound using lip or tongue movements.
[0873] Step 2:
[0874] The device activates its built-in camera to capture high-resolution video of the user's lip and tongue movements, with the camera set to a high frame rate to ensure smooth recording of the mouth movements.
[0875] Step 3:
[0876] The device encodes the captured video data in real time, optimizing the data volume and converting it into a suitable format for transmission.
[0877] Step 4:
[0878] The device then transmits the encoded data to the server over a secure communication channel, where the communication is encrypted and privacy-protected.
[0879] Step 5:
[0880] The server decodes the received data, returning the encoded video data to a format that can be analyzed.
[0881] Step 6:
[0882] The server's image processing AI model analyzes the decoded video data and converts lip and tongue movements into text data.
[0883] Step 7:
[0884] The server's emotion engine recognizes emotions from video data and user speech, and this emotion information is used for subsequent processing.
[0885] Step 8:
[0886] The server receives text candidates generated by the image processing AI model, and uses the generative AI model to generate multiple utterance candidates based on the conversation context and surrounding content.
[0887] Step 9:
[0888] The server compares the generated utterance candidates and determines the most appropriate text, taking into account the context and past conversation history to select the most natural utterance.
[0889] Step 10:
[0890] The server adjusts the optimized text based on the emotional information from the emotion engine, adjusting the content and tone of the text according to the emotional information.
[0891] Step 11:
[0892] The server then passes the determined text and emotional information to a text-to-speech (TTS) model, which uses this information to generate a voice that closely resembles the user's voice in real time.
[0893] Step 12:
[0894] The server encodes the generated audio data and sends it to the device, where it is converted into a format suitable for playback.
[0895] Step 13:
[0896] The device decodes the received audio data, which is then in a format that can be played back.
[0897] Step 14:
[0898] The terminal plays back the decoded voice data and outputs it as voice to the other party, who can then recognize the content of the user's silent utterance as voice.
[0899] In this way, not only can the user's unvoiced utterances be conveyed to the other party in real time, but the user's emotions can also be reflected.
[0900] Example 2
[0901] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0902] Conventional communication systems make it difficult to ensure privacy through speech, making smooth communication difficult in public places or environments where silence is desired. Furthermore, when silent speech is used, there are no systems that can accurately convert speech from lip and tongue movements alone, and generate speech that reflects the user's emotions. Therefore, a new communication method that combines real-time conversion of silent speech into text and emotion recognition is needed.
[0903] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for capturing a user's silent utterance at high resolution, means for encoding the captured data, means for transmitting the encoded data to the server, means for decoding and analyzing the transmitted data, means for converting the analyzed data into text, means for generating a plurality of utterance candidates using a generative AI model, means for collating the generated utterance candidates to determine the optimal text, means for recognizing the user's emotion using an emotion engine, means for adjusting the text and voice data based on the recognized emotion information, means for converting the determined text into voice data, and means for transmitting the voice data to the other party. This enables smooth communication using silent utterances, converting them into text and voice in real time, and further reflecting the user's emotions in the voice.
[0904] "Capture" means recording the movements of the user's lips and tongue as high-resolution video.
[0905] "Encoding" refers to converting captured data into an optimized format.
[0906] "Sending to the server" means sending the encoded data to the server while taking security into consideration.
[0907] "Decoding" means restoring data received by the server to its original form.
[0908] "Analysis" refers to analyzing the movements of the user's lips and tongue based on the decoded data.
[0909] "Convert to text" means generating appropriate sentences from the analyzed data.
[0910] A "generative AI model" is an artificial intelligence that generates multiple utterance candidates based on textual data.
[0911] "Matching utterance candidates" refers to comparing the generated utterance candidates and selecting the most suitable text.
[0912] The "emotion engine" is a system that recognizes emotions from the user's silent utterances.
[0913] "Adjusting based on emotional information" means adjusting the tone and nuance of the generated text or audio data based on the emotions recognized by the emotion engine.
[0914] "Conversion to audio data" refers to converting the determined text into audio format.
[0915] "Sending voice data to the other party" means delivering voice data to the other party in real time.
[0916] This invention relates to a system that converts a user's silent speech into text in real time and generates and transmits voice data. Furthermore, by combining it with an emotion engine, it is possible to recognize the emotion of the speech and reflect it in the text and voice. This system exchanges data between the user, terminal, and server to complete the overall processing.
[0917] The user puts on a speakerphone or headset and begins to speak silently. The camera captures the user's lip and tongue movements as high-resolution video. For example, imagine a user silently says, "How's the project going?" during a meeting. The user's mouth movements are recorded in real time.
[0918] The device encodes the captured video data into an optimized format using video encoding software (e.g., FFmpeg), and then transmits the encoded data to the server over a secure communication channel (e.g., HTTPS or WebSocket).
[0919] The server first decodes the received data. Next, it uses an image processing AI model (e.g., OpenCV or TensorFlow) to analyze the user's lip and tongue movements. Text data is generated based on this analysis data. Next, it uses a generative AI model (e.g., GPT-3) to generate multiple utterance candidates based on the context of the conversation and the content before and after. These candidates are compared to determine the most appropriate text.
[0920] Next, the server's emotion engine (e.g., IBM Watson Emotion Analysis) recognizes the user's emotions from their silent speech and captured data. Based on this emotion information, the generated text and voice data are adjusted. For example, if the user is nervous, the generated text and voice tone are adjusted to match that emotion.
[0921] Finally, the server passes the determined text and emotional information to a speech generation AI (TTS) model (e.g., Google Text-to-Speech API), which generates a voice that closely resembles the user's voice in real time. This voice data is then transmitted to the other party via a secure communication channel.
[0922] For example, if a user silently says, "How's the project going?" and sounds a little nervous, the camera captures this movement and sends it from the device to the server. The server analyzes and converts the text, and a generative AI model generates utterance candidates that capture the context of the conversation and determines the most appropriate text. The emotion engine recognizes the user's nervousness and adds a voice tone that reflects this. Based on this text and emotional information, the voice generation AI generates a voice that closely resembles the user's voice quality and sends it to the other party.
[0923] An example of a prompt sentence might be:
[0924] "The user silently utters, 'How's the project going?' Please convert this into text. Also, the user seems a little nervous, so please generate a voice that reflects that emotion."
[0925] This allows users to maintain privacy and quiet surroundings while still being able to communicate smoothly and express their emotions.
[0926] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0927] Step 1:
[0928] The user puts on a speakerphone or headset and begins to speak silently. The camera captures the user's lip and tongue movements as high-resolution video. The input is the user's lip and tongue movements, and the output is the captured high-resolution video data. For example, if the user silently speaks, "How is the project going?", the camera records this movement in real time.
[0929] Step 2:
[0930] The device encodes the captured video data into an optimized format. The input is high-resolution video data, and the output is the encoded data. This process uses video encoding software (e.g., FFmpeg). Encoding reduces the data size and improves transmission efficiency.
[0931] Step 3:
[0932] The device sends the encoded data to the server via a secure communication channel (e.g., HTTPS or WebSocket). The input is the encoded data, and the output is the data sent to the server. The device encrypts the data when sending it to ensure its security. For example, the device sends the data using the SSL / TLS protocol.
[0933] Step 4:
[0934] The server decodes the received data. The input is the data sent from the device, and the output is the decoded data. The server then uses an image processing AI model (e.g., OpenCV or TensorFlow) to analyze the user's lip and tongue movements. This analysis generates the basic data for identifying what the user is saying.
[0935] Step 5:
[0936] The server converts the analyzed data into text. The input is data obtained by analyzing lip and tongue movements, and the output is text data. Next, a generative AI model (e.g., GPT-3) is used to generate multiple utterance candidates. The text is generated taking into account the context of the conversation and the content before and after.
[0937] Step 6:
[0938] The server compares the generated utterance candidates and determines the optimal text. The input is multiple utterance candidates, and the output is the optimal text. The generative AI model performs context analysis and selects the most appropriate candidate.
[0939] Step 7:
[0940] The server's emotion engine (e.g., IBM Watson Emotion Analysis) recognizes the user's emotions from the captured data of the user's silent speech. The input is real-time speech data and candidate utterance text, and the output is emotional information. Based on this emotional information, the generated text and voice data are adjusted.
[0941] Step 8:
[0942] The server passes the adjusted text and emotional information to a speech generation AI (TTS) model (e.g., Google Text-to-Speech API). The input is the adjusted text and emotional information, and the output is audio data. Based on this, the speech generation AI generates a voice that closely resembles the user's voice quality in real time.
[0943] Step 9:
[0944] The generated voice data is sent to the other party via a secure communication channel. The input is the voice data, and the output is the voice delivered to the other party. By listening to this voice, the other party can understand the silent message and its emotion that the user has made.
[0945] This allows users to communicate smoothly and convey their emotions without speaking.
[0946] (Application example 2)
[0947] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0948] When users communicate efficiently using silent speech in autonomous vehicles, they need a way to fully convey their emotions without being distracted by surrounding noise. However, existing technologies have had challenges, such as difficulty accurately grasping the user's intentions and emotions and generating unnatural speech. Furthermore, there has been a lack of means for quiet communication while driving or in noisy environments.
[0949] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0950] In this invention, the server includes means for capturing the user's silent utterances at high resolution, means for transmitting the captured data to the server, means for analyzing the transmitted data and converting it into text, means for generating a plurality of utterance candidates based on the converted text, means for comparing the generated utterance candidates to determine the most appropriate text, means for converting the determined text into voice data, means for transmitting the voice data to an in-vehicle audio device, and means for analyzing emotional information and reflecting it in the generated voice data. This enables the user to communicate effectively in an autonomous vehicle even without vocal utterances, and by generating natural voice that reflects emotions, smooth conversations can be held while maintaining quietness during the ride.
[0951] "User" means an end user of the System.
[0952] "Silent speech" is the act of communicating without making any sounds, but through the movement of the lips and tongue.
[0953] "High-resolution capture means" refers to equipment or methods for capturing the minute movements of a user's lips and tongue as high-resolution video data.
[0954] A "server" is a computer system for processing, storing, and analyzing data.
[0955] The "analysis means" refers to the algorithms or techniques used to analyze the received data and convert it into a particular format.
[0956] "Means for converting into text" refers to a technique or device that converts the analyzed data into text data in sentence format.
[0957] The "means for generating utterance candidates" refers to a technique or method for generating multiple candidate texts intended by the user based on the converted text data.
[0958] The "means for determining the optimal text" refers to an algorithm or process for selecting the optimal utterance from among the multiple generated candidate utterances.
[0959] The "means for converting into voice data" refers to a technique or method for generating voice data from the determined text.
[0960] "In-vehicle sound equipment" refers to speakers and sound systems installed inside autonomous vehicles.
[0961] "Means for analyzing emotional information" refers to technologies and algorithms that extract and analyze emotions from users' silent speech and captured data.
[0962] "Means for reflecting in voice data" refers to techniques or methods for reflecting analyzed emotional information in the tone and intonation of voice data.
[0963] The present invention is a system for enabling a user to communicate effectively using silent speech within an autonomous vehicle, and can be implemented using the following specific procedures and configurations.
[0964] System Configuration
[0965] User behavior
[0966] The user speaks silently. For example, if they want to give silent instructions while driving, they speak only with their mouths. The movements of the user's lips and tongue are captured by a high-resolution camera installed inside the car.
[0967] Device behavior
[0968] A system inside the autonomous vehicle captures the user's lip and tongue movements in real time and transmits the captured data to a server, where it is encoded in an optimized format and transmitted over a secure communication channel.
[0969] Server Operation
[0970] The server decodes the received data and uses an image processing AI model to analyze the user's lip and tongue movements. This analysis data is used to generate text data. A generative AI model is then used to generate multiple utterance candidates based on the conversation context and surrounding content. These candidates are then collated to determine the most suitable text.
[0971] Emotion Engine Operation
[0972] The server's emotion engine recognizes the user's emotions from silent speech and captured data. Based on this emotion information, the generated text and voice data are adjusted. For example, if the user is nervous or happy, the tone of the generated voice will be adjusted to match that emotion.
[0973] Audio Generation and Output
[0974] The server then passes the determined text and emotional information to a text-to-speech (TTS) model, which then generates a voice in real time that closely resembles the user's voice. This voice data is then output through the car's audio system.
[0975] Hardware and Software
[0976] This system uses the following hardware and software:
[0977] High-resolution camera: Used to capture the user's lip and tongue movements.
[0978] Server: Handles all of the analysis of received data, text generation, emotion recognition, and speech generation.
[0979] Image processing AI model: A model for analyzing the movements of the user's lips and tongue.
[0980] Generative AI model: A model for generating utterance candidates that capture the context of the conversation.
[0981] Emotion engine: A system for recognizing user emotions and reflecting them in voice data.
[0982] Speech generation AI (TTS) model: Generates speech based on generated text and emotional information.
[0983] Specific examples
[0984] For example, consider a case where a user silently utters, "Please tell me the distance to the next rest area." The user's lip and tongue movements are captured by a high-resolution camera inside the vehicle, and the data is sent to the server. The server's image processing AI model analyzes this movement and generates text data, "Please tell me the distance to the next rest area." Next, the generative AI model generates utterance candidates that capture the context, and the optimal text is determined. The emotion engine also recognizes the user's relaxed emotion and reflects that emotion in the generated voice data. Finally, the voice generation AI generates speech based on this text and emotional information, which is output through the vehicle's sound system.
[0985] Prompt Sentence Examples
[0986] "Please tell me how far it is to the next rest area."
[0987] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0988] Processing flow
[0989] Step 1:
[0990] The user makes silent speech. The user communicates without vocalization, using only lip and tongue movements. The input to this process is the user's lip and tongue movements, and the output is the captured video data.
[0991] Step 2:
[0992] The device uses a high-resolution camera to capture the user's lip and tongue movements. The input is the user's movements, and the output is high-resolution video data. The captured data is processed in real time.
[0993] Step 3:
[0994] The device sends the captured video data to the server. The input is the captured video data, and the output is the encoded data sent to the server over a secure communication channel.
[0995] Step 4:
[0996] The server decodes the received data and uses an image processing AI model to analyze the user's lip and tongue movements. The input is encoded video data, and the output is analyzed movement data. Text data is generated based on this analyzed data.
[0997] Step 5:
[0998] The server uses a generative AI model to generate multiple utterance candidates based on the converted text data. The input is text from the analyzed data, and the output is multiple generated utterance candidates. The utterance candidates are generated based on the context and surrounding content.
[0999] Step 6:
[1000] The server matches the generated candidate utterances and determines the best text. The input is multiple candidate utterances, and the output is the text that is judged to be the best. The matching process includes contextual analysis.
[1001] Step 7:
[1002] The emotion engine on the server recognizes the user's emotion from the user's silent utterances and captured data. The input is the analyzed data and text, and the output is the recognized emotion information.
[1003] Step 8:
[1004] The server passes the determined text and emotional information to a speech generation AI (TTS) model to generate speech data. The input is the optimal text and emotional information, and the output is the generated speech data. The speech data reflects the recognized emotion.
[1005] Step 9:
[1006] The server sends the generated voice data to the in-car audio system, which then outputs the voice from the car's speakers. The input is the generated voice data, and the output is the voice played in the car. This voice is natural and reflects the user's emotions.
[1007] Specific actions
[1008] As a specific example, when a user makes a silent utterance such as "Please tell me the distance to the next rest area," the process proceeds as follows:
[1009] 1. The user makes lip and tongue movements without speaking.
[1010] 2. The device captures the movement with a high-resolution camera.
[1011] 3. The device encodes the captured video data and sends it to the server.
[1012] 4. The server decodes and parses the received data.
[1013] 5. The server generates text data based on the analysis data.
[1014] 6. The generative AI model generates multiple utterance candidates.
[1015] 7. The server determines the best text.
[1016] 8. The emotion engine recognizes the user's emotions.
[1017] 9. Voice generation AI generates voice data based on optimal text and emotional information.
[1018] 10. The server sends the audio data to the car's audio device, which outputs it as audio from the car's speakers.
[1019] Prompt Sentence Examples
[1020] "Please tell me how far it is to the next rest area."
[1021] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1022] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1023] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1024] [Fourth embodiment]
[1025] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1026] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1027] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1028] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1029] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1030] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1031] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1032] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1033] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1034] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1035] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1036] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1037] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1038] The present invention relates to a system that converts a user's silent speech into text in real time, and generates and transmits voice data. This system exchanges data between the user, the terminal, and a server to complete the overall process.
[1039] Program processing explanation
[1040] When a user makes a silent speech, the system starts by capturing the movements of the user's lips and tongue. For example, consider a situation where a user wants to have a confidential conversation with a business partner at a cafe. In this case, the user speaks using only the movements of their mouth.
[1041] User behavior
[1042] The user puts on a speakerphone or headset and begins to speak silently, and the camera responds, capturing high-resolution video of the user's lip and tongue movements.
[1043] Device behavior
[1044] The device captures the user's speech in real time and transmits the data to a server, which encodes the captured data into an optimized format and sends it to the server over a secure communication channel.
[1045] Server Operation
[1046] The server decodes the received data and uses an image processing AI model to analyze the user's lip and tongue movements. This analyzed data is converted into text data. A generative AI model is then used to generate multiple utterance candidates based on the context of the conversation and the content before and after. These candidates are compared to determine the most suitable text. The determined text data is passed to a speech generation AI (TTS) model, which generates speech in real time that closely resembles the user's voice quality. The generated voice data is then encoded and sent to the device.
[1047] Specific examples
[1048] For example, suppose a user makes a silent utterance such as "How is the project progressing?" The user's lip and tongue movements are captured on the device and sent to the server. The server's image processing AI analyzes this movement and generates the text candidate "How is the project progressing?" The generative AI model then generates utterance candidates that capture the context of the conversation and determines the same text as the best candidate. Based on this text, the speech generation AI generates speech that closely resembles the user's voice quality and sends it to the other party.
[1049] The person on the other end of the line will feel as if the user is silently asking, "How's the project going?" This system allows users to maintain privacy and communicate smoothly while maintaining the surrounding silence.
[1050] The processing flow will be explained below.
[1051] Step 1:
[1052] The user initiates silent speech: The user silently initiates speech using lip and tongue movements.
[1053] Step 2:
[1054] The device activates its built-in camera to capture high-resolution video of the user's lip and tongue movements, with the camera set to a high frame rate to ensure smooth recording.
[1055] Step 3:
[1056] The device encodes the captured video data, optimizing the amount of data and converting it into a suitable format for transmission.
[1057] Step 4:
[1058] The device then transmits the encoded data to the server over a secure communication channel, where the communication is encrypted and privacy-protected.
[1059] Step 5:
[1060] The server decodes the received data, returning the encoded video data to a format that can be analyzed.
[1061] Step 6:
[1062] The server passes the decoded data to an image-processing AI model, which analyzes the user's lip and tongue movements and converts them into text suggestions.
[1063] Step 7:
[1064] The server receives text candidates generated by the image processing AI model, and uses the generative AI model to generate multiple utterance candidates based on the conversation context and surrounding content.
[1065] Step 8:
[1066] The server compares the generated utterance candidates and determines the most appropriate text, taking into account the context and past conversation history to select the most natural utterance.
[1067] Step 9:
[1068] The server then passes the determined text to a text-to-speech (TTS) model, which generates a voice in real time that closely resembles the user's voice quality.
[1069] Step 10:
[1070] The server encodes the generated audio data and sends it to the device, where it is converted into a format suitable for playback.
[1071] Step 11:
[1072] The device decodes the received audio data, which is then in a format that can be played back.
[1073] Step 12:
[1074] The terminal plays back the decoded voice data and outputs it as voice to the other party, who can then recognize the content of the user's silent utterance as voice.
[1075] In this way, the user's unvoiced speech can be transmitted to the other party in real time.
[1076] Example 1
[1077] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1078] Conventional voice communication systems require users to speak, resulting in issues of ambient noise and privacy. Furthermore, when silent speech is used, there is a lack of technology to accurately convert the content into text or speech, making smooth communication difficult. Therefore, there is a need for technology that can accurately convert silent speech into text and speech for communication.
[1079] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1080] In this invention, the server includes means for capturing a user's silent utterances at high resolution, means for encoding the captured data and transmitting it to the server, means for decoding the transmitted data, means for analyzing the user's lip and tongue movements using image processing on the decoded data, means for converting the analyzed data into text, means for generating a plurality of utterance candidates based on the converted text using a generative AI model, means for comparing the generated utterance candidates to determine the most appropriate text, means for converting the determined text into voice data using a voice generation AI model, and means for transmitting the voice data to the other party. This allows the user to accurately convert silent utterances into text and voice, enabling smooth communication.
[1081] "Silent speech" is a method in which a user speaks without making a sound, by moving their lips and tongue.
[1082] "High-resolution capture means" refers to high-precision cameras and sensors used to capture the movements of a user's lips and tongue in detail.
[1083] "Encoding" refers to the process of converting captured data into an optimized format, which preserves high quality information while minimizing data volume.
[1084] "Decoding" refers to the process of returning encoded data to its original form, making it possible to analyze the data.
[1085] "Image processing" refers to a series of calculations and algorithms used to extract and analyze useful information from captured video data, specifically analyzing the movements of the user's lips and tongue.
[1086] A "generative AI model" refers to an artificial intelligence model that generates text, speech, etc. based on given input data.
[1087] "Speech generation AI model" refers to an artificial intelligence model for generating natural-sounding speech based on text data.
[1088] The "means for converting into text" refers to a process for converting the analyzed motion data into a character string.
[1089] "Means for generating utterance candidates" refers to the process of using a generative AI model to create multiple utterance candidates based on input text data.
[1090] "Means for converting into audio data" refers to the process of generating audio based on text data and converting it into audio data format.
[1091] "Means for transmitting to the other party" refers to the communication means for delivering the final generated voice data to the other party via the terminal.
[1092] The present invention relates to a system that converts a user's silent speech into text in real time, and generates and transmits voice data. This system exchanges data between the user, the terminal, and a server to complete the overall process.
[1093] First, when a user makes a silent speech, the user uses a device equipped with a high-resolution camera to capture the speech. When the user puts on a speakerphone or headset and starts speaking silently, the camera captures this movement as high-resolution video. For example, imagine a situation where a user wants to have a confidential conversation with a business partner in a cafe. In this case, the user speaks only with the movement of their lips and tongue.
[1094] The device encodes the captured video data in real time and converts it into an optimized format (e.g., H.264). This encoding process uses a video processing library (e.g., FFmpeg). The encoded data is then sent to the server via a secure communication protocol (e.g., HTTPS).
[1095] The server decodes the received video data and uses image processing to analyze the user's lip and tongue movements. This analysis uses an image processing AI model (e.g., TensorFlow model). The analyzed movement data is converted into text data. For example, if a specific lip movement is determined to be a "P," the server adds a "P" to the text data.
[1096] Next, the server generates utterance candidates using a generative AI model (e.g., generative AI) based on the converted text data. The prompt text is entered as follows: "Generate a dialogue based on the content uttered silently by the user. The user is trying to say, 'How is the project progressing?'"
[1097] The generative AI model generates multiple utterance candidates and selects the most appropriate one. For example, it selects the most appropriate utterance based on the user's past conversation patterns and context. The selected text data is passed to a speech generation AI model (e.g., speech generation AI), which generates speech that closely resembles the user's voice quality.
[1098] The generated voice data is encoded (e.g., MP3 format) and sent to the device again via a secure communication protocol. The device decodes the received voice data and converts it into a playable format (e.g., WAV format). The decoded voice is played through the speaker and transmitted to the other party. For example, if a user utters a silent utterance such as "How is the project progressing?", the other party will feel as if the user is asking the question without making any sound.
[1099] According to the present invention, even if a user speaks silently, the content can be accurately converted into text and voice, enabling smooth communication.
[1100] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1101] Step 1:
[1102] The user initiates a silent utterance. The user puts on a speakerphone or headset and makes lip and tongue movements in front of a high-resolution camera. For example, the user might say, "How's the project going?" The input is lip and tongue movements, and the output is video data.
[1103] Step 2:
[1104] The device captures the user's lip and tongue movements as high-resolution video. The device's high-resolution camera captures the user's movements in real time and stores the captured video in a specific memory area. The input is the lip and tongue movements, and the output is high-resolution video data.
[1105] Step 3:
[1106] The device encodes the captured video data and sends it to the server via a secure communication protocol. The device uses a video processing library such as FFmpeg to encode the video data into H.264 format. The input is high-resolution video data, and the output is the encoded video data.
[1107] Step 4:
[1108] The server receives the encoded video data and decodes it. The server then converts the video data back to its original format using libraries such as FFmpeg, and passes it to the image processing AI model. The input is the encoded video data, and the output is the decoded video data.
[1109] Step 5:
[1110] The server analyzes the decoded video data and extracts the user's lip and tongue movements. The server uses an image processing AI model (e.g., TensorFlow model) to analyze the movements of each video frame. For example, a specific lip movement is determined to be "P." The input is the decoded video data, and the output is lip and tongue movement data.
[1111] Step 6:
[1112] The server converts lip and tongue movement data into text data. The server then maps the analyzed movement data into a string of characters and generates text data. For example, if it identifies the movement of the letter "P," it outputs the letter "P." The input is lip and tongue movement data, and the output is text data.
[1113] Step 7:
[1114] The server uses a generative AI model to generate utterance candidates based on text data. The prompt text is input as "Generate a dialogue based on the content uttered silently by the user. The user is trying to say, 'How is the project progressing?'" The generative AI model (e.g., generative AI) generates multiple utterance candidates based on the text data. The input is text data, and the output is a list of utterance candidates.
[1115] Step 8:
[1116] The server compares the generated utterance candidates and determines the optimal text. The server evaluates multiple candidates obtained from the generative AI model and selects the most appropriate utterance based on the user's past conversation patterns and context. The input is a list of utterance candidates, and the output is the optimal text.
[1117] Step 9:
[1118] The server uses a speech generation AI model to convert the optimal text into speech data. The server then passes the determined text to a speech generation AI model (e.g., speech generation AI) to generate speech that closely resembles the user's voice quality. The input is the optimal text, and the output is the generated speech data.
[1119] Step 10:
[1120] The server encodes the generated audio data and sends it to the device. The server encodes the generated audio data into MP3 format and sends it to the device via a secure communication protocol. The input is the generated audio data and the output is the encoded audio data.
[1121] Step 11:
[1122] The device decodes the audio data and plays it through the speaker. The device decodes the received audio data and converts it into a playable format (for example, WAV format). The decoded audio is played through the speaker and transmitted to the other party, making it sound as if the user is saying, "How's the project going?" The input is encoded audio data, and the output is the played audio.
[1123] (Application example 1)
[1124] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1125] In today's virtual stores, staff need to have strong communication skills to respond to customers smoothly and quickly. However, direct verbal communication can be difficult due to environmental noise and privacy issues. Furthermore, a method for silent communication using silent speech has yet to be established. This presents a challenge in providing effective customer service.
[1126] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1127] In this invention, the server includes means for capturing a user's silent utterances at high resolution, means for transmitting the captured data to the server, means for analyzing the transmitted data and converting it into text, means for generating a plurality of utterance candidates based on the converted text, means for comparing the generated utterance candidates and determining the most appropriate text, means for converting the determined text into voice data, means for transmitting the voice data to the other party, and means for the user to interact with customers silently in a virtual store to provide customer service support. This enables store staff to maintain a quiet environment while quickly and accurately interacting with customers through voiceless utterances.
[1128] "Silent speech" is speech that is made without producing any sound using lip or tongue movements.
[1129] "High-resolution capture means" refers to camera and sensor technology that accurately captures the subtle movements of the user's lips and tongue.
[1130] The "means for transmitting to the server" is a technology for transmitting the captured data to the server using a secure data communication protocol (e.g., gRPC or HTTP).
[1131] The "means of analyzing and converting into text" is a process of converting the received lip and tongue movements into text in natural language using image processing technology and natural language processing technology.
[1132] The "means for generating multiple candidate utterances" is a technology that uses a generative AI model to output multiple candidate utterances based on the context from text data.
[1133] The "means for determining the most appropriate text" is a process for selecting the most appropriate text for the context and situation from the multiple utterance candidates generated.
[1134] The "means of converting into voice data" refers to the process of generating voice that closely resembles the user's voice quality using voice synthesis technology based on the determined text data.
[1135] The "means for transmitting to the other party" is a communication technology for transferring the generated voice data to the other party in real time.
[1136] A "virtual store" is a store environment that uses virtual reality and augmented reality technology to offer products and services online.
[1137] The "means for users to provide customer service without vocal utterances to support customer service" is a technique in which staff members provide guidance and assistance to customers using vocal utterances.
[1138] The system embodying this invention converts a user's silent speech into text in real time, then converts it into voice data, and responds to customers in a virtual store. The system is composed of a user, a terminal, and a server.
[1139] The user puts on the smart glasses and begins silent speech. The smart glasses' camera captures the user's lip and tongue movements in high resolution. This data is encoded by the device in real time and transmitted to the server over a secure communication channel.
[1140] The server decodes the received data and uses an image processing AI model to analyze the user's lip and tongue movements. This analyzed data is converted into text data using a natural language processing library (e.g., spaCy or NLTK). A generative AI model is then used to generate multiple utterance candidates based on the context of the conversation and the surrounding content. These candidates are compared to determine the most suitable text. Next, a voice generation AI (Text-to-Speech, e.g., Google Text-to-Speech or Amazon Polly) uses the determined text to generate audio that closely resembles the user's voice quality. This audio data is encoded and sent back to the device via a secure communication channel, where it is played back to the customer.
[1141] As a specific example, consider the case where a customer asks about a new product in a virtual store. When the customer asks, "Do you have this product in other colors?", the staff user silently speaks into the smart glasses, "Do you have this product in other colors?" This speech is captured by a camera, and the encoded data is sent to the server. The server analyzes it and generates a similar text, "Do you have this product in other colors?" The generative AI model then determines that this text is appropriate for the context, and the speech generation AI generates a voice that is close to the staff member's voice quality. This voice data is sent to the device and played back to the customer in a natural way.
[1142] An example prompt for a generative AI model might look something like this:
[1143] "Convert the silent speech into text and generate voice data based on the following requirements: 'Speech content', Task: 'Guidance to customer', Voice quality: 'User's voice quality'"
[1144] This system allows store staff to provide quick and accurate customer service while maintaining a quiet and private environment.
[1145] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1146] Step 1:
[1147] The user puts on the smart glasses and begins silent speech, while the smart glasses' camera captures the user's lip and tongue movements in high resolution.
[1148] Input: User's lip and tongue movements
[1149] Output: High-resolution image data
[1150] How it works: The camera in the smart glasses captures the movements of the user's lips and tongue.
[1151] Step 2:
[1152] The device encodes the captured data in real time and transmits it to the server over a secure communication channel.
[1153] Input: High-resolution image data
[1154] Output: Encoded data
[1155] What it does: The data encoding module converts the image data into a compact format and sends it over the Internet.
[1156] Step 3:
[1157] The server decodes the received data and uses an image processing AI model to analyze the movements of the user's lips and tongue.
[1158] Input: Encoded data
[1159] Output: Analysis results (lip and tongue position information)
[1160] Specific operation: The data decoding module converts the encoded data back into the original image data, which is then analyzed by the image processing AI.
[1161] Step 4:
[1162] The server converts the analyzed data into text data using a natural language processing library.
[1163] Input: Analysis results (lip and tongue position information)
[1164] Output: Unedited text data
[1165] What it does: A natural language processing (NLP) library converts the parsed data into text.
[1166] Step 5:
[1167] The server uses a generative AI model to generate multiple utterance candidates based on the context of the conversation and the content before and after.
[1168] Input: Unedited text data
[1169] Output: Multiple utterance candidates
[1170] How it works: The generative AI model analyzes raw text and outputs multiple utterance candidates.
[1171] Step 6:
[1172] The server selects the most suitable text from the multiple generated utterance candidates.
[1173] Input: Multiple utterance candidates
[1174] Output: Optimal text
[1175] What it does: A contextual analysis algorithm selects the most appropriate text based on the context.
[1176] Step 7:
[1177] The server uses speech generation AI to convert the optimal text into a voice that closely resembles the user's voice quality.
[1178] Input: Optimal text
[1179] Output: Audio data
[1180] What it does: A text-to-speech engine converts text into speech.
[1181] Step 8:
[1182] The terminal receives the audio data from the server and plays it back to the customer.
[1183] Input: Audio data
[1184] Output: The audio to be played
[1185] Specific operation: The speaker of the smart glasses plays the audio data.
[1186] This process allows store staff to provide quick and accurate customer service while maintaining a quiet and private environment.
[1187] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1188] This invention relates to a system that converts a user's silent speech into text in real time and generates and transmits voice data. Furthermore, by combining it with an emotion engine, it is possible to recognize the emotion of the speech and reflect it in the text and voice. This system exchanges data between the user, terminal, and server to complete the overall processing.
[1189] Program processing explanation
[1190] When a user makes a silent speech, the system starts by capturing the movements of the user's lips and tongue. For example, consider a user who wants to express their opinion during a meeting. In this case, the user speaks only with the movements of their mouth.
[1191] User behavior
[1192] The user puts on a speakerphone or headset and begins to speak silently, and the camera responds, capturing high-resolution video of the user's lip and tongue movements.
[1193] Device behavior
[1194] The device captures the user's speech in real time and transmits the data to a server, which encodes the captured data into an optimized format and sends it to the server over a secure communication channel.
[1195] Server Operation
[1196] The server decodes the received data and uses an image processing AI model to analyze the user's lip and tongue movements. This analyzed data is converted into text. A generative AI model is then used to generate multiple utterance candidates based on the context of the conversation and the surrounding content. These candidates are then collated to determine the most suitable text.
[1197] Emotion Engine Operation
[1198] The server's emotion engine recognizes the user's emotions from silent speech and captured data. Based on this emotion information, the generated text and voice data are adjusted. For example, if the user is excited or depressed, the generated text and voice tone are adjusted to match that emotion.
[1199] Voice generation
[1200] The server then passes the determined text and emotional information to a text-to-speech (TTS) model, which then generates a voice that closely resembles the user's voice in real time. This voice data also reflects the recognized emotions.
[1201] Specific examples
[1202] For example, suppose a user makes a silent utterance such as, "How is the project progressing?". At this time, the user is slightly nervous. The movements of the user's lips and tongue are captured by the device and sent to the server. The server's image processing AI analyzes this movement and generates the text candidate, "How is the project progressing?". The generative AI model then generates utterance candidates that capture the context of the conversation and determines the most suitable text. The emotion engine recognizes the user's nervousness and adds a voice tone that reflects this. Based on this text and emotional information, the voice generation AI generates a voice that closely resembles the user's voice quality and sends it to the other party.
[1203] The person on the other end of the line can recognize that the user is silently asking, "How's the project going?" and sense the tension in their voice. This system allows users to maintain privacy and smooth communication that conveys emotion while maintaining the silence of the surrounding area.
[1204] The processing flow will be explained below.
[1205] Step 1:
[1206] The user initiates silent speech: the user begins speaking without making any sound using lip or tongue movements.
[1207] Step 2:
[1208] The device activates its built-in camera to capture high-resolution video of the user's lip and tongue movements, with the camera set to a high frame rate to ensure smooth recording of the mouth movements.
[1209] Step 3:
[1210] The device encodes the captured video data in real time, optimizing the data volume and converting it into a suitable format for transmission.
[1211] Step 4:
[1212] The device then transmits the encoded data to the server over a secure communication channel, where the communication is encrypted and privacy-protected.
[1213] Step 5:
[1214] The server decodes the received data, returning the encoded video data to a format that can be analyzed.
[1215] Step 6:
[1216] The server's image processing AI model analyzes the decoded video data and converts lip and tongue movements into text data.
[1217] Step 7:
[1218] The server's emotion engine recognizes emotions from video data and user speech, and this emotion information is used for subsequent processing.
[1219] Step 8:
[1220] The server receives text candidates generated by the image processing AI model, and uses the generative AI model to generate multiple utterance candidates based on the conversation context and surrounding content.
[1221] Step 9:
[1222] The server compares the generated utterance candidates and determines the most appropriate text, taking into account the context and past conversation history to select the most natural utterance.
[1223] Step 10:
[1224] The server adjusts the optimized text based on the emotional information from the emotion engine, adjusting the content and tone of the text according to the emotional information.
[1225] Step 11:
[1226] The server then passes the determined text and emotional information to a text-to-speech (TTS) model, which uses this information to generate a voice that closely resembles the user's voice in real time.
[1227] Step 12:
[1228] The server encodes the generated audio data and sends it to the device, where it is converted into a format suitable for playback.
[1229] Step 13:
[1230] The device decodes the received audio data, which is then in a format that can be played back.
[1231] Step 14:
[1232] The terminal plays back the decoded voice data and outputs it as voice to the other party, who can then recognize the content of the user's silent utterance as voice.
[1233] In this way, not only can the user's unvoiced utterances be conveyed to the other party in real time, but the user's emotions can also be reflected.
[1234] Example 2
[1235] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1236] Conventional communication systems make it difficult to ensure privacy through speech, making smooth communication difficult in public places or environments where silence is desired. Furthermore, when silent speech is used, there are no systems that can accurately convert speech from lip and tongue movements alone, and generate speech that reflects the user's emotions. Therefore, a new communication method that combines real-time conversion of silent speech into text and emotion recognition is needed.
[1237] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for capturing a user's silent utterance at high resolution, means for encoding the captured data, means for transmitting the encoded data to the server, means for decoding and analyzing the transmitted data, means for converting the analyzed data into text, means for generating a plurality of utterance candidates using a generative AI model, means for collating the generated utterance candidates to determine the optimal text, means for recognizing the user's emotion using an emotion engine, means for adjusting the text and voice data based on the recognized emotion information, means for converting the determined text into voice data, and means for transmitting the voice data to the other party. This enables smooth communication using silent utterances, converting them into text and voice in real time, and further reflecting the user's emotions in the voice.
[1238] "Capture" means recording the movements of the user's lips and tongue as high-resolution video.
[1239] "Encoding" refers to converting captured data into an optimized format.
[1240] "Sending to the server" means sending the encoded data to the server while taking security into consideration.
[1241] "Decoding" means restoring data received by the server to its original form.
[1242] "Analysis" refers to analyzing the movements of the user's lips and tongue based on the decoded data.
[1243] "Convert to text" means generating appropriate sentences from the analyzed data.
[1244] A "generative AI model" is an artificial intelligence that generates multiple utterance candidates based on textual data.
[1245] "Matching utterance candidates" refers to comparing the generated utterance candidates and selecting the most suitable text.
[1246] The "emotion engine" is a system that recognizes emotions from the user's silent utterances.
[1247] "Adjusting based on emotional information" means adjusting the tone and nuance of the generated text or audio data based on the emotions recognized by the emotion engine.
[1248] "Conversion to audio data" refers to converting the determined text into audio format.
[1249] "Sending voice data to the other party" means delivering voice data to the other party in real time.
[1250] This invention relates to a system that converts a user's silent speech into text in real time and generates and transmits voice data. Furthermore, by combining it with an emotion engine, it is possible to recognize the emotion of the speech and reflect it in the text and voice. This system exchanges data between the user, terminal, and server to complete the overall processing.
[1251] The user puts on a speakerphone or headset and begins to speak silently. The camera captures the user's lip and tongue movements as high-resolution video. For example, imagine a user silently says, "How's the project going?" during a meeting. The user's mouth movements are recorded in real time.
[1252] The device encodes the captured video data into an optimized format using video encoding software (e.g., FFmpeg), and then transmits the encoded data to the server over a secure communication channel (e.g., HTTPS or WebSocket).
[1253] The server first decodes the received data. Next, it uses an image processing AI model (e.g., OpenCV or TensorFlow) to analyze the user's lip and tongue movements. Text data is generated based on this analysis data. Next, it uses a generative AI model (e.g., GPT-3) to generate multiple utterance candidates based on the context of the conversation and the content before and after. These candidates are compared to determine the most appropriate text.
[1254] Next, the server's emotion engine (e.g., IBM Watson Emotion Analysis) recognizes the user's emotions from their silent speech and captured data. Based on this emotion information, the generated text and voice data are adjusted. For example, if the user is nervous, the generated text and voice tone are adjusted to match that emotion.
[1255] Finally, the server passes the determined text and emotional information to a speech generation AI (TTS) model (e.g., Google Text-to-Speech API), which generates a voice that closely resembles the user's voice in real time. This voice data is then transmitted to the other party via a secure communication channel.
[1256] For example, if a user silently says, "How's the project going?" and sounds a little nervous, the camera captures this movement and sends it from the device to the server. The server analyzes and converts the text, and a generative AI model generates utterance candidates that capture the context of the conversation and determines the most appropriate text. The emotion engine recognizes the user's nervousness and adds a voice tone that reflects this. Based on this text and emotional information, the voice generation AI generates a voice that closely resembles the user's voice quality and sends it to the other party.
[1257] An example of a prompt sentence might be:
[1258] "The user silently utters, 'How's the project going?' Please convert this into text. Also, the user seems a little nervous, so please generate a voice that reflects that emotion."
[1259] This allows users to maintain privacy and quiet surroundings while still being able to communicate smoothly and express their emotions.
[1260] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1261] Step 1:
[1262] The user puts on a speakerphone or headset and begins to speak silently. The camera captures the user's lip and tongue movements as high-resolution video. The input is the user's lip and tongue movements, and the output is the captured high-resolution video data. For example, if the user silently speaks, "How is the project going?", the camera records this movement in real time.
[1263] Step 2:
[1264] The device encodes the captured video data into an optimized format. The input is high-resolution video data, and the output is the encoded data. This process uses video encoding software (e.g., FFmpeg). Encoding reduces the data size and improves transmission efficiency.
[1265] Step 3:
[1266] The device sends the encoded data to the server via a secure communication channel (e.g., HTTPS or WebSocket). The input is the encoded data, and the output is the data sent to the server. The device encrypts the data when sending it to ensure its security. For example, the device sends the data using the SSL / TLS protocol.
[1267] Step 4:
[1268] The server decodes the received data. The input is the data sent from the device, and the output is the decoded data. The server then uses an image processing AI model (e.g., OpenCV or TensorFlow) to analyze the user's lip and tongue movements. This analysis generates the basic data for identifying what the user is saying.
[1269] Step 5:
[1270] The server converts the analyzed data into text. The input is data obtained by analyzing lip and tongue movements, and the output is text data. Next, a generative AI model (e.g., GPT-3) is used to generate multiple utterance candidates. The text is generated taking into account the context of the conversation and the content before and after.
[1271] Step 6:
[1272] The server compares the generated utterance candidates and determines the optimal text. The input is multiple utterance candidates, and the output is the optimal text. The generative AI model performs context analysis and selects the most appropriate candidate.
[1273] Step 7:
[1274] The server's emotion engine (e.g., IBM Watson Emotion Analysis) recognizes the user's emotions from the captured data of the user's silent speech. The input is real-time speech data and candidate utterance text, and the output is emotional information. Based on this emotional information, the generated text and voice data are adjusted.
[1275] Step 8:
[1276] The server passes the adjusted text and emotional information to a speech generation AI (TTS) model (e.g., Google Text-to-Speech API). The input is the adjusted text and emotional information, and the output is audio data. Based on this, the speech generation AI generates a voice that closely resembles the user's voice quality in real time.
[1277] Step 9:
[1278] The generated voice data is sent to the other party via a secure communication channel. The input is the voice data, and the output is the voice delivered to the other party. By listening to this voice, the other party can understand the silent message and its emotion that the user has made.
[1279] This allows users to communicate smoothly and convey their emotions without speaking.
[1280] (Application example 2)
[1281] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1282] When users communicate efficiently using silent speech in autonomous vehicles, they need a way to fully convey their emotions without being distracted by surrounding noise. However, existing technologies have had challenges, such as difficulty accurately grasping the user's intentions and emotions and generating unnatural speech. Furthermore, there has been a lack of means for quiet communication while driving or in noisy environments.
[1283] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1284] In this invention, the server includes means for capturing the user's silent utterances at high resolution, means for transmitting the captured data to the server, means for analyzing the transmitted data and converting it into text, means for generating a plurality of utterance candidates based on the converted text, means for comparing the generated utterance candidates to determine the most appropriate text, means for converting the determined text into voice data, means for transmitting the voice data to an in-vehicle audio device, and means for analyzing emotional information and reflecting it in the generated voice data. This enables the user to communicate effectively in an autonomous vehicle even without vocal utterances, and by generating natural voice that reflects emotions, smooth conversations can be held while maintaining quietness during the ride.
[1285] "User" means an end user of the System.
[1286] "Silent speech" is the act of communicating without making any sounds, but through the movement of the lips and tongue.
[1287] "High-resolution capture means" refers to equipment or methods for capturing the minute movements of a user's lips and tongue as high-resolution video data.
[1288] A "server" is a computer system for processing, storing, and analyzing data.
[1289] The "analysis means" refers to the algorithms or techniques used to analyze the received data and convert it into a particular format.
[1290] "Means for converting into text" refers to a technique or device that converts the analyzed data into text data in sentence format.
[1291] The "means for generating utterance candidates" refers to a technique or method for generating multiple candidate texts intended by the user based on the converted text data.
[1292] The "means for determining the optimal text" refers to an algorithm or process for selecting the optimal utterance from among the multiple generated candidate utterances.
[1293] The "means for converting into voice data" refers to a technique or method for generating voice data from the determined text.
[1294] "In-vehicle sound equipment" refers to speakers and sound systems installed inside autonomous vehicles.
[1295] "Means for analyzing emotional information" refers to technologies and algorithms that extract and analyze emotions from users' silent speech and captured data.
[1296] "Means for reflecting in voice data" refers to techniques or methods for reflecting analyzed emotional information in the tone and intonation of voice data.
[1297] The present invention is a system for enabling a user to communicate effectively using silent speech within an autonomous vehicle, and can be implemented using the following specific procedures and configurations.
[1298] System Configuration
[1299] User behavior
[1300] The user speaks silently. For example, if they want to give silent instructions while driving, they speak only with their mouths. The movements of the user's lips and tongue are captured by a high-resolution camera installed inside the car.
[1301] Device behavior
[1302] A system inside the autonomous vehicle captures the user's lip and tongue movements in real time and transmits the captured data to a server, where it is encoded in an optimized format and transmitted over a secure communication channel.
[1303] Server Operation
[1304] The server decodes the received data and uses an image processing AI model to analyze the user's lip and tongue movements. This analysis data is used to generate text data. A generative AI model is then used to generate multiple utterance candidates based on the conversation context and surrounding content. These candidates are then collated to determine the most suitable text.
[1305] Emotion Engine Operation
[1306] The server's emotion engine recognizes the user's emotions from silent speech and captured data. Based on this emotion information, the generated text and voice data are adjusted. For example, if the user is nervous or happy, the tone of the generated voice will be adjusted to match that emotion.
[1307] Audio Generation and Output
[1308] The server then passes the determined text and emotional information to a text-to-speech (TTS) model, which then generates a voice in real time that closely resembles the user's voice. This voice data is then output through the car's audio system.
[1309] Hardware and Software
[1310] This system uses the following hardware and software:
[1311] High-resolution camera: Used to capture the user's lip and tongue movements.
[1312] Server: Handles all of the analysis of received data, text generation, emotion recognition, and speech generation.
[1313] Image processing AI model: A model for analyzing the movements of the user's lips and tongue.
[1314] Generative AI model: A model for generating utterance candidates that capture the context of the conversation.
[1315] Emotion engine: A system for recognizing user emotions and reflecting them in voice data.
[1316] Speech generation AI (TTS) model: Generates speech based on generated text and emotional information.
[1317] Specific examples
[1318] For example, consider a case where a user silently utters, "Please tell me the distance to the next rest area." The user's lip and tongue movements are captured by a high-resolution camera inside the vehicle, and the data is sent to the server. The server's image processing AI model analyzes this movement and generates text data, "Please tell me the distance to the next rest area." Next, the generative AI model generates utterance candidates that capture the context, and the optimal text is determined. The emotion engine also recognizes the user's relaxed emotion and reflects that emotion in the generated voice data. Finally, the voice generation AI generates speech based on this text and emotional information, which is output through the vehicle's sound system.
[1319] Prompt Sentence Examples
[1320] "Please tell me how far it is to the next rest area."
[1321] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1322] Processing flow
[1323] Step 1:
[1324] The user makes silent speech. The user communicates without vocalization, using only lip and tongue movements. The input to this process is the user's lip and tongue movements, and the output is the captured video data.
[1325] Step 2:
[1326] The device uses a high-resolution camera to capture the user's lip and tongue movements. The input is the user's movements, and the output is high-resolution video data. The captured data is processed in real time.
[1327] Step 3:
[1328] The device sends the captured video data to the server. The input is the captured video data, and the output is the encoded data sent to the server over a secure communication channel.
[1329] Step 4:
[1330] The server decodes the received data and uses an image processing AI model to analyze the user's lip and tongue movements. The input is encoded video data, and the output is analyzed movement data. Text data is generated based on this analyzed data.
[1331] Step 5:
[1332] The server uses a generative AI model to generate multiple utterance candidates based on the converted text data. The input is text from the analyzed data, and the output is multiple generated utterance candidates. The utterance candidates are generated based on the context and surrounding content.
[1333] Step 6:
[1334] The server matches the generated candidate utterances and determines the best text. The input is multiple candidate utterances, and the output is the text that is judged to be the best. The matching process includes contextual analysis.
[1335] Step 7:
[1336] The emotion engine on the server recognizes the user's emotion from the user's silent utterances and captured data. The input is the analyzed data and text, and the output is the recognized emotion information.
[1337] Step 8:
[1338] The server passes the determined text and emotional information to a speech generation AI (TTS) model to generate speech data. The input is the optimal text and emotional information, and the output is the generated speech data. The speech data reflects the recognized emotion.
[1339] Step 9:
[1340] The server sends the generated voice data to the in-car audio system, which then outputs the voice from the car's speakers. The input is the generated voice data, and the output is the voice played in the car. This voice is natural and reflects the user's emotions.
[1341] Specific actions
[1342] As a specific example, when a user makes a silent utterance such as "Please tell me the distance to the next rest area," the process proceeds as follows:
[1343] 1. The user makes lip and tongue movements without speaking.
[1344] 2. The device captures the movement with a high-resolution camera.
[1345] 3. The device encodes the captured video data and sends it to the server.
[1346] 4. The server decodes and parses the received data.
[1347] 5. The server generates text data based on the analysis data.
[1348] 6. The generative AI model generates multiple utterance candidates.
[1349] 7. The server determines the best text.
[1350] 8. The emotion engine recognizes the user's emotions.
[1351] 9. Voice generation AI generates voice data based on optimal text and emotional information.
[1352] 10. The server sends the audio data to the car's audio device, which outputs it as audio from the car's speakers.
[1353] Prompt Sentence Examples
[1354] "Please tell me how far it is to the next rest area."
[1355] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1356] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1357] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1358] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1359] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1360] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1361] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1362] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1363] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1364] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1365] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1366] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1367] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1368] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1369] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1370] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1371] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1372] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1373] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1374] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1375] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1376] The following is further disclosed regarding the above embodiment.
[1377] (Claim 1)
[1378] a means for capturing a user's unspoken speech in high resolution;
[1379] means for transmitting the captured data to a server;
[1380] a means for parsing and converting the transmitted data into text;
[1381] a means for generating a plurality of utterance candidates based on the converted text;
[1382] a means for matching the generated candidate utterances to determine the best text;
[1383] means for converting the determined text into audio data;
[1384] means for transmitting voice data to the other party;
[1385] A system including:
[1386] (Claim 2)
[1387] 10. The system of claim 1, further comprising means for image processing lip and tongue movement data of the user's silent speech.
[1388] (Claim 3)
[1389] 10. The system of claim 1, further comprising means for generating a voice that approximates the user's voice quality in real time.
[1390] "Example 1"
[1391] (Claim 1)
[1392] a means for capturing a user's unspoken speech in high resolution;
[1393] means for encoding and transmitting the captured data to a server;
[1394] means for decoding the transmitted data;
[1395] A means of analyzing the movements of the user's lips and tongue using image processing of the decoded data;
[1396] a means for converting the parsed data into text;
[1397] A means for generating a plurality of candidate utterances based on the converted text using a generative AI model; and
[1398] a means for matching the generated candidate utterances to determine the best text;
[1399] A means for converting the determined text into speech data using a speech generation AI model;
[1400] means for transmitting voice data to the other party;
[1401] A system including:
[1402] (Claim 2)
[1403] 10. The system of claim 1, further comprising means for capturing and analyzing lip and tongue movement data of a user's silent speech in real time.
[1404] (Claim 3)
[1405] 10. The system of claim 1, further comprising means for generating a voice that approximates the user's voice quality.
[1406] "Application Example 1"
[1407] (Claim 1)
[1408] a means for capturing a user's unspoken speech in high resolution;
[1409] means for transmitting the captured data to a server;
[1410] a means for parsing and converting the transmitted data into text;
[1411] a means for generating a plurality of utterance candidates based on the converted text;
[1412] a means for matching the generated candidate utterances to determine the best text;
[1413] means for converting the determined text into audio data;
[1414] means for transmitting voice data to the other party;
[1415] A means for a user to respond to customers without vocal utterances in a virtual store to assist with customer service;
[1416] A system including:
[1417] (Claim 2)
[1418] 10. The system of claim 1, further comprising means for image processing lip and tongue movement data of the user's silent speech.
[1419] (Claim 3)
[1420] 10. The system of claim 1, further comprising means for generating a voice that approximates the user's voice quality in real time.
[1421] "Example 2: Combining Emotion Engines"
[1422] (Claim 1)
[1423] a means for capturing a user's unspoken speech in high resolution;
[1424] means for encoding the captured data;
[1425] means for transmitting the encoded data to a server;
[1426] means for decoding and analyzing the transmitted data;
[1427] A means for converting the analyzed data into text;
[1428] A means for generating a plurality of utterance candidates using a generative AI model;
[1429] a means for matching the generated candidate utterances to determine the best text;
[1430] a means for recognizing user emotions using an emotion engine;
[1431] a means for adjusting the text or audio data based on the recognized emotion information;
[1432] means for converting the determined text into audio data;
[1433] means for transmitting voice data to the other party;
[1434] A system including:
[1435] (Claim 2)
[1436] 10. The system of claim 1, further comprising means for image processing lip and tongue movement data of the user's silent speech.
[1437] (Claim 3)
[1438] 10. The system of claim 1, further comprising means for generating a voice that approximates the user's voice quality in real time.
[1439] "Application example 2 when combining emotion engines"
[1440] Claims based on new inventions
[1441] (Claim 1)
[1442] a means for capturing a user's unspoken speech in high resolution;
[1443] means for transmitting the captured data to a server;
[1444] a means for parsing and converting the transmitted data into text;
[1445] a means for generating a plurality of utterance candidates based on the converted text;
[1446] a means for matching the generated candidate utterances to determine the best text;
[1447] means for converting the determined text into audio data;
[1448] means for transmitting audio data to an in-vehicle audio device;
[1449] A means for analyzing emotional information and reflecting it in the generated voice data;
[1450] A system including:
[1451] (Claim 2)
[1452] 10. The system of claim 1, further comprising means for image processing lip and tongue movement data of the user's silent speech.
[1453] (Claim 3)
[1454] 10. The system of claim 1, further comprising means for generating a voice that approximates the user's voice quality in real time. [Explanation of symbols]
[1455] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. a means for capturing a user's unspoken speech in high resolution; means for transmitting the captured data to a server; a means for parsing and converting the transmitted data into text; a means for generating a plurality of utterance candidates based on the converted text; a means for matching the generated candidate utterances to determine the best text; means for converting the determined text into audio data; means for transmitting voice data to the other party; A system including:
2. 10. The system of claim 1, further comprising means for image processing lip and tongue movement data of the user's silent speech.
3. 10. The system of claim 1, further comprising means for generating a voice that approximates the user's voice quality in real time.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A