system
The system addresses language translation inaccuracies and real-time performance issues by employing speech recognition, a generative model, and speech synthesis, ensuring high accuracy and natural-sounding communication through user feedback integration.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-10-10
- Publication Date
- 2026-04-22
AI Technical Summary
Existing systems face challenges in achieving high accuracy and real-time performance for language translation, particularly in facilitating smooth communication between users speaking different languages, and lack efficient methods for incorporating user feedback to improve translation quality.
A system utilizing speech recognition for noise reduction and sound quality improvement, a generative model for real-time translation with user feedback integration, and speech synthesis for natural-sounding output, all processed on a cloud server to ensure real-time performance.
Enables seamless, real-time communication with high translation accuracy, allowing users to engage in natural-sounding conversations across languages by continuously improving translation quality through user feedback.
Smart Images

Figure 2026068445000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor and includes steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003] ]>
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] There is a need to provide a means for smoothly performing real-time communication between users who speak different languages. For this reason, a system for translating voice and text with high accuracy is required. However, in existing technologies, the translation accuracy is often low or the real-time performance is lacking, making sufficient communication difficult. And regarding the improvement of translation quality, there is also the problem of how to efficiently incorporate feedback from users.
Means for Solving the Problems
[0005] This invention provides a speech recognition means that converts speech data into text with high accuracy. This speech recognition means includes signal processing for noise reduction and sound quality improvement. It also translates text data into different languages in real time using a generative model means. This generative model means improves translation accuracy by updating model parameters using user feedback. Furthermore, it includes a speech synthesis means that converts the translated text into speech data, enabling communication via voice. Since speech recognition and generative model processing are performed on a cloud server via a communication means, real-time performance of the entire system is ensured.
[0006] "Speech recognition means" refers to a technology that has the function of analyzing speech data and converting it into corresponding text data.
[0007] A "generative model means" is a technology that provides an algorithm for generating text data in different languages based on input text data.
[0008] "Speech synthesis means" refers to a technology that has the function of generating speech data based on text data and outputting it as speech.
[0009] "Communication means" refers to the technology, including the hardware and software necessary to send and receive data via a computer network.
[0010] An "interface means" is a technology that provides an operating environment for a user to input or receive information from a system. [Brief explanation of the drawing]
[0011] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3]This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] This is a sequence diagram showing the processing flow of the data processing system in Example 2, which incorporates an emotion engine. [Figure 14] This is a sequence diagram showing the processing flow of the data processing system in Application Example 2, which combines an emotion engine. [Modes for carrying out the invention]
[0012] Hereinafter, an example of an embodiment of the system relating to the technology of this disclosure will be described with reference to the attached drawings.
[0013] First, let's explain the terminology used in the following explanation.
[0014] In the following embodiments, the labeled processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0015] In the following embodiments, the labeled RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0016] In the following embodiments, the labeled storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, etc.
[0017] In the following embodiments, the labeled communication I / F (Interface) is an interface including a communication processor and an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark), etc.
[0018] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0019] [First Embodiment]
[0020] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0021] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0022] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0023] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0024] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0025] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0026] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0027] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0028] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0029] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0030] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0031] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0032] This invention provides a voice and text translation system that enables smooth, real-time communication between users who speak different languages. Specific embodiments thereof are described below.
[0033] When a user begins speaking in their first language using the device, the device acquires audio data through its microphone. This audio data is analyzed by the device's speech recognition system, processed for noise reduction and sound quality improvement, and then converted into text data. The converted text data is then transmitted to a server via a communication device.
[0034] The server inputs the received text data into a generative model and performs translation into a second language. This generative model continuously updates its model parameters using user feedback, improving translation accuracy. The translated text data is then sent back to the terminal via communication.
[0035] The terminal processes the received translated text using speech synthesis and outputs it as synthesized speech through the speaker. This allows the user to hear the translated content in a natural-sounding voice. Furthermore, the user can input feedback on the translation quality through the terminal's interface. The input feedback is sent back to the server and used to improve translation accuracy in the future.
[0036] In this way, the system provides users with a real-time, highly accurate translation function, facilitating smooth communication. A concrete example of this embodiment is when a Japanese-speaking user says, "It's a nice day today," the speech recognition means converts this into text, the generative model means translates it into English as "It's a nice day today," and the speech synthesis means outputs it as English speech. Through this process, users can achieve seamless communication using different languages.
[0037] The following describes the processing flow.
[0038] Step 1:
[0039] When the user speaks in their first language, the device captures the audio data through the microphone. The captured audio data is immediately sent to the speech recognition system.
[0040] Step 2:
[0041] The terminal's voice recognition means analyzes the voice data and performs noise reduction and sound quality improvement. Subsequently, based on the analysis results, the voice is converted into text data. This text data is sent to the server via the communication means for processing by the generative model means.
[0042] Step 3:
[0043] The server inputs text data received via communication into a generation model. The generation model translates this text into a second language and generates highly accurate translated text data. The generated translated data is then transmitted back to the terminal via communication.
[0044] Step 4:
[0045] The terminal receives the translated text data from the server and inputs it into a speech synthesis device. The speech synthesis device converts the text data into speech data and generates synthesized speech. This speech data is output through the speaker, allowing the user to hear the translated content.
[0046] Step 5:
[0047] Users input feedback on translation quality using an interface on their device. This feedback information is sent to the server via a communication method.
[0048] Step 6:
[0049] The server incorporates the received feedback information into the generation model and updates the model parameters, thereby improving accuracy in subsequent translation processes.
[0050] (Example 1)
[0051] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0052] In communication between users who speak different languages, there is a need for real-time, highly accurate translation. However, conventional technologies have faced challenges in smooth communication due to insufficient speech recognition and translation accuracy, or unnaturalness in the resulting speech synthesis. Furthermore, there have been insufficient methods for efficiently utilizing user feedback to improve translation accuracy.
[0053] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0054] In this invention, the server includes speech recognition means for collecting audio data, performing noise reduction and sound quality improvement, and converting the results into text data; generative model means for translating the text data into different languages, and means for updating model parameters using user feedback; and speech synthesis means for converting the translated text data into audio data and outputting it as natural speech. This enables users to engage in real-time, effective, and natural two-way communication.
[0055] "Speech recognition means" refers to means that have the function of collecting speech data, removing noise and improving sound quality, and then converting that speech data into text data.
[0056] A "generative model means" is a means of translating input text data into different languages, and has the function of updating model parameters using user feedback to improve translation accuracy.
[0057] "Speech synthesis means" refers to a method that uses technology to convert translated text data into speech data and output it as natural-sounding speech.
[0058] "Communication means" refers to means for sending and receiving data obtained by speech recognition means and generative model means via a digital communication network.
[0059] "Interface means" refers to means that have input means for users to provide feedback on the quality of translations, and means for appropriately processing that information.
[0060] This invention provides a voice and text translation system for facilitating smooth, real-time communication between users who speak different languages. When a user uses a terminal and begins speaking in their first language, the terminal acquires voice data through its built-in microphone. The acquired voice data is analyzed within the terminal using speech recognition means, and after noise reduction and sound quality improvement, it is converted into text data. Speech recognition software such as Google® Speech-to-Text API can be used in this process. The converted text data is transmitted to a server via the terminal's communication means.
[0061] The server inputs the received text data into a generative model and performs translation into a second language. AI models such as OpenAI® GPT or DeepL API can be used for this process. The generative model has a mechanism to continuously improve translation accuracy by updating model parameters based on user feedback. The translated text data is sent back from the server to the terminal, where it is processed by a speech synthesis system and output as synthesized speech through the speaker. Technologies such as Google Text-to-Speech or Amazon Polly can be used for speech synthesis. This output allows the user to hear the translated content in a natural-sounding voice.
[0062] Furthermore, users can use an interface to input feedback on translation quality and send it to the server via their device. This feedback will be used to improve translation accuracy in the future.
[0063] As a concrete example, if a Japanese-speaking user says "It's a nice day today," the device can convert this into text, translate it into English as "It's a nice day today" using a generative model, and output it as English speech using a speech synthesis tool. An example of a prompt message could be, "Please convert the audio data into text in real time, translate it into a different language, and output it as speech." This system enables users to communicate seamlessly between different languages.
[0064] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0065] Step 1:
[0066] When a user begins speaking in their first language using the device, the device acquires audio data through its microphone. The input is the user's voice, and the output is digital audio data. This audio data is then prepared for processing by the device's built-in speech recognition system.
[0067] Step 2:
[0068] The terminal performs signal processing on the acquired audio data to remove noise and improve sound quality. This step typically utilizes a digital signal processor (DSP). The input is raw audio data, while the output is audio data with noise removed and improved sound quality. This signal processing ensures that high-quality audio data is available for speech recognition.
[0069] Step 3:
[0070] The device uses speech recognition to convert the improved audio data into text data. The input here is the processed audio data, and the output is the corresponding text data. For speech recognition, for example, the Google Speech-to-Text API is used. This text data forms the basis for subsequent translation processing.
[0071] Step 4:
[0072] The terminal sends the converted text data to the server using a communication method. The input is text data, and the output is the arrival of that text data at the server. Network connections such as Wi-Fi or mobile data are used for communication.
[0073] Step 5:
[0074] The server inputs the received text data into a generative AI model and performs translation into a second language. The input is the transmitted text data, and the output is the translated text data. The generative AI model employs AI technologies such as OpenAI GPT, which results in highly accurate translations.
[0075] Step 6:
[0076] The server sends the translated text data back to the terminal. The input is the translated text generated by the generative AI model, and the output is the terminal receiving that data. This communication also takes place over Wi-Fi or a mobile data network.
[0077] Step 7:
[0078] The terminal processes the received translated text data using a speech synthesis system and outputs it as synthesized speech through the speaker. The input is the translated text data sent from the server, and the output is the speech heard by the user. This speech synthesis uses tools such as Amazon Polly or Google Text-to-Speech to generate natural-sounding speech.
[0079] Step 8:
[0080] Users input feedback on translation quality using the terminal interface. The input is the user's rating, and the output is the transmission of this feedback data to the server. The server uses this feedback to improve the accuracy of the generating AI model.
[0081] (Application Example 1)
[0082] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0083] Smooth two-way communication between users who speak different languages in a virtual sales environment presents challenges, such as misunderstandings and communication stagnation due to language barriers, which affect user satisfaction and sales efficiency. Furthermore, while natural-sounding real-time translated speech output requires high accuracy and speed, translation systems must also be able to generate translations flexibly, reflecting the user's intent and context.
[0084] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0085] In this invention, the server includes speech recognition means for converting speech expressions into text expressions, generative model means for converting the text expressions into different symbolic expressions, speech synthesis means for converting the converted symbolic expressions into speech data, and conversation support means for realizing two-way communication in different languages in a virtual sales environment. This enables real-time and highly accurate speech translation between users who speak different languages, eliminates language barriers in a virtual sales environment, and makes communication more efficient for users.
[0086] "Speech representation" refers to information in a form that can be perceived as sound waves, and is primarily generated by humans or machines for the purpose of linguistic communication.
[0087] "Written representation" refers to a form of expression that uses letters or symbols to represent sounds or concepts, and is used for recording and transmitting information.
[0088] "Different symbolic representation" refers to a representation using characters or symbols in a language different from the original language, and its purpose is to make information understandable to speakers of other languages.
[0089] A "generative model" is a computational algorithm used to translate or generate natural language based on input data, and is implemented using neural networks and other methods.
[0090] "Audio data" refers to audio signals that have been processed and analyzed in a digital format, and is primarily used to facilitate communication and recording.
[0091] A "computer network" is an infrastructure for the mutual exchange of information and data between computers, and the internet is one example of this.
[0092] A "translation symbolic representation" is the result of accurately converting the written representation of the original language into another language, and is generated through the translation process.
[0093] "User" refers to an individual or organization that operates a system or device, and is an entity that utilizes the system's functions to achieve a specific purpose.
[0094] "Evaluation information" refers to the feedback and reviews that users provide to the system, and this data is used for future improvements and accuracy enhancements.
[0095] The system for implementing this invention combines various means to enable communication between users in different languages. First, the terminal uses speech recognition software to accept voice input. Specifically, it utilizes the speech_recognition library to convert voice data into text. This allows users to naturally begin a conversation.
[0096] Once the audio data is converted into text, that text is sent to a server via a communication method. The server uses a generative AI model to translate the received text into a different symbolic representation. Here, machine learning is used in the generative model to improve translation accuracy, and GPT-3® or similar models may be applied. The translated symbolic representation is then sent back to the terminal.
[0097] The terminal converts the received symbolic representations into speech data using speech synthesis means and provides it to the user in natural-sounding speech using a speech synthesis library such as gTTS. Furthermore, the user can evaluate the quality of the translated content through an interface means, and this evaluation is sent to the server and used to improve translation accuracy in the future.
[0098] For example, if a Japanese-speaking user visiting a virtual sales environment asks, "Could you tell me about the features of this product?", the salesperson will respond in English, "This product is eco-friendly and cost-effective." The terminal then translates this into Japanese and outputs the translated audio in Japanese. In this way, smooth communication is achieved between users who speak different languages.
[0099] An example of a prompt message would be, "Recognize the user's question, translate it into English, answer the question, translate that answer back into Japanese, and play it aloud." This allows the system to operate as planned and support the user's intended communication.
[0100] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0101] Step 1:
[0102] The device receives voice input from the user. It acquires voice data through the microphone and converts this voice data into text using the speech_recognition library. The input is the user's voice, and the output is the textual representation of that voice. Specifically, the voice signal is captured, and speech recognition is performed in the backend process.
[0103] Step 2:
[0104] After the character representation is generated, the terminal sends this data to the server via a communication method. The input is the character representation generated in the previous step, and the output is the completion of the data transmission to the server. In this operation, data is securely transferred using a network protocol.
[0105] Step 3:
[0106] The server uses a generative AI model to translate the received text representation into a symbolic representation in a different language. The input is the received text representation, and the output is the translated symbolic representation. For data processing, machine learning algorithms are used, and the weights within the model are utilized to perform highly accurate language translation.
[0107] Step 4:
[0108] The server sends the translated symbolic representation back to the terminal. Here, the input is the translated symbolic representation, and the output is the completion of the transmission to the terminal. Specifically, the data is packaged according to the server's transmission protocol and transferred over the network.
[0109] Step 5:
[0110] The terminal reconstructs the received symbolic representation as audio data using speech synthesis technology and outputs the audio using gTTS or similar methods. The input is the symbolic representation received from the server, and the output is the synthesized speech. During the conversion to audio data, speech with natural intonation is generated and played back through the speaker.
[0111] Step 6:
[0112] Users provide feedback on translation quality through the terminal interface. This feedback is sent to the server and used to improve translation accuracy in the future. The input is the user's feedback data, and the output is the completion of sending the feedback to the server. Specifically, the user input on the interface is captured and then sent back to the server.
[0113] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0114] This invention is a real-time translation system that takes user emotions into account. This system is comprised of a combination of speech recognition, a generative model, speech synthesis, communication, an interface, and an emotion engine.
[0115] When a user begins speaking in their first language, the device captures their voice using its microphone. The captured audio data is converted into text by a speech recognition system, and then an emotion engine extracts additional information, including the user's emotions. This emotion information is sent to the server along with the text data.
[0116] The server uses the received text data and sentiment information to perform translation into a second language using a generative model. During this process, the sentiment engine provides sentiment information to the generative model, and the translated expression is adjusted to best reflect the user's emotions. The translated text data is then transmitted to the terminal via a communication device.
[0117] The device converts this translated text into speech data using speech synthesis technology and outputs it through the speaker as natural-sounding speech that takes emotions into consideration. This allows the user to convey their intended emotions and nuances.
[0118] Furthermore, the system includes a function to receive user feedback and send it to the server. This feedback is used to refine the generative model and sentiment engine, improving the accuracy of translation and sentiment recognition.
[0119] For example, if a user gets a little excited and says "I'm so happy today!" in Japanese, the system recognizes that emotion and translates it into English as "I'm super happy today!" with emotion, outputting it in appropriate voice. This ensures that the user's emotions are accurately conveyed to others, facilitating smoother communication.
[0120] The following describes the processing flow.
[0121] Step 1:
[0122] When a user speaks in their first language, the device uses its microphone to capture the audio. This allows the spoken content to be obtained as digital audio data.
[0123] Step 2:
[0124] The device analyzes the captured audio data using speech recognition and converts it into corresponding text data. During this process, signal processing is also performed to remove noise and improve sound quality.
[0125] Step 3:
[0126] The device passes the converted text data to an emotion engine to analyze the user's emotions and extract emotional information. This emotional information indicates what emotions are conveyed in the utterance.
[0127] Step 4:
[0128] The device sends a dataset containing text data and sentiment information to the server via a communication method. This enables translation processing on the server side.
[0129] Step 5:
[0130] The server inputs the received text data into a generative model and performs translation into a second language. During this process, the emotional information provided by the emotion engine is taken into consideration, and the translated text is adjusted to appropriately express the user's emotions.
[0131] Step 6:
[0132] The server returns the translated text data to the terminal. The returned data includes expressions that reflect emotions.
[0133] Step 7:
[0134] The device processes the translated text data using speech synthesis technology to generate emotionally appropriate audio data. The generated audio is then delivered to the user through the speaker.
[0135] Step 8:
[0136] Users review the translated audio and provide quality feedback through their device's interface. This feedback is then sent back to the server via communication channels and used to improve the system in the future.
[0137] (Example 2)
[0138] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0139] In translation systems, it is difficult to appropriately reflect user emotions and incorporate natural emotional expressions into the translation results. Furthermore, conventional translation systems have the challenge of not being able to effectively utilize user feedback to improve translation accuracy.
[0140] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0141] In this invention, the server includes acoustic analysis means, emotion analysis means, and generative model means. This enables natural translation expressions that reflect the user's emotions.
[0142] "Acoustic analysis means" refers to technology that has the function of collecting audio data and converting it into text data.
[0143] "Emotional analysis means" refers to a technology that processes data to extract user emotional information from text data.
[0144] A "generative model means" is a technology that uses received text data and sentiment information to translate into different languages and express emotions appropriately.
[0145] "Sound generation means" refers to a technology that generates audio data based on translated text data and emotional information, and outputs it as audio that takes emotional expression into account.
[0146] "Communication control means" refers to technology that controls the processing for sending and receiving voice data and text data over a network.
[0147] "Input device means" refers to technology that provides an interface for receiving feedback from users.
[0148] This invention provides a system that enables advanced real-time translation that takes user emotions into consideration. Specific embodiments are shown below.
[0149] First, the device uses a microphone to capture audio data of the user's first language. Speech recognition software is then used as an acoustic analysis tool to convert this audio data into text. Commercially available speech recognition technologies, such as the Google Speech-to-Text API, can be used for this process.
[0150] Next, the terminal extracts user emotion information from the converted text data using emotion analysis tools. Here, natural language processing models and machine learning algorithms can be utilized as software widely used for emotion analysis. The emotion information obtained through this process is sent to the server along with the text data.
[0151] The server uses a generative model to perform advanced translation based on the received text data and sentiment information. In this embodiment, the generative AI model can be, for example, an open-source model or a commercial advanced natural language processing model. The generative model utilizes sentiment information to adjust the expression of the translation, enabling it to accurately convey the emotions intended by the user.
[0152] The translation results are sent from the server to the terminal and converted into audio data using speech synthesis software as a means of generating sound. Natural-sounding speech can be achieved by using Amazon Polly or other commercial speech synthesis services. The output is then delivered to the user through the speaker with appropriate intonation.
[0153] As a concrete example, if a user says "I'm so happy today!", the system takes the emotion into consideration and uses the prompt "Translate the following excited expression from Japanese to English while preserving the user's emotional excitement" to generate the translation "I'm super happy today!", which is then appropriately output as speech.
[0154] Furthermore, the terminal receives feedback from the user regarding the translation results and sends this feedback to the server using communication control means. This feedback is used to improve the accuracy of the generative model and sentiment analysis. In this way, the system can continuously improve the quality of translations and the accuracy of sentiment.
[0155] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0156] Step 1:
[0157] The device captures the user's voice using a microphone. The user's voice signal is used as input. This voice signal is acquired as analog data and converted to a digital format. The output is digitized voice data. This data is used for acoustic analysis in the next step.
[0158] Step 2:
[0159] The terminal inputs the captured audio data into speech recognition software, which converts it into text data. This uses acoustic analysis, and the specific data processing involves breaking down the audio signal into phonemes and converting them into strings. The output is text data representing what the user said as a string.
[0160] Step 3:
[0161] The terminal passes text data to an emotion analysis device to extract the user's emotions. The input in this step is text data obtained from a speech recognition device. Specifically, a natural language processing algorithm is used to analyze words and contexts that represent specific emotions within the text, and the emotion information is output as data. The output is metadata that includes emotional elements such as joy, anger, sadness, and happiness.
[0162] Step 4:
[0163] The terminal transmits text data and sentiment information to the server via a communication control means. The input here is text data and its sentiment metadata. A secure communication protocol is used to ensure data integrity when transmitting to the server. As output, the data received on the server side is provided to the generation model means.
[0164] Step 5:
[0165] The server inputs the received text data and sentiment information into a generative AI model to perform translation. In this step, the translation model performs data calculations to adjust the translation expression to reflect sentiment while performing language conversion based on the input data. Specifically, it uses prompt sentences to activate the generative AI model and outputs the optimal translation result. The output is a sentiment-sensitive translated text in a second language.
[0166] Step 6:
[0167] The server sends the translated text data to the terminal. The input is the translated text generated by the generative AI model. Communication control means are used for data transfer. The output is the translated text received on the terminal side.
[0168] Step 7:
[0169] The terminal inputs translated text data into a speech synthesis system and converts it into speech data. The input is the translated text data. Specifically, the speech synthesis engine converts the text data into a speech waveform and synthesizes it as a natural-sounding voice with emotion. The output is the speech data played back through the speaker.
[0170] Step 8:
[0171] The terminal receives feedback from the user and sends it to the server. An interface is used for inputting the user's opinion and emotional response to the translation results. The feedback data sent to the server contributes to improving the generative AI model and sentiment analysis methods. The output is improvement information that will be useful for future translation accuracy improvements.
[0172] (Application Example 2)
[0173] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".
[0174] Conventional translation systems can perform simple conversions between languages, but they have the problem of failing to adequately reflect the speaker's emotions and nuances, making natural communication difficult. In particular, in the field of content distribution, there is a demand for translations that allow viewers to accurately understand and enjoy the emotions of movies and dramas, but current technology is unable to adequately meet this requirement.
[0175] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0176] In this invention, the server includes speech recognition means for converting speech data into text data, generative model means for translating text data into different languages based on emotional information, and speech synthesis means for converting translated text data that reflects emotional information into speech data. This enables natural translation that takes the user's emotions into account.
[0177] "Speech recognition means" refers to a technology that converts speech data into text data, making it possible to process the user's speech content as text information.
[0178] A "generative model means" is a technology for translating text into different languages using text data and sentiment information obtained by speech recognition means.
[0179] "Speech synthesis means" refers to a technology that converts translated text data into speech data and provides the user with natural-sounding speech output.
[0180] "Communication means" refers to the technology for sending and receiving text data obtained by speech recognition means, emotion information, and translated text data obtained by generative model means via a communication network.
[0181] An "interface means" is a means for users to input feedback regarding the quality and emotional expression of translations, and enables interaction with the user.
[0182] "Emotional information" refers to information extracted by analyzing the emotions and nuances contained in the user's utterance, and is an important element in the translation process.
[0183] The system for implementing this invention consists of a user terminal and a server. The terminal is equipped with speech recognition means, communication means, and speech synthesis means, which convert the user's voice into text data in real time and extract emotional information. The voice data captured on the terminal is then converted into text data through speech recognition software. At this time, an emotional recognition algorithm is used to analyze the user's emotions and generate corresponding emotional information.
[0184] Next, this text data and sentiment information are sent to the server via a communication method. The server uses a generative AI model to translate the text data into different languages. At this time, the sentiment information is input into the generative model, so the translated text takes sentiment into account. The translated text data is then sent back to the terminal.
[0185] The device uses the received translated text data to create natural-sounding, emotionally resonant voice data through speech synthesis. The user can then listen to the translated content through the audio output from the speaker.
[0186] As a concrete example, when a user is watching a movie, the emotionally rich lines spoken by the characters can be instantly translated into different languages, taking emotional information into account, and delivered to the viewer. Through this process, users can accurately understand the emotions and deepen their overall impression of the movie.
[0187] Examples of prompt messages are as follows:
[0188] "When Elsa begins to sing in the snowy mountains, feeling alone but yearning for freedom, please generate a translation that reflects her emotions and strength."
[0189] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0190] Step 1:
[0191] The terminal receives the user's voice as input and captures the audio data using the microphone. The captured audio data is sent to a speech recognition system. The speech recognition system processes the audio data to convert it into text data and passes this converted text data to the next step.
[0192] Step 2:
[0193] The device uses text data obtained by speech recognition to run an emotion recognition algorithm. This algorithm receives text data as input, analyzes the user's emotions from it, and extracts emotion information. The emotion information is output as data representing the emotions expressed in the user's speech.
[0194] Step 3:
[0195] Text data and sentiment information are transmitted from the terminal to the server using a communication method. The server then uses the acquired data as input and runs a generative AI model. The generative AI model translates the text into different languages based on the input data. In this process, sentiment information is reflected in the translation result, and the output is translated text data with adjusted expression.
[0196] Step 4:
[0197] The server sends the translated text data to the terminal. The terminal passes this received data as input to the speech synthesis system. The speech synthesis system converts the translated text data into speech data and generates speech that has been adjusted to convey emotions naturally.
[0198] Step 5:
[0199] The device outputs audio data synthesized by a speech synthesis system to the user through its speaker. This allows the user to hear the translated content in audio format.
[0200] Step 6:
[0201] Users input feedback on the provided translation into their device. This feedback is collected through the interface as opinions on the translation quality and sentiment expression, and is sent back to the server. The server uses this feedback information to improve the accuracy of its generative AI model and sentiment recognition.
[0202] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0203] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0204] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0205] [Second Embodiment]
[0206] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0207] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0208] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0209] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0210] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0211] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0212] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0213] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0214] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0215] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0216] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0217] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0218] This invention provides a voice and text translation system that enables smooth, real-time communication between users who speak different languages. Specific embodiments thereof are described below.
[0219] When a user begins speaking in their first language using the device, the device acquires audio data through its microphone. This audio data is analyzed by the device's speech recognition system, processed for noise reduction and sound quality improvement, and then converted into text data. The converted text data is then transmitted to a server via a communication device.
[0220] The server inputs the received text data into a generative model and performs translation into a second language. This generative model continuously updates its model parameters using user feedback, improving translation accuracy. The translated text data is then sent back to the terminal via communication.
[0221] The terminal processes the received translated text using speech synthesis and outputs it as synthesized speech through the speaker. This allows the user to hear the translated content in a natural-sounding voice. Furthermore, the user can input feedback on the translation quality through the terminal's interface. The input feedback is sent back to the server and used to improve translation accuracy in the future.
[0222] In this way, the system provides users with a real-time, highly accurate translation function, facilitating smooth communication. A concrete example of this embodiment is when a Japanese-speaking user says, "It's a nice day today," the speech recognition means converts this into text, the generative model means translates it into English as "It's a nice day today," and the speech synthesis means outputs it as English speech. Through this process, users can achieve seamless communication using different languages.
[0223] The following describes the processing flow.
[0224] Step 1:
[0225] When the user speaks in their first language, the device captures the audio data through the microphone. The captured audio data is immediately sent to the speech recognition system.
[0226] Step 2:
[0227] The terminal's voice recognition means analyzes the voice data and performs noise reduction and sound quality improvement. Subsequently, based on the analysis results, the voice is converted into text data. This text data is sent to the server via the communication means for processing by the generative model means.
[0228] Step 3:
[0229] The server inputs text data received via communication into a generation model. The generation model translates this text into a second language and generates highly accurate translated text data. The generated translated data is then transmitted back to the terminal via communication.
[0230] Step 4:
[0231] The terminal receives the translated text data from the server and inputs it into a speech synthesis device. The speech synthesis device converts the text data into speech data and generates synthesized speech. This speech data is output through the speaker, allowing the user to hear the translated content.
[0232] Step 5:
[0233] Users input feedback on translation quality using an interface on their device. This feedback information is sent to the server via a communication method.
[0234] Step 6:
[0235] The server incorporates the received feedback information into the generation model and updates the model parameters, thereby improving accuracy in subsequent translation processes.
[0236] (Example 1)
[0237] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0238] In communication between users who speak different languages, there is a need for real-time, highly accurate translation. However, conventional technologies have faced challenges in smooth communication due to insufficient speech recognition and translation accuracy, or unnaturalness in the resulting speech synthesis. Furthermore, there have been insufficient methods for efficiently utilizing user feedback to improve translation accuracy.
[0239] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0240] In this invention, the server includes speech recognition means for collecting audio data, performing noise reduction and sound quality improvement, and converting the results into text data; generative model means for translating the text data into different languages, and means for updating model parameters using user feedback; and speech synthesis means for converting the translated text data into audio data and outputting it as natural speech. This enables users to engage in real-time, effective, and natural two-way communication.
[0241] "Speech recognition means" refers to means that have the function of collecting speech data, removing noise and improving sound quality, and then converting that speech data into text data.
[0242] A "generative model means" is a means of translating input text data into different languages, and has the function of updating model parameters using user feedback to improve translation accuracy.
[0243] "Speech synthesis means" refers to a method that uses technology to convert translated text data into speech data and output it as natural-sounding speech.
[0244] "Communication means" refers to means for sending and receiving data obtained by speech recognition means and generative model means via a digital communication network.
[0245] "Interface means" refers to means that have input means for users to provide feedback on the quality of translations, and means for appropriately processing that information.
[0246] This invention provides a voice and text translation system for facilitating smooth, real-time communication between users who speak different languages. When a user uses a terminal and begins speaking in their first language, the terminal acquires voice data through its built-in microphone. The acquired voice data is analyzed using speech recognition means within the terminal, and after noise reduction and sound quality improvement, it is converted into text data. Speech recognition software such as the Google Speech-to-Text API can be used in this process. The converted text data is transmitted to a server via the terminal's communication means.
[0247] The server inputs the received text data into a generative model and performs translation into a second language. AI models such as OpenAI GPT or DeepL API can be used for this process. The generative model has a mechanism to continuously improve translation accuracy by updating its model parameters based on user feedback. The translated text data is sent back from the server to the terminal, where it is processed by a speech synthesis system and output as synthesized speech through the speaker. Technologies such as Google Text-to-Speech or Amazon Polly can be used for speech synthesis. This output allows the user to hear the translated content in a natural-sounding voice.
[0248] Furthermore, users can use an interface to input feedback on translation quality and send it to the server via their device. This feedback will be used to improve translation accuracy in the future.
[0249] As a concrete example, if a Japanese-speaking user says "It's a nice day today," the device can convert this into text, translate it into English as "It's a nice day today" using a generative model, and output it as English speech using a speech synthesis tool. An example of a prompt message could be, "Please convert the audio data into text in real time, translate it into a different language, and output it as speech." This system enables users to communicate seamlessly between different languages.
[0250] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0251] Step 1:
[0252] When a user begins speaking in their first language using the device, the device acquires audio data through its microphone. The input is the user's voice, and the output is digital audio data. This audio data is then prepared for processing by the device's built-in speech recognition system.
[0253] Step 2:
[0254] The terminal performs signal processing on the acquired audio data to remove noise and improve sound quality. This step typically utilizes a digital signal processor (DSP). The input is raw audio data, while the output is audio data with noise removed and improved sound quality. This signal processing ensures that high-quality audio data is available for speech recognition.
[0255] Step 3:
[0256] The device uses speech recognition to convert the improved audio data into text data. The input here is the processed audio data, and the output is the corresponding text data. For speech recognition, for example, the Google Speech-to-Text API is used. This text data forms the basis for subsequent translation processing.
[0257] Step 4:
[0258] The terminal sends the converted text data to the server using a communication method. The input is text data, and the output is the arrival of that text data at the server. Network connections such as Wi-Fi or mobile data are used for communication.
[0259] Step 5:
[0260] The server inputs the received text data into a generative AI model and performs translation into a second language. The input is the transmitted text data, and the output is the translated text data. The generative AI model employs AI technologies such as OpenAI GPT, which results in highly accurate translations.
[0261] Step 6:
[0262] The server sends the translated text data back to the terminal. The input is the translated text generated by the generative AI model, and the output is the terminal receiving that data. This communication also takes place over Wi-Fi or a mobile data network.
[0263] Step 7:
[0264] The terminal processes the received translated text data using a speech synthesis system and outputs it as synthesized speech through the speaker. The input is the translated text data sent from the server, and the output is the speech heard by the user. This speech synthesis uses tools such as Amazon Polly or Google Text-to-Speech to generate natural-sounding speech.
[0265] Step 8:
[0266] Users input feedback on translation quality using the terminal interface. The input is the user's rating, and the output is the transmission of this feedback data to the server. The server uses this feedback to improve the accuracy of the generating AI model.
[0267] (Application Example 1)
[0268] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0269] Smooth two-way communication between users who speak different languages in a virtual sales environment presents challenges, such as misunderstandings and communication stagnation due to language barriers, which affect user satisfaction and sales efficiency. Furthermore, while natural-sounding real-time translated speech output requires high accuracy and speed, translation systems must also be able to generate translations flexibly, reflecting the user's intent and context.
[0270] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0271] In this invention, the server includes speech recognition means for converting speech expressions into text expressions, generative model means for converting the text expressions into different symbolic expressions, speech synthesis means for converting the converted symbolic expressions into speech data, and conversation support means for realizing two-way communication in different languages in a virtual sales environment. This enables real-time and highly accurate speech translation between users who speak different languages, eliminates language barriers in a virtual sales environment, and makes communication more efficient for users.
[0272] "Speech representation" refers to information in a form that can be perceived as sound waves, and is primarily generated by humans or machines for the purpose of linguistic communication.
[0273] "Written representation" refers to a form of expression that uses letters or symbols to represent sounds or concepts, and is used for recording and transmitting information.
[0274] "Different symbolic representation" refers to a representation using characters or symbols in a language different from the original language, and its purpose is to make information understandable to speakers of other languages.
[0275] A "generative model" is a computational algorithm used to translate or generate natural language based on input data, and is implemented using neural networks and other methods.
[0276] "Audio data" refers to audio signals that have been processed and analyzed in a digital format, and is primarily used to facilitate communication and recording.
[0277] A "computer network" is an infrastructure for the mutual exchange of information and data between computers, and the internet is one example of this.
[0278] A "translation symbolic representation" is the result of accurately converting the written representation of the original language into another language, and is generated through the translation process.
[0279] "User" refers to an individual or organization that operates a system or device, and is an entity that utilizes the system's functions to achieve a specific purpose.
[0280] "Evaluation information" refers to the feedback and reviews that users provide to the system, and this data is used for future improvements and accuracy enhancements.
[0281] The system for implementing this invention combines various means to enable communication in different languages among users. First, the terminal uses speech recognition software to receive voice input. Specifically, it utilizes the speech_recognition library to convert voice data into a character representation. This enables users to start conversations naturally.
[0282] Once the voice data is converted into a character representation, that character representation is transmitted to the server through the communication means. The server translates the received character representation into a different symbol representation using a generative AI model. Here, to improve translation accuracy, the generative model uses machine learning, and models such as GPT-3 or similar ones may be applied. The translated symbol representation is transmitted back to the terminal.
[0283] The terminal converts the received symbol representation into voice data by voice synthesis means and provides it to the user as natural voice using a voice synthesis library such as gTTS. Also, the user can evaluate the quality of the translated content via the interface means, and this evaluation is transmitted to the server and utilized to improve subsequent translation accuracy.
[0284] As a specific example, when a Japanese-speaking user visiting a virtual sales environment asks, "Please tell me the features of this product," the store clerk replies in English, "This product is eco-friendly and cost-effective." In response, the terminal translates it into Japanese and outputs the translated voice in Japanese. In this way, smooth communication is realized among users speaking different languages.
[0285] An example of a prompt sentence is, "Recognize what the user asks, translate it into English, answer it, and then translate that answer back into Japanese and play it as voice." This enables the system to operate as planned and support the communication intended by the user.
[0286] The flow of the specific process in Application Example 1 will be described using FIG. 12.
[0287] Step 1:
[0288] The terminal receives the user's voice input. It acquires voice data through a microphone and uses the speech_recognition library to convert this voice data into a character representation. The input is the user's voice, and the output is the character representation of that voice. As a specific operation, the voice signal is captured, and voice recognition is performed in the backend process.
[0289] Step 2:
[0290] After the character representation is generated, the terminal sends this data to the server through the communication means. The input is the character representation generated in the previous step, and the output is the completion of data transmission to the server. As an operation, the data is securely transferred using the network protocol.
[0291] Step 3:
[0292] Based on the received character representation, the server uses the generation AI model to translate it into symbol representations in different languages. The input is the received character representation, and the output is the translated symbol representation. As data processing, a machine learning algorithm is used, and the weights inside the model are utilized to perform high-precision language translation.
[0293] Step 4:
[0294] The server sends the translated symbol representation back to the terminal again. Here, the input is the translated symbol representation, and the output is the completion of transmission to the terminal. As a specific operation, the data is packaged according to the server's transmission protocol and transferred via the network.
[0295] Step 5:
[0296] The terminal reconstructs the received symbolic representation as audio data using speech synthesis technology and outputs the audio using gTTS or similar methods. The input is the symbolic representation received from the server, and the output is the synthesized speech. During the conversion to audio data, speech with natural intonation is generated and played back through the speaker.
[0297] Step 6:
[0298] Users provide feedback on translation quality through the terminal interface. This feedback is sent to the server and used to improve translation accuracy in the future. The input is the user's feedback data, and the output is the completion of sending the feedback to the server. Specifically, the user input on the interface is captured and then sent back to the server.
[0299] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0300] This invention is a real-time translation system that takes user emotions into account. This system is comprised of a combination of speech recognition, a generative model, speech synthesis, communication, an interface, and an emotion engine.
[0301] When a user begins speaking in their first language, the device captures their voice using its microphone. The captured audio data is converted into text by a speech recognition system, and then an emotion engine extracts additional information, including the user's emotions. This emotion information is sent to the server along with the text data.
[0302] The server uses the received text data and sentiment information to perform translation into a second language by means of a generation model. At this time, the sentiment information is provided to the generation model means by the sentiment engine, and the expression of the translation is adjusted in a form most suitable for the user's sentiment. The translated text data is transmitted to the terminal via the communication means.
[0303] The terminal converts this translated text into voice data by means of voice synthesis and outputs it from the speaker as voice with a natural expression considering sentiment. As a result, the user can convey the intended sentiment and nuance.
[0304] Furthermore, this system includes a function of receiving feedback from the user and transmitting it to the server. The feedback is utilized for the adjustment of the generation model and the sentiment engine, improving the translation and sentiment recognition accuracy.
[0305] As a specific example, when the user speaks a little excitedly "I'm very happy today!" in Japanese, the system recognizes the sentiment, translates it into English rich in sentiment as "I'm super happy today!", and outputs it with appropriate voice. As a result, the user's sentiment can be accurately conveyed to others, and communication can be made smooth.
[0306] The process flow will be described below.
[0307] Step 1:
[0308] When the user speaks in the first language, the terminal uses the microphone to capture the voice. As a result, the utterance content is acquired as digital voice data.
[0309] Step 2:
[0310] The terminal analyzes the voice data captured using voice recognition means and converts it into corresponding text data. In this process, signal processing for noise removal and sound quality improvement is also performed.
[0311] Step 3:
[0312] The device passes the converted text data to an emotion engine to analyze the user's emotions and extract emotional information. This emotional information indicates what emotions are conveyed in the utterance.
[0313] Step 4:
[0314] The device sends a dataset containing text data and sentiment information to the server via a communication method. This enables translation processing on the server side.
[0315] Step 5:
[0316] The server inputs the received text data into a generative model and performs translation into a second language. During this process, the emotional information provided by the emotion engine is taken into consideration, and the translated text is adjusted to appropriately express the user's emotions.
[0317] Step 6:
[0318] The server returns the translated text data to the terminal. The returned data includes expressions that reflect emotions.
[0319] Step 7:
[0320] The device processes the translated text data using speech synthesis technology to generate emotionally appropriate audio data. The generated audio is then delivered to the user through the speaker.
[0321] Step 8:
[0322] Users review the translated audio and provide quality feedback through their device's interface. This feedback is then sent back to the server via communication channels and used to improve the system in the future.
[0323] (Example 2)
[0324] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0325] In translation systems, it is difficult to appropriately reflect user emotions and incorporate natural emotional expressions into the translation results. Furthermore, conventional translation systems have the challenge of not being able to effectively utilize user feedback to improve translation accuracy.
[0326] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0327] In this invention, the server includes acoustic analysis means, emotion analysis means, and generative model means. This enables natural translation expressions that reflect the user's emotions.
[0328] "Acoustic analysis means" refers to technology that has the function of collecting audio data and converting it into text data.
[0329] "Emotional analysis means" refers to a technology that processes data to extract user emotional information from text data.
[0330] A "generative model means" is a technology that uses received text data and sentiment information to translate into different languages and express emotions appropriately.
[0331] "Sound generation means" refers to a technology that generates audio data based on translated text data and emotional information, and outputs it as audio that takes emotional expression into account.
[0332] "Communication control means" refers to technology that controls the processing for sending and receiving voice data and text data over a network.
[0333] "Input device means" refers to technology that provides an interface for receiving feedback from users.
[0334] This invention provides a system that enables advanced real-time translation that takes user emotions into consideration. Specific embodiments are shown below.
[0335] First, the device uses a microphone to capture audio data of the user's first language. Speech recognition software is then used as an acoustic analysis tool to convert this audio data into text. Commercially available speech recognition technologies, such as the Google Speech-to-Text API, can be used for this process.
[0336] Next, the terminal extracts user emotion information from the converted text data using emotion analysis tools. Here, natural language processing models and machine learning algorithms can be utilized as software widely used for emotion analysis. The emotion information obtained through this process is sent to the server along with the text data.
[0337] The server uses a generative model to perform advanced translation based on the received text data and sentiment information. In this embodiment, the generative AI model can be, for example, an open-source model or a commercial advanced natural language processing model. The generative model utilizes sentiment information to adjust the expression of the translation, enabling it to accurately convey the emotions intended by the user.
[0338] The translation results are sent from the server to the terminal and converted into audio data using speech synthesis software as a means of generating sound. Natural-sounding speech can be achieved by using Amazon Polly or other commercial speech synthesis services. The output is then delivered to the user through the speaker with appropriate intonation.
[0339] As a concrete example, if a user says "I'm so happy today!", the system takes the emotion into consideration and uses the prompt "Translate the following excited expression from Japanese to English while preserving the user's emotional excitement" to generate the translation "I'm super happy today!", which is then appropriately output as speech.
[0340] Furthermore, the terminal receives feedback from the user regarding the translation results and sends this feedback to the server using communication control means. This feedback is used to improve the accuracy of the generative model and sentiment analysis. In this way, the system can continuously improve the quality of translations and the accuracy of sentiment.
[0341] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0342] Step 1:
[0343] The device captures the user's voice using a microphone. The user's voice signal is used as input. This voice signal is acquired as analog data and converted to a digital format. The output is digitized voice data. This data is used for acoustic analysis in the next step.
[0344] Step 2:
[0345] The terminal inputs the captured audio data into speech recognition software, which converts it into text data. This uses acoustic analysis, and the specific data processing involves breaking down the audio signal into phonemes and converting them into strings. The output is text data representing what the user said as a string.
[0346] Step 3:
[0347] The terminal passes text data to an emotion analysis device to extract the user's emotions. The input in this step is text data obtained from a speech recognition device. Specifically, a natural language processing algorithm is used to analyze words and contexts that represent specific emotions within the text, and the emotion information is output as data. The output is metadata that includes emotional elements such as joy, anger, sadness, and happiness.
[0348] Step 4:
[0349] The terminal transmits text data and sentiment information to the server via a communication control means. The input here is text data and its sentiment metadata. A secure communication protocol is used to ensure data integrity when transmitting to the server. As output, the data received on the server side is provided to the generation model means.
[0350] Step 5:
[0351] The server inputs the received text data and sentiment information into a generative AI model to perform translation. In this step, the translation model performs data calculations to adjust the translation expression to reflect sentiment while performing language conversion based on the input data. Specifically, it uses prompt sentences to activate the generative AI model and outputs the optimal translation result. The output is a sentiment-sensitive translated text in a second language.
[0352] Step 6:
[0353] The server sends the translated text data to the terminal. The input is the translated text generated by the generative AI model. Communication control means are used for data transfer. The output is the translated text received on the terminal side.
[0354] Step 7:
[0355] The terminal inputs translated text data into a speech synthesis system and converts it into speech data. The input is the translated text data. Specifically, the speech synthesis engine converts the text data into a speech waveform and synthesizes it as a natural-sounding voice with emotion. The output is the speech data played back through the speaker.
[0356] Step 8:
[0357] The terminal receives feedback from the user and sends it to the server. An interface is used for inputting the user's opinion and emotional response to the translation results. The feedback data sent to the server contributes to improving the generative AI model and sentiment analysis methods. The output is improvement information that will be useful for future translation accuracy improvements.
[0358] (Application Example 2)
[0359] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0360] Conventional translation systems can perform simple conversions between languages, but they have the problem of failing to adequately reflect the speaker's emotions and nuances, making natural communication difficult. In particular, in the field of content distribution, there is a demand for translations that allow viewers to accurately understand and enjoy the emotions of movies and dramas, but current technology is unable to adequately meet this requirement.
[0361] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0362] In this invention, the server includes speech recognition means for converting speech data into text data, generative model means for translating text data into different languages based on emotional information, and speech synthesis means for converting translated text data that reflects emotional information into speech data. This enables natural translation that takes the user's emotions into account.
[0363] "Speech recognition means" refers to a technology that converts speech data into text data, making it possible to process the user's speech content as text information.
[0364] A "generative model means" is a technology for translating text into different languages using text data and sentiment information obtained by speech recognition means.
[0365] "Speech synthesis means" refers to a technology that converts translated text data into speech data and provides the user with natural-sounding speech output.
[0366] "Communication means" refers to the technology for sending and receiving text data obtained by speech recognition means, emotion information, and translated text data obtained by generative model means via a communication network.
[0367] An "interface means" is a means for users to input feedback regarding the quality and emotional expression of translations, and enables interaction with the user.
[0368] "Emotional information" refers to information extracted by analyzing the emotions and nuances contained in the user's utterance, and is an important element in the translation process.
[0369] The system for implementing this invention consists of a user terminal and a server. The terminal is equipped with speech recognition means, communication means, and speech synthesis means, which convert the user's voice into text data in real time and extract emotional information. The voice data captured on the terminal is then converted into text data through speech recognition software. At this time, an emotional recognition algorithm is used to analyze the user's emotions and generate corresponding emotional information.
[0370] Next, this text data and sentiment information are sent to the server via a communication method. The server uses a generative AI model to translate the text data into different languages. At this time, the sentiment information is input into the generative model, so the translated text takes sentiment into account. The translated text data is then sent back to the terminal.
[0371] The device uses the received translated text data to create natural-sounding, emotionally resonant voice data through speech synthesis. The user can then listen to the translated content through the audio output from the speaker.
[0372] As a concrete example, when a user is watching a movie, the emotionally rich lines spoken by the characters can be instantly translated into different languages, taking emotional information into account, and delivered to the viewer. Through this process, users can accurately understand the emotions and deepen their overall impression of the movie.
[0373] Examples of prompt messages are as follows:
[0374] "When Elsa begins to sing in the snowy mountains, feeling alone but yearning for freedom, please generate a translation that reflects her emotions and strength."
[0375] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0376] Step 1:
[0377] The terminal receives the user's voice as input and captures the audio data using the microphone. The captured audio data is sent to a speech recognition system. The speech recognition system processes the audio data to convert it into text data and passes this converted text data to the next step.
[0378] Step 2:
[0379] The device uses text data obtained by speech recognition to run an emotion recognition algorithm. This algorithm receives text data as input, analyzes the user's emotions from it, and extracts emotion information. The emotion information is output as data representing the emotions expressed in the user's speech.
[0380] Step 3:
[0381] Text data and sentiment information are transmitted from the terminal to the server using a communication method. The server then uses the acquired data as input and runs a generative AI model. The generative AI model translates the text into different languages based on the input data. In this process, sentiment information is reflected in the translation result, and the output is translated text data with adjusted expression.
[0382] Step 4:
[0383] The server sends the translated text data to the terminal. The terminal passes this received data as input to the speech synthesis system. The speech synthesis system converts the translated text data into speech data and generates speech that has been adjusted to convey emotions naturally.
[0384] Step 5:
[0385] The device outputs audio data synthesized by a speech synthesis system to the user through its speaker. This allows the user to hear the translated content in audio format.
[0386] Step 6:
[0387] Users input feedback on the provided translation into their device. This feedback is collected through the interface as opinions on the translation quality and sentiment expression, and is sent back to the server. The server uses this feedback information to improve the accuracy of its generative AI model and sentiment recognition.
[0388] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0389] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0390] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0391] [Third Embodiment]
[0392] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0393] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0394] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0395] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0396] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0397] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0398] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0399] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0400] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0401] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0402] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0403] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0404] This invention provides a voice and text translation system that enables smooth, real-time communication between users who speak different languages. Specific embodiments thereof are described below.
[0405] When a user begins speaking in their first language using the device, the device acquires audio data through its microphone. This audio data is analyzed by the device's speech recognition system, processed for noise reduction and sound quality improvement, and then converted into text data. The converted text data is then transmitted to a server via a communication device.
[0406] The server inputs the received text data into a generative model and performs translation into a second language. This generative model continuously updates its model parameters using user feedback, improving translation accuracy. The translated text data is then sent back to the terminal via communication.
[0407] The terminal processes the received translated text using speech synthesis and outputs it as synthesized speech through the speaker. This allows the user to hear the translated content in a natural-sounding voice. Furthermore, the user can input feedback on the translation quality through the terminal's interface. The input feedback is sent back to the server and used to improve translation accuracy in the future.
[0408] In this way, the system provides users with a real-time, highly accurate translation function, facilitating smooth communication. A concrete example of this embodiment is when a Japanese-speaking user says, "It's a nice day today," the speech recognition means converts this into text, the generative model means translates it into English as "It's a nice day today," and the speech synthesis means outputs it as English speech. Through this process, users can achieve seamless communication using different languages.
[0409] The following describes the processing flow.
[0410] Step 1:
[0411] When the user speaks in their first language, the device captures the audio data through the microphone. The captured audio data is immediately sent to the speech recognition system.
[0412] Step 2:
[0413] The terminal's voice recognition means analyzes the voice data and performs noise reduction and sound quality improvement. Subsequently, based on the analysis results, the voice is converted into text data. This text data is sent to the server via the communication means for processing by the generative model means.
[0414] Step 3:
[0415] The server inputs text data received via communication into a generation model. The generation model translates this text into a second language and generates highly accurate translated text data. The generated translated data is then transmitted back to the terminal via communication.
[0416] Step 4:
[0417] The terminal receives the translated text data from the server and inputs it into a speech synthesis device. The speech synthesis device converts the text data into speech data and generates synthesized speech. This speech data is output through the speaker, allowing the user to hear the translated content.
[0418] Step 5:
[0419] Users input feedback on translation quality using an interface on their device. This feedback information is sent to the server via a communication method.
[0420] Step 6:
[0421] The server incorporates the received feedback information into the generation model and updates the model parameters, thereby improving accuracy in subsequent translation processes.
[0422] (Example 1)
[0423] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0424] In communication between users who speak different languages, there is a need for real-time, highly accurate translation. However, conventional technologies have faced challenges in smooth communication due to insufficient speech recognition and translation accuracy, or unnaturalness in the resulting speech synthesis. Furthermore, there have been insufficient methods for efficiently utilizing user feedback to improve translation accuracy.
[0425] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0426] In this invention, the server includes speech recognition means for collecting audio data, performing noise reduction and sound quality improvement, and converting the results into text data; generative model means for translating the text data into different languages, and means for updating model parameters using user feedback; and speech synthesis means for converting the translated text data into audio data and outputting it as natural speech. This enables users to engage in real-time, effective, and natural two-way communication.
[0427] "Speech recognition means" refers to means that have the function of collecting speech data, removing noise and improving sound quality, and then converting that speech data into text data.
[0428] A "generative model means" is a means of translating input text data into different languages, and has the function of updating model parameters using user feedback to improve translation accuracy.
[0429] "Speech synthesis means" refers to a method that uses technology to convert translated text data into speech data and output it as natural-sounding speech.
[0430] "Communication means" refers to means for sending and receiving data obtained by speech recognition means and generative model means via a digital communication network.
[0431] "Interface means" refers to means that have input means for users to provide feedback on the quality of translations, and means for appropriately processing that information.
[0432] This invention provides a voice and text translation system for facilitating smooth, real-time communication between users who speak different languages. When a user uses a terminal and begins speaking in their first language, the terminal acquires voice data through its built-in microphone. The acquired voice data is analyzed using speech recognition means within the terminal, and after noise reduction and sound quality improvement, it is converted into text data. Speech recognition software such as the Google Speech-to-Text API can be used in this process. The converted text data is transmitted to a server via the terminal's communication means.
[0433] The server inputs the received text data into a generative model and performs translation into a second language. AI models such as OpenAI GPT or DeepL API can be used for this process. The generative model has a mechanism to continuously improve translation accuracy by updating its model parameters based on user feedback. The translated text data is sent back from the server to the terminal, where it is processed by a speech synthesis system and output as synthesized speech through the speaker. Technologies such as Google Text-to-Speech or Amazon Polly can be used for speech synthesis. This output allows the user to hear the translated content in a natural-sounding voice.
[0434] Furthermore, users can use an interface to input feedback on translation quality and send it to the server via their device. This feedback will be used to improve translation accuracy in the future.
[0435] As a concrete example, if a Japanese-speaking user says "It's a nice day today," the device can convert this into text, translate it into English as "It's a nice day today" using a generative model, and output it as English speech using a speech synthesis tool. An example of a prompt message could be, "Please convert the audio data into text in real time, translate it into a different language, and output it as speech." This system enables users to communicate seamlessly between different languages.
[0436] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0437] Step 1:
[0438] When a user begins speaking in their first language using the device, the device acquires audio data through its microphone. The input is the user's voice, and the output is digital audio data. This audio data is then prepared for processing by the device's built-in speech recognition system.
[0439] Step 2:
[0440] The terminal performs signal processing on the acquired audio data to remove noise and improve sound quality. This step typically utilizes a digital signal processor (DSP). The input is raw audio data, while the output is audio data with noise removed and improved sound quality. This signal processing ensures that high-quality audio data is available for speech recognition.
[0441] Step 3:
[0442] The device uses speech recognition to convert the improved audio data into text data. The input here is the processed audio data, and the output is the corresponding text data. For speech recognition, for example, the Google Speech-to-Text API is used. This text data forms the basis for subsequent translation processing.
[0443] Step 4:
[0444] The terminal sends the converted text data to the server using a communication method. The input is text data, and the output is the arrival of that text data at the server. Network connections such as Wi-Fi or mobile data are used for communication.
[0445] Step 5:
[0446] The server inputs the received text data into a generative AI model and performs translation into a second language. The input is the transmitted text data, and the output is the translated text data. The generative AI model employs AI technologies such as OpenAI GPT, which results in highly accurate translations.
[0447] Step 6:
[0448] The server sends the translated text data back to the terminal. The input is the translated text generated by the generative AI model, and the output is the terminal receiving that data. This communication also takes place over Wi-Fi or a mobile data network.
[0449] Step 7:
[0450] The terminal processes the received translated text data using a speech synthesis system and outputs it as synthesized speech through the speaker. The input is the translated text data sent from the server, and the output is the speech heard by the user. This speech synthesis uses tools such as Amazon Polly or Google Text-to-Speech to generate natural-sounding speech.
[0451] Step 8:
[0452] Users input feedback on translation quality using the terminal interface. The input is the user's rating, and the output is the transmission of this feedback data to the server. The server uses this feedback to improve the accuracy of the generating AI model.
[0453] (Application Example 1)
[0454] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0455] Smooth two-way communication between users who speak different languages in a virtual sales environment presents challenges, such as misunderstandings and communication stagnation due to language barriers, which affect user satisfaction and sales efficiency. Furthermore, while natural-sounding real-time translated speech output requires high accuracy and speed, translation systems must also be able to generate translations flexibly, reflecting the user's intent and context.
[0456] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0457] In this invention, the server includes speech recognition means for converting speech expressions into text expressions, generative model means for converting the text expressions into different symbolic expressions, speech synthesis means for converting the converted symbolic expressions into speech data, and conversation support means for realizing two-way communication in different languages in a virtual sales environment. This enables real-time and highly accurate speech translation between users who speak different languages, eliminates language barriers in a virtual sales environment, and makes communication more efficient for users.
[0458] "Speech representation" refers to information in a form that can be perceived as sound waves, and is primarily generated by humans or machines for the purpose of linguistic communication.
[0459] "Written representation" refers to a form of expression that uses letters or symbols to represent sounds or concepts, and is used for recording and transmitting information.
[0460] "Different symbolic representation" refers to a representation using characters or symbols in a language different from the original language, and its purpose is to make information understandable to speakers of other languages.
[0461] A "generative model" is a computational algorithm used to translate or generate natural language based on input data, and is implemented using neural networks and other methods.
[0462] "Audio data" refers to audio signals that have been processed and analyzed in a digital format, and is primarily used to facilitate communication and recording.
[0463] A "computer network" is an infrastructure for the mutual exchange of information and data between computers, and the internet is one example of this.
[0464] A "translation symbolic representation" is the result of accurately converting the written representation of the original language into another language, and is generated through the translation process.
[0465] "User" refers to an individual or organization that operates a system or device, and is an entity that utilizes the system's functions to achieve a specific purpose.
[0466] "Evaluation information" refers to the feedback and reviews that users provide to the system, and this data is used for future improvements and accuracy enhancements.
[0467] The system for implementing this invention combines various means to enable communication between users in different languages. First, the terminal uses speech recognition software to accept voice input. Specifically, it utilizes the speech_recognition library to convert voice data into text. This allows users to naturally begin a conversation.
[0468] Once the audio data is converted into text, that text is sent to a server via a communication method. The server uses a generative AI model to translate the received text into a different symbolic representation. Here, machine learning is used in the generative model to improve translation accuracy, and GPT-3 or similar models may be applied. The translated symbolic representation is then sent back to the terminal.
[0469] The terminal converts the received symbolic representations into speech data using speech synthesis means and provides it to the user in natural-sounding speech using a speech synthesis library such as gTTS. Furthermore, the user can evaluate the quality of the translated content through an interface means, and this evaluation is sent to the server and used to improve translation accuracy in the future.
[0470] For example, if a Japanese-speaking user visiting a virtual sales environment asks, "Could you tell me about the features of this product?", the salesperson will respond in English, "This product is eco-friendly and cost-effective." The terminal then translates this into Japanese and outputs the translated audio in Japanese. In this way, smooth communication is achieved between users who speak different languages.
[0471] An example of a prompt message would be, "Recognize the user's question, translate it into English, answer the question, translate that answer back into Japanese, and play it aloud." This allows the system to operate as planned and support the user's intended communication.
[0472] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0473] Step 1:
[0474] The device receives voice input from the user. It acquires voice data through the microphone and converts this voice data into text using the speech_recognition library. The input is the user's voice, and the output is the textual representation of that voice. Specifically, the voice signal is captured, and speech recognition is performed in the backend process.
[0475] Step 2:
[0476] After the character representation is generated, the terminal sends this data to the server via a communication method. The input is the character representation generated in the previous step, and the output is the completion of the data transmission to the server. In this operation, data is securely transferred using a network protocol.
[0477] Step 3:
[0478] The server uses a generative AI model to translate the received text representation into a symbolic representation in a different language. The input is the received text representation, and the output is the translated symbolic representation. For data processing, machine learning algorithms are used, and the weights within the model are utilized to perform highly accurate language translation.
[0479] Step 4:
[0480] The server sends the translated symbolic representation back to the terminal. Here, the input is the translated symbolic representation, and the output is the completion of the transmission to the terminal. Specifically, the data is packaged according to the server's transmission protocol and transferred over the network.
[0481] Step 5:
[0482] The terminal reconstructs the received symbolic representation as audio data using speech synthesis technology and outputs the audio using gTTS or similar methods. The input is the symbolic representation received from the server, and the output is the synthesized speech. During the conversion to audio data, speech with natural intonation is generated and played back through the speaker.
[0483] Step 6:
[0484] Users provide feedback on translation quality through the terminal interface. This feedback is sent to the server and used to improve translation accuracy in the future. The input is the user's feedback data, and the output is the completion of sending the feedback to the server. Specifically, the user input on the interface is captured and then sent back to the server.
[0485] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0486] This invention is a real-time translation system that takes user emotions into account. This system is comprised of a combination of speech recognition, a generative model, speech synthesis, communication, an interface, and an emotion engine.
[0487] When a user begins speaking in their first language, the device captures their voice using its microphone. The captured audio data is converted into text by a speech recognition system, and then an emotion engine extracts additional information, including the user's emotions. This emotion information is sent to the server along with the text data.
[0488] The server uses the received text data and sentiment information to perform translation into a second language using a generative model. During this process, the sentiment engine provides sentiment information to the generative model, and the translated expression is adjusted to best reflect the user's emotions. The translated text data is then transmitted to the terminal via a communication device.
[0489] The device converts this translated text into speech data using speech synthesis technology and outputs it through the speaker as natural-sounding speech that takes emotions into consideration. This allows the user to convey their intended emotions and nuances.
[0490] Furthermore, the system includes a function to receive user feedback and send it to the server. This feedback is used to refine the generative model and sentiment engine, improving the accuracy of translation and sentiment recognition.
[0491] For example, if a user gets a little excited and says "I'm so happy today!" in Japanese, the system recognizes that emotion and translates it into English as "I'm super happy today!" with emotion, outputting it in appropriate voice. This ensures that the user's emotions are accurately conveyed to others, facilitating smoother communication.
[0492] The following describes the processing flow.
[0493] Step 1:
[0494] When a user speaks in their first language, the device uses its microphone to capture the audio. This allows the spoken content to be obtained as digital audio data.
[0495] Step 2:
[0496] The device analyzes the captured audio data using speech recognition and converts it into corresponding text data. During this process, signal processing is also performed to remove noise and improve sound quality.
[0497] Step 3:
[0498] The device passes the converted text data to an emotion engine to analyze the user's emotions and extract emotional information. This emotional information indicates what emotions are conveyed in the utterance.
[0499] Step 4:
[0500] The device sends a dataset containing text data and sentiment information to the server via a communication method. This enables translation processing on the server side.
[0501] Step 5:
[0502] The server inputs the received text data into a generative model and performs translation into a second language. During this process, the emotional information provided by the emotion engine is taken into consideration, and the translated text is adjusted to appropriately express the user's emotions.
[0503] Step 6:
[0504] The server returns the translated text data to the terminal. The returned data includes expressions that reflect emotions.
[0505] Step 7:
[0506] The device processes the translated text data using speech synthesis technology to generate emotionally appropriate audio data. The generated audio is then delivered to the user through the speaker.
[0507] Step 8:
[0508] Users review the translated audio and provide quality feedback through their device's interface. This feedback is then sent back to the server via communication channels and used to improve the system in the future.
[0509] (Example 2)
[0510] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0511] In translation systems, it is difficult to appropriately reflect user emotions and incorporate natural emotional expressions into the translation results. Furthermore, conventional translation systems have the challenge of not being able to effectively utilize user feedback to improve translation accuracy.
[0512] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0513] In this invention, the server includes acoustic analysis means, emotion analysis means, and generative model means. This enables natural translation expressions that reflect the user's emotions.
[0514] "Acoustic analysis means" refers to technology that has the function of collecting audio data and converting it into text data.
[0515] "Emotional analysis means" refers to a technology that processes data to extract user emotional information from text data.
[0516] A "generative model means" is a technology that uses received text data and sentiment information to translate into different languages and express emotions appropriately.
[0517] "Sound generation means" refers to a technology that generates audio data based on translated text data and emotional information, and outputs it as audio that takes emotional expression into account.
[0518] "Communication control means" refers to technology that controls the processing for sending and receiving voice data and text data over a network.
[0519] "Input device means" refers to technology that provides an interface for receiving feedback from users.
[0520] This invention provides a system that enables advanced real-time translation that takes user emotions into consideration. Specific embodiments are shown below.
[0521] First, the device uses a microphone to capture audio data of the user's first language. Speech recognition software is then used as an acoustic analysis tool to convert this audio data into text. Commercially available speech recognition technologies, such as the Google Speech-to-Text API, can be used for this process.
[0522] Next, the terminal extracts user emotion information from the converted text data using emotion analysis tools. Here, natural language processing models and machine learning algorithms can be utilized as software widely used for emotion analysis. The emotion information obtained through this process is sent to the server along with the text data.
[0523] The server uses a generative model to perform advanced translation based on the received text data and sentiment information. In this embodiment, the generative AI model can be, for example, an open-source model or a commercial advanced natural language processing model. The generative model utilizes sentiment information to adjust the expression of the translation, enabling it to accurately convey the emotions intended by the user.
[0524] The translation results are sent from the server to the terminal and converted into audio data using speech synthesis software as a means of generating sound. Natural-sounding speech can be achieved by using Amazon Polly or other commercial speech synthesis services. The output is then delivered to the user through the speaker with appropriate intonation.
[0525] As a concrete example, if a user says "I'm so happy today!", the system takes the emotion into consideration and uses the prompt "Translate the following excited expression from Japanese to English while preserving the user's emotional excitement" to generate the translation "I'm super happy today!", which is then appropriately output as speech.
[0526] Furthermore, the terminal receives feedback from the user regarding the translation results and sends this feedback to the server using communication control means. This feedback is used to improve the accuracy of the generative model and sentiment analysis. In this way, the system can continuously improve the quality of translations and the accuracy of sentiment.
[0527] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0528] Step 1:
[0529] The device captures the user's voice using a microphone. The user's voice signal is used as input. This voice signal is acquired as analog data and converted to a digital format. The output is digitized voice data. This data is used for acoustic analysis in the next step.
[0530] Step 2:
[0531] The terminal inputs the captured audio data into speech recognition software, which converts it into text data. This uses acoustic analysis, and the specific data processing involves breaking down the audio signal into phonemes and converting them into strings. The output is text data representing what the user said as a string.
[0532] Step 3:
[0533] The terminal passes text data to an emotion analysis device to extract the user's emotions. The input in this step is text data obtained from a speech recognition device. Specifically, a natural language processing algorithm is used to analyze words and contexts that represent specific emotions within the text, and the emotion information is output as data. The output is metadata that includes emotional elements such as joy, anger, sadness, and happiness.
[0534] Step 4:
[0535] The terminal transmits text data and sentiment information to the server via a communication control means. The input here is text data and its sentiment metadata. A secure communication protocol is used to ensure data integrity when transmitting to the server. As output, the data received on the server side is provided to the generation model means.
[0536] Step 5:
[0537] The server inputs the received text data and sentiment information into a generative AI model to perform translation. In this step, the translation model performs data calculations to adjust the translation expression to reflect sentiment while performing language conversion based on the input data. Specifically, it uses prompt sentences to activate the generative AI model and outputs the optimal translation result. The output is a sentiment-sensitive translated text in a second language.
[0538] Step 6:
[0539] The server sends the translated text data to the terminal. The input is the translated text generated by the generative AI model. Communication control means are used for data transfer. The output is the translated text received on the terminal side.
[0540] Step 7:
[0541] The terminal inputs translated text data into a speech synthesis system and converts it into speech data. The input is the translated text data. Specifically, the speech synthesis engine converts the text data into a speech waveform and synthesizes it as a natural-sounding voice with emotion. The output is the speech data played back through the speaker.
[0542] Step 8:
[0543] The terminal receives feedback from the user and sends it to the server. An interface is used for inputting the user's opinion and emotional response to the translation results. The feedback data sent to the server contributes to improving the generative AI model and sentiment analysis methods. The output is improvement information that will be useful for future translation accuracy improvements.
[0544] (Application Example 2)
[0545] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0546] Conventional translation systems can perform simple conversions between languages, but they have the problem of failing to adequately reflect the speaker's emotions and nuances, making natural communication difficult. In particular, in the field of content distribution, there is a demand for translations that allow viewers to accurately understand and enjoy the emotions of movies and dramas, but current technology is unable to adequately meet this requirement.
[0547] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0548] In this invention, the server includes speech recognition means for converting speech data into text data, generative model means for translating text data into different languages based on emotional information, and speech synthesis means for converting translated text data that reflects emotional information into speech data. This enables natural translation that takes the user's emotions into account.
[0549] "Speech recognition means" refers to a technology that converts speech data into text data, making it possible to process the user's speech content as text information.
[0550] A "generative model means" is a technology for translating text into different languages using text data and sentiment information obtained by speech recognition means.
[0551] "Speech synthesis means" refers to a technology that converts translated text data into speech data and provides the user with natural-sounding speech output.
[0552] "Communication means" refers to the technology for sending and receiving text data obtained by speech recognition means, emotion information, and translated text data obtained by generative model means via a communication network.
[0553] An "interface means" is a means for users to input feedback regarding the quality and emotional expression of translations, and enables interaction with the user.
[0554] "Emotional information" refers to information extracted by analyzing the emotions and nuances contained in the user's utterance, and is an important element in the translation process.
[0555] The system for implementing this invention consists of a user terminal and a server. The terminal is equipped with speech recognition means, communication means, and speech synthesis means, which convert the user's voice into text data in real time and extract emotional information. The voice data captured on the terminal is then converted into text data through speech recognition software. At this time, an emotional recognition algorithm is used to analyze the user's emotions and generate corresponding emotional information.
[0556] Next, this text data and sentiment information are sent to the server via a communication method. The server uses a generative AI model to translate the text data into different languages. At this time, the sentiment information is input into the generative model, so the translated text takes sentiment into account. The translated text data is then sent back to the terminal.
[0557] The device uses the received translated text data to create natural-sounding, emotionally resonant voice data through speech synthesis. The user can then listen to the translated content through the audio output from the speaker.
[0558] As a concrete example, when a user is watching a movie, the emotionally rich lines spoken by the characters can be instantly translated into different languages, taking emotional information into account, and delivered to the viewer. Through this process, users can accurately understand the emotions and deepen their overall impression of the movie.
[0559] Examples of prompt messages are as follows:
[0560] "When Elsa begins to sing in the snowy mountains, feeling alone but yearning for freedom, please generate a translation that reflects her emotions and strength."
[0561] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0562] Step 1:
[0563] The terminal receives the user's voice as input and captures the audio data using the microphone. The captured audio data is sent to a speech recognition system. The speech recognition system processes the audio data to convert it into text data and passes this converted text data to the next step.
[0564] Step 2:
[0565] The device uses text data obtained by speech recognition to run an emotion recognition algorithm. This algorithm receives text data as input, analyzes the user's emotions from it, and extracts emotion information. The emotion information is output as data representing the emotions expressed in the user's speech.
[0566] Step 3:
[0567] Text data and sentiment information are transmitted from the terminal to the server using a communication method. The server then uses the acquired data as input and runs a generative AI model. The generative AI model translates the text into different languages based on the input data. In this process, sentiment information is reflected in the translation result, and the output is translated text data with adjusted expression.
[0568] Step 4:
[0569] The server sends the translated text data to the terminal. The terminal passes this received data as input to the speech synthesis system. The speech synthesis system converts the translated text data into speech data and generates speech that has been adjusted to convey emotions naturally.
[0570] Step 5:
[0571] The device outputs audio data synthesized by a speech synthesis system to the user through its speaker. This allows the user to hear the translated content in audio format.
[0572] Step 6:
[0573] Users input feedback on the provided translation into their device. This feedback is collected through the interface as opinions on the translation quality and sentiment expression, and is sent back to the server. The server uses this feedback information to improve the accuracy of its generative AI model and sentiment recognition.
[0574] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0575] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0576] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[0577] [Fourth Embodiment]
[0578] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0579] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0580] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0581] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0582] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0583] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0584] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0585] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0586] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0587] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0588] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0589] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0590] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0591] This invention provides a voice and text translation system that enables smooth, real-time communication between users who speak different languages. Specific embodiments thereof are described below.
[0592] When a user begins speaking in their first language using the device, the device acquires audio data through its microphone. This audio data is analyzed by the device's speech recognition system, processed for noise reduction and sound quality improvement, and then converted into text data. The converted text data is then transmitted to a server via a communication device.
[0593] The server inputs the received text data into a generative model and performs translation into a second language. This generative model continuously updates its model parameters using user feedback, improving translation accuracy. The translated text data is then sent back to the terminal via communication.
[0594] The terminal processes the received translated text using speech synthesis and outputs it as synthesized speech through the speaker. This allows the user to hear the translated content in a natural-sounding voice. Furthermore, the user can input feedback on the translation quality through the terminal's interface. The input feedback is sent back to the server and used to improve translation accuracy in the future.
[0595] In this way, the system provides users with a real-time, highly accurate translation function, facilitating smooth communication. A concrete example of this embodiment is when a Japanese-speaking user says, "It's a nice day today," the speech recognition means converts this into text, the generative model means translates it into English as "It's a nice day today," and the speech synthesis means outputs it as English speech. Through this process, users can achieve seamless communication using different languages.
[0596] The following describes the processing flow.
[0597] Step 1:
[0598] When the user speaks in their first language, the device captures the audio data through the microphone. The captured audio data is immediately sent to the speech recognition system.
[0599] Step 2:
[0600] The terminal's voice recognition means analyzes the voice data and performs noise reduction and sound quality improvement. Subsequently, based on the analysis results, the voice is converted into text data. This text data is sent to the server via the communication means for processing by the generative model means.
[0601] Step 3:
[0602] The server inputs text data received via communication into a generation model. The generation model translates this text into a second language and generates highly accurate translated text data. The generated translated data is then transmitted back to the terminal via communication.
[0603] Step 4:
[0604] The terminal receives the translated text data from the server and inputs it into a speech synthesis device. The speech synthesis device converts the text data into speech data and generates synthesized speech. This speech data is output through the speaker, allowing the user to hear the translated content.
[0605] Step 5:
[0606] Users input feedback on translation quality using an interface on their device. This feedback information is sent to the server via a communication method.
[0607] Step 6:
[0608] The server incorporates the received feedback information into the generation model and updates the model parameters, thereby improving accuracy in subsequent translation processes.
[0609] (Example 1)
[0610] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0611] In communication between users who speak different languages, there is a need for real-time, highly accurate translation. However, conventional technologies have faced challenges in smooth communication due to insufficient speech recognition and translation accuracy, or unnaturalness in the resulting speech synthesis. Furthermore, there have been insufficient methods for efficiently utilizing user feedback to improve translation accuracy.
[0612] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0613] In this invention, the server includes speech recognition means for collecting audio data, performing noise reduction and sound quality improvement, and converting the results into text data; generative model means for translating the text data into different languages, and means for updating model parameters using user feedback; and speech synthesis means for converting the translated text data into audio data and outputting it as natural speech. This enables users to engage in real-time, effective, and natural two-way communication.
[0614] "Speech recognition means" refers to means that have the function of collecting speech data, removing noise and improving sound quality, and then converting that speech data into text data.
[0615] A "generative model means" is a means of translating input text data into different languages, and has the function of updating model parameters using user feedback to improve translation accuracy.
[0616] "Speech synthesis means" refers to a method that uses technology to convert translated text data into speech data and output it as natural-sounding speech.
[0617] "Communication means" refers to means for sending and receiving data obtained by speech recognition means and generative model means via a digital communication network.
[0618] "Interface means" refers to means that have input means for users to provide feedback on the quality of translations, and means for appropriately processing that information.
[0619] This invention provides a voice and text translation system for facilitating smooth, real-time communication between users who speak different languages. When a user uses a terminal and begins speaking in their first language, the terminal acquires voice data through its built-in microphone. The acquired voice data is analyzed using speech recognition means within the terminal, and after noise reduction and sound quality improvement, it is converted into text data. Speech recognition software such as the Google Speech-to-Text API can be used in this process. The converted text data is transmitted to a server via the terminal's communication means.
[0620] The server inputs the received text data into a generative model and performs translation into a second language. AI models such as OpenAI GPT or DeepL API can be used for this process. The generative model has a mechanism to continuously improve translation accuracy by updating its model parameters based on user feedback. The translated text data is sent back from the server to the terminal, where it is processed by a speech synthesis system and output as synthesized speech through the speaker. Technologies such as Google Text-to-Speech or Amazon Polly can be used for speech synthesis. This output allows the user to hear the translated content in a natural-sounding voice.
[0621] Furthermore, users can use an interface to input feedback on translation quality and send it to the server via their device. This feedback will be used to improve translation accuracy in the future.
[0622] As a concrete example, if a Japanese-speaking user says "It's a nice day today," the device can convert this into text, translate it into English as "It's a nice day today" using a generative model, and output it as English speech using a speech synthesis tool. An example of a prompt message could be, "Please convert the audio data into text in real time, translate it into a different language, and output it as speech." This system enables users to communicate seamlessly between different languages.
[0623] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0624] Step 1:
[0625] When a user begins speaking in their first language using the device, the device acquires audio data through its microphone. The input is the user's voice, and the output is digital audio data. This audio data is then prepared for processing by the device's built-in speech recognition system.
[0626] Step 2:
[0627] The terminal performs signal processing on the acquired audio data to remove noise and improve sound quality. This step typically utilizes a digital signal processor (DSP). The input is raw audio data, while the output is audio data with noise removed and improved sound quality. This signal processing ensures that high-quality audio data is available for speech recognition.
[0628] Step 3:
[0629] The device uses speech recognition to convert the improved audio data into text data. The input here is the processed audio data, and the output is the corresponding text data. For speech recognition, for example, the Google Speech-to-Text API is used. This text data forms the basis for subsequent translation processing.
[0630] Step 4:
[0631] The terminal sends the converted text data to the server using a communication method. The input is text data, and the output is the arrival of that text data at the server. Network connections such as Wi-Fi or mobile data are used for communication.
[0632] Step 5:
[0633] The server inputs the received text data into a generative AI model and performs translation into a second language. The input is the transmitted text data, and the output is the translated text data. The generative AI model employs AI technologies such as OpenAI GPT, which results in highly accurate translations.
[0634] Step 6:
[0635] The server sends the translated text data back to the terminal. The input is the translated text generated by the generative AI model, and the output is the terminal receiving that data. This communication also takes place over Wi-Fi or a mobile data network.
[0636] Step 7:
[0637] The terminal processes the received translated text data using a speech synthesis system and outputs it as synthesized speech through the speaker. The input is the translated text data sent from the server, and the output is the speech heard by the user. This speech synthesis uses tools such as Amazon Polly or Google Text-to-Speech to generate natural-sounding speech.
[0638] Step 8:
[0639] Users input feedback on translation quality using the terminal interface. The input is the user's rating, and the output is the transmission of this feedback data to the server. The server uses this feedback to improve the accuracy of the generating AI model.
[0640] (Application Example 1)
[0641] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0642] Smooth two-way communication between users who speak different languages in a virtual sales environment presents challenges, such as misunderstandings and communication stagnation due to language barriers, which affect user satisfaction and sales efficiency. Furthermore, while natural-sounding real-time translated speech output requires high accuracy and speed, translation systems must also be able to generate translations flexibly, reflecting the user's intent and context.
[0643] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0644] In this invention, the server includes speech recognition means for converting speech expressions into text expressions, generative model means for converting the text expressions into different symbolic expressions, speech synthesis means for converting the converted symbolic expressions into speech data, and conversation support means for realizing two-way communication in different languages in a virtual sales environment. This enables real-time and highly accurate speech translation between users who speak different languages, eliminates language barriers in a virtual sales environment, and makes communication more efficient for users.
[0645] "Speech representation" refers to information in a form that can be perceived as sound waves, and is primarily generated by humans or machines for the purpose of linguistic communication.
[0646] "Written representation" refers to a form of expression that uses letters or symbols to represent sounds or concepts, and is used for recording and transmitting information.
[0647] "Different symbolic representation" refers to a representation using characters or symbols in a language different from the original language, and its purpose is to make information understandable to speakers of other languages.
[0648] A "generative model" is a computational algorithm used to translate or generate natural language based on input data, and is implemented using neural networks and other methods.
[0649] "Audio data" refers to audio signals that have been processed and analyzed in a digital format, and is primarily used to facilitate communication and recording.
[0650] A "computer network" is an infrastructure for the mutual exchange of information and data between computers, and the internet is one example of this.
[0651] A "translation symbolic representation" is the result of accurately converting the written representation of the original language into another language, and is generated through the translation process.
[0652] "User" refers to an individual or organization that operates a system or device, and is an entity that utilizes the system's functions to achieve a specific purpose.
[0653] "Evaluation information" refers to the feedback and reviews that users provide to the system, and this data is used for future improvements and accuracy enhancements.
[0654] The system for implementing this invention combines various means to enable communication between users in different languages. First, the terminal uses speech recognition software to accept voice input. Specifically, it utilizes the speech_recognition library to convert voice data into text. This allows users to naturally begin a conversation.
[0655] Once the audio data is converted into text, that text is sent to a server via a communication method. The server uses a generative AI model to translate the received text into a different symbolic representation. Here, machine learning is used in the generative model to improve translation accuracy, and GPT-3 or similar models may be applied. The translated symbolic representation is then sent back to the terminal.
[0656] The terminal converts the received symbolic representations into speech data using speech synthesis means and provides it to the user in natural-sounding speech using a speech synthesis library such as gTTS. Furthermore, the user can evaluate the quality of the translated content through an interface means, and this evaluation is sent to the server and used to improve translation accuracy in the future.
[0657] For example, if a Japanese-speaking user visiting a virtual sales environment asks, "Could you tell me about the features of this product?", the salesperson will respond in English, "This product is eco-friendly and cost-effective." The terminal then translates this into Japanese and outputs the translated audio in Japanese. In this way, smooth communication is achieved between users who speak different languages.
[0658] An example of a prompt message would be, "Recognize the user's question, translate it into English, answer the question, translate that answer back into Japanese, and play it aloud." This allows the system to operate as planned and support the user's intended communication.
[0659] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0660] Step 1:
[0661] The device receives voice input from the user. It acquires voice data through the microphone and converts this voice data into text using the speech_recognition library. The input is the user's voice, and the output is the textual representation of that voice. Specifically, the voice signal is captured, and speech recognition is performed in the backend process.
[0662] Step 2:
[0663] After the character representation is generated, the terminal sends this data to the server via a communication method. The input is the character representation generated in the previous step, and the output is the completion of the data transmission to the server. In this operation, data is securely transferred using a network protocol.
[0664] Step 3:
[0665] The server uses a generative AI model to translate the received text representation into a symbolic representation in a different language. The input is the received text representation, and the output is the translated symbolic representation. For data processing, machine learning algorithms are used, and the weights within the model are utilized to perform highly accurate language translation.
[0666] Step 4:
[0667] The server sends the translated symbolic representation back to the terminal. Here, the input is the translated symbolic representation, and the output is the completion of the transmission to the terminal. Specifically, the data is packaged according to the server's transmission protocol and transferred over the network.
[0668] Step 5:
[0669] The terminal reconstructs the received symbolic representation as audio data using speech synthesis technology and outputs the audio using gTTS or similar methods. The input is the symbolic representation received from the server, and the output is the synthesized speech. During the conversion to audio data, speech with natural intonation is generated and played back through the speaker.
[0670] Step 6:
[0671] Users provide feedback on translation quality through the terminal interface. This feedback is sent to the server and used to improve translation accuracy in the future. The input is the user's feedback data, and the output is the completion of sending the feedback to the server. Specifically, the user input on the interface is captured and then sent back to the server.
[0672] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0673] This invention is a real-time translation system that takes user emotions into account. This system is comprised of a combination of speech recognition, a generative model, speech synthesis, communication, an interface, and an emotion engine.
[0674] When a user begins speaking in their first language, the device captures their voice using its microphone. The captured audio data is converted into text by a speech recognition system, and then an emotion engine extracts additional information, including the user's emotions. This emotion information is sent to the server along with the text data.
[0675] The server uses the received text data and sentiment information to perform translation into a second language using a generative model. During this process, the sentiment engine provides sentiment information to the generative model, and the translated expression is adjusted to best reflect the user's emotions. The translated text data is then transmitted to the terminal via a communication device.
[0676] The device converts this translated text into speech data using speech synthesis technology and outputs it through the speaker as natural-sounding speech that takes emotions into consideration. This allows the user to convey their intended emotions and nuances.
[0677] Furthermore, the system includes a function to receive user feedback and send it to the server. This feedback is used to refine the generative model and sentiment engine, improving the accuracy of translation and sentiment recognition.
[0678] For example, if a user gets a little excited and says "I'm so happy today!" in Japanese, the system recognizes that emotion and translates it into English as "I'm super happy today!" with emotion, outputting it in appropriate voice. This ensures that the user's emotions are accurately conveyed to others, facilitating smoother communication.
[0679] The following describes the processing flow.
[0680] Step 1:
[0681] When a user speaks in their first language, the device uses its microphone to capture the audio. This allows the spoken content to be obtained as digital audio data.
[0682] Step 2:
[0683] The device analyzes the captured audio data using speech recognition and converts it into corresponding text data. During this process, signal processing is also performed to remove noise and improve sound quality.
[0684] Step 3:
[0685] The device passes the converted text data to an emotion engine to analyze the user's emotions and extract emotional information. This emotional information indicates what emotions are conveyed in the utterance.
[0686] Step 4:
[0687] The device sends a dataset containing text data and sentiment information to the server via a communication method. This enables translation processing on the server side.
[0688] Step 5:
[0689] The server inputs the received text data into a generative model and performs translation into a second language. During this process, the emotional information provided by the emotion engine is taken into consideration, and the translated text is adjusted to appropriately express the user's emotions.
[0690] Step 6:
[0691] The server returns the translated text data to the terminal. The returned data includes expressions that reflect emotions.
[0692] Step 7:
[0693] The device processes the translated text data using speech synthesis technology to generate emotionally appropriate audio data. The generated audio is then delivered to the user through the speaker.
[0694] Step 8:
[0695] Users review the translated audio and provide quality feedback through their device's interface. This feedback is then sent back to the server via communication channels and used to improve the system in the future.
[0696] (Example 2)
[0697] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0698] In translation systems, it is difficult to appropriately reflect user emotions and incorporate natural emotional expressions into the translation results. Furthermore, conventional translation systems have the challenge of not being able to effectively utilize user feedback to improve translation accuracy.
[0699] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0700] In this invention, the server includes acoustic analysis means, emotion analysis means, and generative model means. This enables natural translation expressions that reflect the user's emotions.
[0701] "Acoustic analysis means" refers to technology that has the function of collecting audio data and converting it into text data.
[0702] "Emotional analysis means" refers to a technology that processes data to extract user emotional information from text data.
[0703] A "generative model means" is a technology that uses received text data and sentiment information to translate into different languages and express emotions appropriately.
[0704] "Sound generation means" refers to a technology that generates audio data based on translated text data and emotional information, and outputs it as audio that takes emotional expression into account.
[0705] "Communication control means" refers to technology that controls the processing for sending and receiving voice data and text data over a network.
[0706] "Input device means" refers to technology that provides an interface for receiving feedback from users.
[0707] This invention provides a system that enables advanced real-time translation that takes user emotions into consideration. Specific embodiments are shown below.
[0708] First, the device uses a microphone to capture audio data of the user's first language. Speech recognition software is then used as an acoustic analysis tool to convert this audio data into text. Commercially available speech recognition technologies, such as the Google Speech-to-Text API, can be used for this process.
[0709] Next, the terminal extracts user emotion information from the converted text data using emotion analysis tools. Here, natural language processing models and machine learning algorithms can be utilized as software widely used for emotion analysis. The emotion information obtained through this process is sent to the server along with the text data.
[0710] The server uses a generative model to perform advanced translation based on the received text data and sentiment information. In this embodiment, the generative AI model can be, for example, an open-source model or a commercial advanced natural language processing model. The generative model utilizes sentiment information to adjust the expression of the translation, enabling it to accurately convey the emotions intended by the user.
[0711] The translation results are sent from the server to the terminal and converted into audio data using speech synthesis software as a means of generating sound. Natural-sounding speech can be achieved by using Amazon Polly or other commercial speech synthesis services. The output is then delivered to the user through the speaker with appropriate intonation.
[0712] As a concrete example, if a user says "I'm so happy today!", the system takes the emotion into consideration and uses the prompt "Translate the following excited expression from Japanese to English while preserving the user's emotional excitement" to generate the translation "I'm super happy today!", which is then appropriately output as speech.
[0713] Furthermore, the terminal receives feedback from the user regarding the translation results and sends this feedback to the server using communication control means. This feedback is used to improve the accuracy of the generative model and sentiment analysis. In this way, the system can continuously improve the quality of translations and the accuracy of sentiment.
[0714] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0715] Step 1:
[0716] The device captures the user's voice using a microphone. The user's voice signal is used as input. This voice signal is acquired as analog data and converted to a digital format. The output is digitized voice data. This data is used for acoustic analysis in the next step.
[0717] Step 2:
[0718] The terminal inputs the captured audio data into speech recognition software, which converts it into text data. This uses acoustic analysis, and the specific data processing involves breaking down the audio signal into phonemes and converting them into strings. The output is text data representing what the user said as a string.
[0719] Step 3:
[0720] The terminal passes text data to an emotion analysis device to extract the user's emotions. The input in this step is text data obtained from a speech recognition device. Specifically, a natural language processing algorithm is used to analyze words and contexts that represent specific emotions within the text, and the emotion information is output as data. The output is metadata that includes emotional elements such as joy, anger, sadness, and happiness.
[0721] Step 4:
[0722] The terminal transmits text data and sentiment information to the server via a communication control means. The input here is text data and its sentiment metadata. A secure communication protocol is used to ensure data integrity when transmitting to the server. As output, the data received on the server side is provided to the generation model means.
[0723] Step 5:
[0724] The server inputs the received text data and sentiment information into a generative AI model to perform translation. In this step, the translation model performs data calculations to adjust the translation expression to reflect sentiment while performing language conversion based on the input data. Specifically, it uses prompt sentences to activate the generative AI model and outputs the optimal translation result. The output is a sentiment-sensitive translated text in a second language.
[0725] Step 6:
[0726] The server sends the translated text data to the terminal. The input is the translated text generated by the generative AI model. Communication control means are used for data transfer. The output is the translated text received on the terminal side.
[0727] Step 7:
[0728] The terminal inputs translated text data into a speech synthesis system and converts it into speech data. The input is the translated text data. Specifically, the speech synthesis engine converts the text data into a speech waveform and synthesizes it as a natural-sounding voice with emotion. The output is the speech data played back through the speaker.
[0729] Step 8:
[0730] The terminal receives feedback from the user and sends it to the server. An interface is used for inputting the user's opinion and emotional response to the translation results. The feedback data sent to the server contributes to improving the generative AI model and sentiment analysis methods. The output is improvement information that will be useful for future translation accuracy improvements.
[0731] (Application Example 2)
[0732] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0733] Conventional translation systems can perform simple conversions between languages, but they have the problem of failing to adequately reflect the speaker's emotions and nuances, making natural communication difficult. In particular, in the field of content distribution, there is a demand for translations that allow viewers to accurately understand and enjoy the emotions of movies and dramas, but current technology is unable to adequately meet this requirement.
[0734] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0735] In this invention, the server includes speech recognition means for converting speech data into text data, generative model means for translating text data into different languages based on emotional information, and speech synthesis means for converting translated text data that reflects emotional information into speech data. This enables natural translation that takes the user's emotions into account.
[0736] "Speech recognition means" refers to a technology that converts speech data into text data, making it possible to process the user's speech content as text information.
[0737] A "generative model means" is a technology for translating text into different languages using text data and sentiment information obtained by speech recognition means.
[0738] "Speech synthesis means" refers to a technology that converts translated text data into speech data and provides the user with natural-sounding speech output.
[0739] "Communication means" refers to the technology for sending and receiving text data obtained by speech recognition means, emotion information, and translated text data obtained by generative model means via a communication network.
[0740] An "interface means" is a means for users to input feedback regarding the quality and emotional expression of translations, and enables interaction with the user.
[0741] "Emotional information" refers to information extracted by analyzing the emotions and nuances contained in the user's utterance, and is an important element in the translation process.
[0742] The system for implementing this invention consists of a user terminal and a server. The terminal is equipped with speech recognition means, communication means, and speech synthesis means, which convert the user's voice into text data in real time and extract emotional information. The voice data captured on the terminal is then converted into text data through speech recognition software. At this time, an emotional recognition algorithm is used to analyze the user's emotions and generate corresponding emotional information.
[0743] Next, this text data and sentiment information are sent to the server via a communication method. The server uses a generative AI model to translate the text data into different languages. At this time, the sentiment information is input into the generative model, so the translated text takes sentiment into account. The translated text data is then sent back to the terminal.
[0744] The device uses the received translated text data to create natural-sounding, emotionally resonant voice data through speech synthesis. The user can then listen to the translated content through the audio output from the speaker.
[0745] As a concrete example, when a user is watching a movie, the emotionally rich lines spoken by the characters can be instantly translated into different languages, taking emotional information into account, and delivered to the viewer. Through this process, users can accurately understand the emotions and deepen their overall impression of the movie.
[0746] Examples of prompt messages are as follows:
[0747] "When Elsa begins to sing in the snowy mountains, feeling alone but yearning for freedom, please generate a translation that reflects her emotions and strength."
[0748] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0749] Step 1:
[0750] The terminal receives the user's voice as input and captures the audio data using the microphone. The captured audio data is sent to a speech recognition system. The speech recognition system processes the audio data to convert it into text data and passes this converted text data to the next step.
[0751] Step 2:
[0752] The device uses text data obtained by speech recognition to run an emotion recognition algorithm. This algorithm receives text data as input, analyzes the user's emotions from it, and extracts emotion information. The emotion information is output as data representing the emotions expressed in the user's speech.
[0753] Step 3:
[0754] Text data and sentiment information are transmitted from the terminal to the server using a communication method. The server then uses the acquired data as input and runs a generative AI model. The generative AI model translates the text into different languages based on the input data. In this process, sentiment information is reflected in the translation result, and the output is translated text data with adjusted expression.
[0755] Step 4:
[0756] The server sends the translated text data to the terminal. The terminal passes this received data as input to the speech synthesis system. The speech synthesis system converts the translated text data into speech data and generates speech that has been adjusted to convey emotions naturally.
[0757] Step 5:
[0758] The device outputs audio data synthesized by a speech synthesis system to the user through its speaker. This allows the user to hear the translated content in audio format.
[0759] Step 6:
[0760] Users input feedback on the provided translation into their device. This feedback is collected through the interface as opinions on the translation quality and sentiment expression, and is sent back to the server. The server uses this feedback information to improve the accuracy of its generative AI model and sentiment recognition.
[0761] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0762] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0763] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[0764] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0765] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. In the upper and lower directions of the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. Also, the upper side of the concentric circles is where "pleasant" emotions are located, and the lower side is where "unpleasant" emotions are located. In this way, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0766] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0767] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0768] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0769] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0770] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0771] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[0772] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[0773] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0774] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0775] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0776] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0777] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0778] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0779] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0780] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0781] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[0782] The following is further disclosed regarding the embodiments described above.
[0783] (Claim 1)
[0784] A speech recognition means for converting audio data into text data,
[0785] A generative model means for translating the aforementioned text data into different languages,
[0786] A speech synthesis means for converting the translated text data into speech data,
[0787] A communication means for sending and receiving text data obtained by the speech recognition means and translated text data obtained by the generative model means via a computer network,
[0788] The aforementioned interface means for the user to input feedback regarding the quality of the translation,
[0789] A system that includes this.
[0790] (Claim 2)
[0791] The system according to claim 1, wherein the generation model means updates model parameters using user feedback information to improve translation accuracy.
[0792] (Claim 3)
[0793] The system according to claim 1, wherein the speech recognition means performs signal processing for noise reduction and sound quality improvement.
[0794] "Example 1"
[0795] (Claim 1)
[0796] A speech recognition means that collects audio data, performs noise reduction and sound quality improvement, and converts the results into text data,
[0797] A generative model means for translating the aforementioned text data into different languages, comprising means for updating model parameters using user feedback,
[0798] A speech synthesis means that converts the translated text data into audio data and outputs it as natural-sounding speech,
[0799] A communication means for transmitting and receiving data obtained by the speech recognition means and generative model means via a digital communication network,
[0800] The aforementioned interface means for the user to provide feedback on the quality of the translation,
[0801] A system that includes this.
[0802] (Claim 2)
[0803] The system according to claim 1, wherein the generation model means evolves translation accuracy using user feedback information.
[0804] (Claim 3)
[0805] The system according to claim 1, wherein the speech recognition means performs signal processing for noise reduction and sound quality improvement.
[0806] "Application Example 1"
[0807] (Claim 1)
[0808] A speech recognition means for converting speech expressions into text expressions,
[0809] A generative model means for converting the aforementioned character representation into a different symbolic representation,
[0810] A speech synthesis means that converts the converted symbolic representation into speech data,
[0811] A communication means for transmitting and receiving character representations obtained by the speech recognition means and translated symbol representations obtained by the generative model means via a computer network,
[0812] The aforementioned interface means for the user to input an evaluation regarding the quality of the translation,
[0813] A conversation support method for enabling two-way communication in different languages in a virtual sales environment,
[0814] A system that includes this.
[0815] (Claim 2)
[0816] The system according to claim 1, wherein the generation model means updates the generation model parameters using evaluation information from users to improve translation accuracy.
[0817] (Claim 3)
[0818] The system according to claim 1, wherein the voice recognition means performs signal processing for suppressing unwanted signals and improving sound quality.
[0819] "Example 2 of combining an emotion engine"
[0820] (Claim 1)
[0821] An acoustic analysis means for converting audio data into text data,
[0822] A means for analyzing emotions to extract emotional information from the aforementioned text data,
[0823] A generative model means for translating the aforementioned text data and sentiment information into different languages,
[0824] Sound generation means that converts the translated text data and emotional information into audio data,
[0825] A communication control means that transmits and receives text data obtained by the acoustic analysis means and translation data obtained by the generation model means via a communication path,
[0826] The aforementioned input device means for the user to input feedback regarding the quality and emotional expression of the translation,
[0827] A system that includes this.
[0828] (Claim 2)
[0829] The system according to claim 1, wherein the generation model means takes the emotion information into consideration, updates the model structure using user feedback information, and improves the accuracy of translation and emotion expression.
[0830] (Claim 3)
[0831] The system according to claim 1, wherein the acoustic analysis means performs signal analysis for external noise reduction and sound quality improvement.
[0832] "Application example 2 when combining with an emotional engine"
[0833] (Claim 1)
[0834] A speech recognition means for converting audio data into text data,
[0835] A generative model means for translating the aforementioned text data and emotional information into different languages,
[0836] A speech synthesis means that converts translated text data into speech data using emotional information,
[0837] A communication means for transmitting and receiving text data obtained by the speech recognition means, emotion information, and translated text data obtained by the generative model means via a communication network.
[0838] The aforementioned interface means for the user to input feedback regarding the quality and emotional expression of the translation,
[0839] A system that includes this.
[0840] (Claim 2)
[0841] The system according to claim 1, wherein the generation model means updates model parameters using user feedback information and sentiment information to improve translation accuracy and the accuracy of sentiment expression.
[0842] (Claim 3)
[0843] The system according to claim 1, wherein the speech recognition means performs signal processing for noise reduction and sound quality improvement, and also extracts emotional information. [Explanation of Symbols]
[0844] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. A speech recognition means for converting audio data into text data, A generative model means for translating the aforementioned text data into different languages, A speech synthesis means for converting the translated text data into speech data, A communication means for sending and receiving text data obtained by the speech recognition means and translated text data obtained by the generation model means via a computer network, The aforementioned interface means for the user to input feedback regarding the quality of the translation, A system that includes this.
2. The system according to claim 1, wherein the generation model means updates model parameters using user feedback information to improve translation accuracy.
3. The system according to claim 1, wherein the speech recognition means performs signal processing for noise reduction and sound quality improvement.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A