Two-way voice communication system "Words are IKIGAI"

The two-way voice communication system automatically translates and verifies translations in real-time, addressing user-friendliness and accuracy issues of existing devices, facilitating seamless communication for foreign visitors and staff.

JP3252523UActive Publication Date: 2025-08-22村田 聡
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025001404U
Authority / Receiving Office
JP · JP
Patent Type
Utility models
Current Assignee / Owner
Filing Date
2025-05-03
Publication Date
2025-08-22
Estimated Expiration
2035-05-03

AI Technical Summary

Technical Problem

Existing translation devices require users to manually select languages and often display only translated text, making it difficult for non-native speakers to verify accuracy and correct mistranslations, and are not user-friendly for hospitality staff.

Method used

A two-way voice communication system that automatically recognizes and translates speech without language selection, displays original and translated text side by side, and provides immediate verification of translations through explicit operation steps.

Benefits of technology

Eliminates language barriers by ensuring accurate and immediate translation verification, enhancing user-friendliness and accessibility, particularly benefiting hospitality staff and foreign visitors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0003252523000001_ABST
    Figure 0003252523000001_ABST
Patent Text Reader

Abstract

To provide a multilingual two-way voice communication device that enables smooth communication with the rapidly increasing number of foreign visitors to Japan and allows users to instantly check and correct the content when misrecognition or mistranslation occurs. [Solution] This device is a two-way voice communication device that automatically identifies and translates spoken foreign languages ​​through voice input, and conveys the information in text and audio. It has a simple mechanism that performs the translation process with three tap steps, from starting voice input to confirmation and transmission, and displays the translated text and the original text side by side on the same screen, allowing both users to instantly check for misrecognition or mistranslation and make corrections by voice if necessary.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This invention relates to a multilingual interactive voice communication support system that can be used to eliminate communication obstacles with foreign visitors to Japan that can occur in hospitality service industries such as restaurants. [Background technology]

[0002] The main background technology of this invention is as follows: 1) Application of ChatGPT (large-scale language model) engine 2) UI structure to prevent misrecognition 3) Operational design based on thorough on-site operational research 4) Dialogue model design based on communication science and neuroscience Examples include: Summary of the Invention [Problem to be solved by the invention]

[0003] Background of the idea According to a survey by the Japan Tourism Agency, the second most common travel problem for foreign visitors to Japan is "communication at stores and facilities." Against this backdrop, there is a demand for devices and systems that allow smooth communication with the rapidly increasing number of foreign visitors. However, many employees in charge of customer service and other communication at stores and facilities have difficulty operating complex devices and systems, and it is thought that there will be an increasing need for easy-to-use communication support devices that can be operated easily and can accommodate foreigners who speak a variety of languages. assignment While existing translation devices may appear convenient at first glance, they are designed specifically for use when traveling abroad, requiring users to select the language to translate into first, making them difficult to use for foreigners who do not know the language they are speaking. Furthermore, many of these devices only display the translated text, making it difficult for users to immediately check and correct the content in the event of a misrecognition or mistranslation. To address these issues, there is a demand for a speech translation support system that is easy for anyone to use, prevents misrecognition, and allows users to immediately check the translation content. [Means for solving the problem]

[0004] In response to the background and issues, the present invention adopts the following technical configuration. Technical requirements and features This invention is a two-way voice communication system that automatically recognizes and translates any spoken language through voice input on a mobile device such as a smartphone, without the user having to select the language. Operation of the device does not require any text input other than tapping on an icon. The system of the present invention has the following features: (1) The translation process is performed through three explicit steps, from the start of voice input to the check icon (confirmation) → up arrow icon (send), preventing mis-sending and mis-recognition. (2) The translated text (translated language) and the original text (original language) are displayed side by side on the same screen, with the translated text in bold to improve visibility and the original text in normal style. (3) By displaying the original text, both users can immediately check whether there are any misunderstandings or mistranslations. (4) When inputting voice, there is no need to manually select the language; the system automatically determines and translates each utterance. (5) A configuration that allows users to operate the system while maintaining portability by using a smartphone. "Actual operation" The operation of this invention is described below. This system, called "Kotoba wa IKIGAI," is configured to run a speech translation support function on a smartphone device (with iOS or Android OS) using ChatGPT Custom GPT as its engine, based on human operations research. The system is started by opening the provided URL link on a smartphone; no special application installation is required. A microphone icon for voice input is displayed at the bottom of the screen, and translation can be performed by following the steps below. (1) Tap the microphone icon to start voice input. (2) Speak. (3) During recording, a wavy animation indicating the audio level is displayed in the input field, visually indicating that audio input is taking place. (4) After you finish speaking, a check icon (confirm) will appear. By tapping this, the speech will be converted into text and displayed in place of the wavy line animation. (5) Check the displayed text and if there are no problems, tap the up arrow icon (Send). (6) The translated text (translated language) is displayed in bold, and the original text (original language) is displayed in normal style on the same screen. (7) At the same time, the translated text is read aloud. You can mute the sound by adjusting the volume of your device. These operations allow users to perform speech translation without having to select a language, and both users can immediately check whether there are any misrecognitions or mistranslations. Furthermore, translation history is temporarily stored on the screen and can be used as a record of the conversation. The history is cleared when the application or system is closed. [Effects of the Invention]

[0005] With this device, users do not need to select a language; speech in any language is automatically recognized and translated, eliminating communication barriers between foreign visitors to Japan and customer service staff. Furthermore, by providing explicit operation methods from the start of voice input to confirmation and transmission, misrecognition and mistransmission are prevented, and both users can instantly check the accuracy of the translation. Furthermore, the translation results are displayed in a highly visible format, making it easy to identify and correct mistranslations. The voice reading function enables communication with the visually impaired and ensures accessibility (barrier-free). In addition, by supporting multiple devices such as smartphones, tablets, and personal computers, users can maintain freedom of device selection and portability, making it possible to easily operate this device in a variety of environments. The system's simple operation design allows users to start using it immediately after a simple training session, which will reduce the workload of employees in the hospitality industry, improve service quality, and contribute to alleviating labor shortages and expanding employment opportunities. In addition, the system of this invention is compatible with not only smartphones, but also tablet devices and personal computers (notebook or desktop). For users who have difficulty reading the small text on smartphones, a tablet or PC can be used, and compatibility with a wide range of devices contributes to barrier-free access for people of all ages. [Brief explanation of the drawings]

[0006] [Figure 1] The icon display area at the bottom right of the screen displays three options for selection: Microphone: Voice input → Check: Voice recognition confirmation → Arrow: Send to system. The window to the left of the icon displays the volume level waveform during voice input, and the recognized words during voice recognition. The recognized words are translated and displayed in bold, and simultaneously displayed in the spoken language in thin type.

Claims

[Claim 1] A two-way voice communication system that automatically determines the language of speech input and performs translation processing on a mobile device such as a smartphone, tablet, or personal computer, without the user having to manually select a language. It comprises the following: (1) A user interface that allows you to complete three steps of voice input, from starting to confirming and sending, simply by tapping an icon. (2) A function that displays the translation and the original text side by side on the same screen, allowing both users to check for misrecognition and mistranslation. (3) A function to transmit the translation content aloud using a voice reading function; (4) A configuration that allows portability for users and compatibility with a variety of devices.