System

The system addresses noise interference and operation challenges in hearing aids by using a voice assistance device with noise cancellation, generative AI, and text conversion, improving communication for elderly users.

JP2026028924APending Publication Date: 2026-02-20SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024131541
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-08-07
Publication Date
2026-02-20

AI Technical Summary

Technical Problem

Conventional hearing aids struggle with noise interference, difficulty in hearing specific frequency bands, cumbersome operation, and lack of text transcription for conversation review, especially affecting elderly users.

Method used

A system comprising a voice assistance device connected to a mobile communication device, with noise cancellation, generative AI for voice correction, and display for text conversion, enabling clear audio, customizable settings, and conversation transcription.

Benefits of technology

Enhances user communication by providing clear audio, customizable settings, and allowing review of past conversations, making it easier for elderly users to participate in discussions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026028924000001_ABST
    Figure 2026028924000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system, comprising: an intelligent voice assistant device; a mobile communication device electrically connected to the voice assistant device; a communication unit configured to customize voice data in the mobile communication device; a processing unit configured to perform noise cancellation processing on the voice data; a correction unit configured to correct the voice data to a natural language, wherein the correction unit comprises generative artificial intelligence; and a display unit configured to transcribe and display the voice data.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] Conventional hearing aids have problems with picking up noise that makes it difficult for users to clearly hear surrounding speech, and with making it difficult to hear speech in certain frequency bands. Furthermore, when elderly people use these hearing aids, they face the problem of cumbersome operation and difficulty adjusting settings. Furthermore, in addition to being able to clearly hear speech, they lack the ability to recognize and save conversation content as text, making it difficult to review past conversations. The present invention aims to solve these problems, promote the use of hearing aids, and improve user communication. [Means for solving the problem]

[0005] The present invention provides a system including a highly functional voice assistance device, a mobile communication device electrically connected to the voice assistance device, a communication means for customizing voice data in the mobile communication device, a processing means for performing noise cancellation processing on the voice data, a correction means equipped with generative artificial intelligence for correcting the voice data into natural language, and a display means for converting the voice data into text and displaying it.

[0006] In particular, the system is equipped with a communication means for transmitting voice data from a mobile communication device to a voice assist device in real time and a user-configurable equalization means, which allows users to easily operate the system and clearly hear voices even in specific frequency bands. Furthermore, by displaying and saving transcribed conversation content, users can review past conversations. These means solve the conventional problems and provide convenience and flexibility for users of voice assist devices.

[0007] An "audio assistance device" is an electronic device that amplifies surrounding sounds and plays them back at a volume and quality suitable for the user.

[0008] A "mobile communications device" is a portable device capable of transmitting and receiving data to and from other devices using wireless communications.

[0009] "Communication means" refers to the communication interface and protocol for sending and receiving audio data to other devices.

[0010] "Processing means" refers to hardware and software that executes particular algorithms on audio data to achieve a desired result.

[0011] The "correction means" is a system that includes artificial intelligence to process voice data so that it is easier for the user to hear and convert it into natural language.

[0012] "Display means" refers to a screen or other display device for converting voice data into text and visually displaying it.

[0013] "Equalizing" is a technology that corrects sound quality by adjusting the strength of audio signals in specific frequency bands.

[0014] "Noise cancellation" is a processing method for removing unwanted noise and chatter from audio data. [Brief explanation of the drawings]

[0015] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13]FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0016] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0017] First, the terms used in the following description will be explained.

[0018] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0019] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0020] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0021] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0022] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0023] [First embodiment]

[0024] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0025] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0026] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0027] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0028] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0029] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0030] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0031] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0032] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0033] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0034] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0035] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0036] The system of the present invention is composed of a highly functional voice assistant device and a mobile communication device electrically connected to it. In this system, the mobile communication device (e.g., a smartphone) plays a central role, providing functions such as voice data customization, noise cancellation, correction using generation AI, and transcription.

[0037] System configuration

[0038] 1. Audio assistants

[0039] The advanced audio assistant effectively amplifies surrounding sounds and has filters to reduce noise according to the user's hearing ability, and can be connected to a mobile communication device via a wireless communication interface such as Bluetooth.

[0040] 2. Mobile Communication Devices

[0041] The device is a portable electronic device such as a smartphone or tablet, which connects to the audio assistant via wireless communication such as Bluetooth and has an application installed to collect, process, and display audio data.

[0042] Program processing flow

[0043] The program processing of this system is carried out as follows.

[0044] 1. Bluetooth connection

[0045] User: Launch the smartphone app and set the audio assistant device to Bluetooth pairing mode.

[0046] Device (smartphone): Search for Bluetooth devices, select the audio assistant device, and establish pairing.

[0047] 2. Collection and transmission of voice data

[0048] Device (smartphone): The microphone collects surrounding sounds and stores them in a buffer. The audio data in the buffer is sent in real time to the audio assistant device via wireless communication.

[0049] Audio assistant: Decodes the received audio data and adjusts the volume and frequency characteristics to suit the user before playing it back.

[0050] 3. Equalization and noise cancellation

[0051] User: Adjust equalization and noise cancellation parameters in the settings screen of the smartphone app.

[0052] Device (smartphone): Receives user settings, updates the filtering algorithm, applies a noise cancellation filter to the audio data, and sends it to the audio assistant device.

[0053] Audio assistant: Apply filtered audio data to reduce external noise before playback.

[0054] 4. Transcribe and save conversations

[0055] Terminal (smartphone): Audio data collected by a microphone is sent to a server, where real-time voice recognition is performed.

[0056] Server: Analyzes the voice data, generates text data, and sends it back to the smartphone.

[0057] Device (smartphone): The generated text is displayed on the screen and saved in the internal storage.

[0058] 5. Language Correction by Generative AI

[0059] Device (smartphone): Sends inaudible audio data to the generating AI.

[0060] Server: Generative AI analyzes the voice data and corrects it into natural language.

[0061] Terminal (smartphone): Receives the corrected data and sends it to the audio assistant device for playback.

[0062] Specific examples

[0063] 1. Use at family dinners

[0064] User: Places smartphone in the center of the table and connects to audio assistant via Bluetooth. Adjusts equalization and noise cancellation settings through the app.

[0065] Device (smartphone): The microphone picks up surrounding conversations and transmits them in real time to the voice assistant. The voice data is sent to the server, which then retrieves the text data and displays it on the screen.

[0066] Server: Analyzes the voice data, generates text in real time, and sends it back to the smartphone.

[0067] Terminal (smartphone): Saves the text data and sends the correction results to the voice assistant device.

[0068] Thus, the system of the present invention allows elderly people to participate more naturally in conversations with their families, providing a practical way to compensate for lost communication.

[0069] The processing flow will be explained below.

[0070] Step 1:

[0071] User: Downloads and installs the smartphone app. Opens the app and creates a new account.

[0072] Step 2:

[0073] User: Power on the audio assistant and set it to Bluetooth pairing mode.

[0074] Step 3:

[0075] Device (smartphone): Open the Bluetooth settings screen and search for the audio assist device. Select the audio assist device from the search results and send a pairing request. If pairing is successful, establish a connection with the audio assist device.

[0076] Step 4:

[0077] Device (smartphone): Once a Bluetooth connection with the audio assistant is established, obtain microphone permission from the app and begin collecting audio.

[0078] Step 5:

[0079] Device (smartphone): The smartphone's microphone collects surrounding sounds and stores them in a buffer in real time.

[0080] Step 6:

[0081] Terminal (smartphone): Periodically performs digital signal processing on the audio data in the buffer and transmits it to the audio assistant device via Bluetooth.

[0082] Step 7:

[0083] Audio assistant: Decodes the received audio data and adjusts the volume and frequency characteristics to suit the user before playing it back.

[0084] Step 8:

[0085] User: Adjust equalization and noise cancellation parameters in the settings screen of the smartphone app.

[0086] Step 9:

[0087] Device (smartphone): Receives user setting changes and updates the voice data filtering algorithm. Digital filtering is applied to reduce noise before sending the audio data to the voice assistant device.

[0088] Step 10:

[0089] Audio assistant: Reconditions the received filtered audio data to play the best audio for the user.

[0090] Step 11:

[0091] Device (smartphone): Audio data during conversation is sent to the server in real time in streaming format.

[0092] Step 12:

[0093] Server: The received voice data is analyzed using a voice recognition engine and text data is generated.

[0094] Step 13:

[0095] Server: The generated text data is sent back to the smartphone.

[0096] Step 14:

[0097] Device (smartphone): The generated text data is displayed within the app and saved in the internal storage.

[0098] Step 15:

[0099] User: Tells the app to mark parts of the audio that are difficult to hear.

[0100] Step 16:

[0101] Device (smartphone): Sends the marked audio data to the server and requests correction by the generation AI.

[0102] Step 17:

[0103] Server: The generative AI analyzes the voice data and corrects it into natural language.

[0104] Step 18:

[0105] Server: Sends the corrected audio data back to the smartphone.

[0106] Step 19:

[0107] Terminal (smartphone): Notifies the user of the received correction data and plays it on the audio assistant device.

[0108] In this way, the system of the present invention provides a device that users can easily operate and set up, and not only improves the quality of speech but also has a transcription function that allows users to review past conversations. This system allows users to achieve more natural and comfortable communication.

[0109] Example 1

[0110] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0111] Conventional speech assist devices simply amplify speech, and are not effective enough in noisy environments or when conversations are difficult to hear. Furthermore, they lack the ability to convert speech into text, making it impossible to visually confirm the content of a conversation. This creates challenges for communication, particularly for the elderly and those with hearing impairments.

[0112] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0113] In this invention, the server includes a high-performance audio amplifier, a portable information terminal electrically connected to the audio amplifier, a communication means for customizing audio data on the portable information terminal, a processing means for performing noise reduction processing on the audio data, a correction means having a generative artificial intelligence for correcting the audio data to natural language, a display means for converting the audio data into text and displaying it, a transmission means for storing the audio data in a buffer on the portable information terminal and transmitting the audio data to the audio amplifier in real time via wireless communication, and a setting means for the user to adjust equalization and noise reduction settings on the portable information terminal. This allows for clear audio even in noisy environments and for smooth communication by converting audio data into text and displaying it in real time.

[0114] An "audio amplifier" is a device for audio assistance that has the function of amplifying sound with high precision and reducing noise.

[0115] "Mobile information terminal" is a general term for portable electronic devices such as smartphones and tablets.

[0116] "Communication means" refers to devices or software that have the functionality to collect, send, and receive voice data.

[0117] "Processing means" refers to hardware or software for performing specific processing on audio data.

[0118] "Generative AI" is an AI technology that analyzes voice data and corrects it to make it more natural.

[0119] "Correction means" refers to a device or software that has the function of converting and correcting voice data into natural language.

[0120] "Display means" refers to a device or software that visually presents the text-converted voice data to the user.

[0121] A "buffer" is a storage area for temporarily storing audio data.

[0122] "Transmission means" refers to a device or software that has the function of transmitting audio data to another device via wireless communication.

[0123] "Setting means" refers to a device or software that has a function that allows the user to adjust sound quality and filtering parameters.

[0124] "Equalizing" is a process that emphasizes or attenuates specific frequency bands in audio.

[0125] "Noise reduction" refers to a processing technique for removing unwanted noise from an audio signal.

[0126] "Text conversion" refers to the process of converting voice data into text data.

[0127] "Wireless communication" refers to a communication technology for sending and receiving data without using cables.

[0128] The system of this invention is primarily composed of a high-performance sound amplifier, a mobile information terminal (such as a smartphone or tablet), and a server. This system provides clear audio even in noisy environments, converts audio data into text and displays it in real time, and enables smooth conversations with elderly people and those with hearing impairments.

[0129] Hardware configuration

[0130] Sound amplification equipment:

[0131] It is equipped with high-precision audio amplification and noise reduction filters.

[0132] It can be connected to a mobile information terminal via wireless communication such as Bluetooth.

[0133] Mobile devices:

[0134] A smartphone or tablet with a voice assistant application installed.

[0135] Surrounding sounds are picked up through the microphone and stored in a buffer.

[0136] Wireless communication (such as Bluetooth) is used for data communication.

[0137] server:

[0138] It analyzes voice data, transcribes it, and corrects it using AI. It has high-performance processing capabilities and performs secure data communication.

[0139] Software configuration

[0140] Voice-assisted applications:

[0141] Customize audio data, set noise reduction, transcribe audio in real time, and save text.

[0142] It has a settings screen that allows you to adjust equalization and noise reduction parameters.

[0143] Generative AI models:

[0144] It has the ability to analyze voice data and correct it into natural language.

[0145] It runs on the server and returns the correction results to the mobile information terminal.

[0146] Specific examples of operation procedures

[0147] Here, as a concrete example of how the system is used, we will show a scene where the system is used at a family dinner party.

[0148] Example of use at a family dinner:

[0149] User: Places smartphone in the center of the table and connects to sound amplification device via Bluetooth. Adjusts equalization and noise cancellation settings through the app.

[0150] Device (smartphone): The microphone collects the conversations during the dinner party and transmits them in real time to an audio amplifier. The audio data is then sent to a server to obtain text data, which is then displayed on the smartphone screen.

[0151] Server: Analyzes the voice data, generates text in real time, and sends it back to the smartphone. For example, the utterance "Mom, would you like another serving?" is generated as text data.

[0152] Device (smartphone): The generated text data is displayed on the screen so the user can check the content of the conversation. The text data is also saved in the internal storage for later reference. If necessary, language correction is performed on the voice data, and the correction results are sent to the audio amplification device.

[0153] Prompt Sentence Examples

[0154] "How exactly do I use a speech assistive device at a family dinner?"

[0155] This system allows elderly and hearing-impaired people to participate more naturally in conversations with their families and compensate for lost communication.

[0156] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0157] Step 1: Bluetooth connection

[0158] Input: Sound amplification device in pairing mode. User launches smartphone app.

[0159] Operation:

[0160] User: Launch the smartphone app and put the audio amplifier into Bluetooth pairing mode by pressing and holding the pairing button on the audio amplifier.

[0161] Device (smartphone): Start Bluetooth device search and detect the sound amplifier. Select the sound amplifier from the device list and establish pairing. The message "Connection completed" will be displayed.

[0162] Output: "Connected" message upon successful pairing connection.

[0163] Step 2: Collect and send audio data

[0164] Input: Smartphone connected to sound amplification device via Bluetooth.

[0165] Operation:

[0166] Device (smartphone): Uses a microphone to collect surrounding sounds in real time. This data is temporarily stored in a buffer. Consider a dinner party scenario, where the sound of conversations is collected.

[0167] Terminal (smartphone): The audio data stored in the buffer is sent to the audio amplifier via wireless communication (Bluetooth). The data is divided and sent in packet format.

[0168] Audio Amplification Device: Decodes the received audio data and adjusts the volume and frequency characteristics, so that the audio is played back to the user in a clear and crisp manner.

[0169] Output: The audio data with adjusted volume and frequency response reaches the user.

[0170] Step 3: Equalization and noise cancellation settings

[0171] Input: A user is accessing the settings screen within a smartphone app.

[0172] Operation:

[0173] Users: Adjust equalization and noise cancellation parameters in the app's settings, for example, by boosting high frequencies.

[0174] Device (smartphone): Receives user settings and updates the filtering algorithm. When the new settings are reflected in the app, the algorithm processes the audio data accordingly.

[0175] Terminal (smartphone): The filtered audio data is sent back to the audio amplifier. After noise removal, the quality of the audio data is improved.

[0176] Sound amplifier: Reproduces the filtered audio data so that the user hears it with external noise reduced.

[0177] Output: The filtered audio data is delivered to the user.

[0178] Step 4: Transcribe and save the conversation

[0179] Input: Audio data is collected by a smartphone and sent to a server via the internet.

[0180] Operation:

[0181] Device (smartphone): Sends voice data collected by the microphone to the server and starts the voice recognition process.

[0182] Server: The received voice data is converted into text data using a generative AI model. The utterance "Hello, how are you?" is generated as text data.

[0183] Server: Sends the generated text data back to the smartphone.

[0184] Device (smartphone): The text data is displayed on the screen so that the user can check the contents of the conversation, and the text data is saved in the internal storage for later reference.

[0185] Output: The textual audio data is displayed on the screen and saved.

[0186] Step 5: Language correction by generative AI

[0187] Input: Difficult-to-hear audio data collected by a smartphone is sent to the generation AI.

[0188] Operation:

[0189] Device (smartphone): Sends inaudible audio data to the generating AI.

[0190] Server: The generation AI analyzes the voice data and corrects it into natural language. Unclear parts are corrected and clear text data is generated.

[0191] Server: Sends the corrected data back to the smartphone.

[0192] Terminal (smartphone): Receives the corrected data, sends it to an audio amplifier, and plays it back. The corrected audio reaches the user, making it easier to hear.

[0193] Output: The corrected audio data is delivered to the user.

[0194] By using specific actions performed at each step, this system can effectively provide communication support to the elderly and the hearing impaired.

[0195] (Application example 1)

[0196] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0197] Elderly people and those with hearing problems often have difficulty hearing in-store announcements and conversations with store clerks when shopping comfortably in physical stores. This not only degrades the quality of the shopping experience, but can also lead to missing important information. Furthermore, existing voice assistants lack effective noise cancellation and voice correction, which can lead to reduced speech recognition accuracy due to environmental noise.

[0198] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0199] In this invention, the server includes a high-performance voice assistance device, a mobile communication device electrically connected to the voice assistance device, a communication means for customizing voice data on the mobile communication device, a processing means for performing noise cancellation processing on the voice data, a correction means having a generation artificial intelligence for correcting the voice data to natural language, a display means for converting the voice data into text and displaying it, a voice output means for playing back the generated text data, and a server communication means for transmitting the voice data to an external server and retrieving the corrected text data. This allows users who require hearing assistance to easily hear in-store announcements and conversations with store clerks while shopping in a physical store, as well as visually confirm the corrected text data. Furthermore, correcting and playing back the generated text data can provide a natural conversation experience.

[0200] A "high-performance audio assist device" is a device that effectively amplifies surrounding sounds and is equipped with a filter to reduce noise according to the user's hearing ability.

[0201] A "mobile communications device" is a portable electronic device such as a smartphone or tablet.

[0202] The "communication means" is a function including a wireless communication interface for connecting a high-performance voice assistant device and a mobile communication device and customizing voice data.

[0203] "Processing means" refers to algorithms or programs for performing noise cancellation processing on speech data in a mobile communication device.

[0204] "Correction means" refers to a function that uses generative AI to correct voice data into natural language.

[0205] "Display means" refers to a display or screen that converts voice data into text and displays it.

[0206] The "audio output means" refers to a function such as a speaker or earphone for reproducing the generated text data as audio.

[0207] The "server communication means" is a function for transmitting voice data from the mobile communication device to an external server and obtaining corrected text data from the server.

[0208] To implement this invention, the following system configuration and operation method are required. A highly functional audio assistance device is used in combination with a mobile communication device (smartphone or tablet). This enables hearing assistance in a store.

[0209] System Configuration

[0210] 1. Advanced audio assistive devices:

[0211] It effectively amplifies surrounding sounds and has a filter to reduce noise according to the user's hearing ability, and can be connected to a mobile communication device via a wireless communication interface such as Bluetooth.

[0212] 2. Mobile communication devices:

[0213] It is a portable electronic device that connects to the audio assistant via wireless communication such as Bluetooth. This device is equipped with an application that customizes audio data, cancels noise, corrects it using generative AI, transcribes it, and plays it back.

[0214] Hardware and software used

[0215] 1. Hardware:

[0216] Smartphone (iOS or Android)

[0217] High-performance Bluetooth-enabled audio assist device

[0218] Your smartphone's built-in microphone and speaker (or earphones)

[0219] 2. Software:

[0220] Speech recognition libraries (e.g., speech_recognition, Google Speech API)

[0221] Text-to-speech engine (e.g. pyttsx3)

[0222] Server communication library (e.g. requests)

[0223] External generation AI server (API)

[0224] Specific examples

[0225] 1. Scenario 1:

[0226] In a physical store, a user stands in front of a shelf and asks a store clerk where a product is located. The user launches an application on their smartphone and types the question into the microphone.

[0227] Example prompt:

[0228] "Please ask the microphone where the item is."

[0229] 2. Scenario 2:

[0230] The user listens to an announcement about a new campaign in front of the cash register at a physical store. The smartphone is placed in a location where the announcement can be heard, and the speech is converted into text, corrected, and the corrected text is output as voice.

[0231] Example prompt:

[0232] Please check the announcement

[0233] Processing flow

[0234] 1. Device (smartphone):

[0235] The microphone picks up surrounding sounds and transmits the audio data via Bluetooth to a high-performance audio assist device, which simultaneously processes the audio data in real time and transmits it to an external server as needed.

[0236] 2. Server:

[0237] It receives voice data, corrects it into natural language using generative AI, and sends the corrected text data back to the smartphone.

[0238] 3. Device (smartphone):

[0239] The received text data is displayed on a display and output as voice using a text-to-speech engine.

[0240] This allows users who require hearing assistance to easily hear in-store announcements and conversations with store clerks while shopping in a physical store, as well as visually confirm the information. Furthermore, by correcting and playing back the generated text data, a natural conversation experience can be provided.

[0241] The system enables people with hearing problems to more comfortably participate in public activities, reducing social barriers.

[0242] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0243] Step 1:

[0244] User: Launches the smartphone app and sets the audio assistant device to Bluetooth pairing mode. Through the smartphone app interface, the smartphone searches for the audio assistant device via Bluetooth and establishes pairing. Based on this input, the smartphone connects to the audio assistant device via Bluetooth communication. When pairing is successful, the device outputs a notification that the connection is complete.

[0245] Step 2:

[0246] Device (smartphone): The smartphone's microphone collects surrounding sounds in real time and stores them in a buffer. This collected sound data is used as input, and the audio data in the buffer is processed in real time and sent via wireless communication to the audio assistance device. The audio assistance device receives the transmitted audio data, converts it into volume and frequency characteristics adjusted to suit the user's hearing, and plays it back.

[0247] Step 3:

[0248] User: Adjusts equalization and noise cancellation parameters using the settings screen in the smartphone app. Using this setting information as input, the smartphone updates the filtering algorithm and resends the audio data with the noise cancellation filter applied to the audio assistive device. This allows the audio assistive device to reproduce audio with reduced external noise.

[0249] Step 4:

[0250] Terminal (smartphone): The voice data collected by the microphone is sent to the server, which is then requested to perform voice recognition processing. In this process, the server analyzes the voice data as input and generates text data. This generated text data is sent back to the smartphone, and is then displayed on the smartphone screen.

[0251] Step 5:

[0252] Server: The server inputs the voice data sent from the smartphone into a generative AI model and corrects difficult-to-hear parts into natural language. The generative AI corrects the voice data by analyzing it and converting it into appropriate language based on the context. The corrected text data is then sent back from the server to the smartphone.

[0253] Step 6:

[0254] Terminal (smartphone): Receives the corrected text data and displays it on the screen. It also uses a text-to-speech engine to output the text data as voice, allowing the user to confirm the corrected information both visually and audibly.

[0255] Step 7:

[0256] User: In a brick-and-mortar store, the user continues to actively use their smartphone, converting surrounding information into text as needed and receiving natural-language enhanced speech information. For example, if a user asks a store clerk where a product is located, the smartphone microphone collects the question as input and receives enhanced speech that has gone through all the processing steps as output, making the clerk's response easier to understand.

[0257] These processing steps enable users to enjoy sophisticated voice assistance and a natural conversational experience in physical stores.

[0258] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0259] The present invention relates to a system that combines a highly functional voice assistant device with an electrically connected mobile communication device and further includes an emotion engine that recognizes the user's emotions. This system customizes voice data, cancels noise, corrects using generative AI, converts text, and optimizes voice data through emotion recognition.

[0260] System configuration

[0261] 1. Audio assistants

[0262] High-performance audio assistants reduce external noise and provide clear audio to users, and can be connected to mobile communication devices via wireless communication interfaces such as Bluetooth.

[0263] 2. Mobile Communication Devices

[0264] The device is a portable electronic device such as a smartphone or tablet, which is connected to the audio assistant via Bluetooth and has an application installed to collect, process, and display audio data.

[0265] 3. Emotion Engine

[0266] The emotion engine identifies emotions from user speech and input data and uses that information to optimize other processes, such as adjusting filtering parameters for specific emotions in the speech data.

[0267] Program processing flow

[0268] The program processing of this system is carried out as follows.

[0269] 1. Bluetooth connection

[0270] User: Launch the smartphone app and set the audio assistant device to Bluetooth pairing mode.

[0271] Device (smartphone): Search for Bluetooth devices, select the audio assistant device, and establish pairing.

[0272] 2. Collection and transmission of voice data

[0273] Device (smartphone): Collects surrounding audio with a microphone and stores it in a buffer in real time.

[0274] Terminal (smartphone): Digital signal processing is performed on the audio data in the buffer, and the data is sent to the audio assistant device via Bluetooth.

[0275] Audio assistant: Decodes the received audio data and adjusts the volume and frequency characteristics to suit the user before playing it back.

[0276] 3. Equalization and noise cancellation

[0277] User: Adjust equalization and noise cancellation parameters in the settings screen of the smartphone app.

[0278] Device (smartphone): Receives user settings, updates the filtering algorithm, and sends noise-reduced audio data to the audio assistant device.

[0279] Audio assistant: Re-adjusts the filtered audio data for optimal playback.

[0280] 4. Transcribe and save conversations

[0281] Device (smartphone): Audio data during conversation is sent to the server in real time in streaming format.

[0282] Server: Analyzes the voice data, generates text data, and sends it back to the smartphone.

[0283] Device (smartphone): Text data is displayed in the app and saved in the internal storage.

[0284] 5. Language Correction by Generative AI

[0285] Device (smartphone): Sends inaudible audio data to the generating AI.

[0286] Server: Generative AI analyzes the voice data and corrects it into natural language.

[0287] Terminal (smartphone): Receives the corrected data and plays it on the audio assistant device.

[0288] 6. Emotion Recognition by Emotion Engine

[0289] Device (smartphone): Sends voice data collected by the microphone to the emotion engine.

[0290] Emotion engine: Analyzes voice data and recognizes the user's emotions. The recognized emotional information is reflected in filtering parameters and generation AI.

[0291] Device (smartphone): Optimize voice data settings based on feedback from the emotion engine.

[0292] Specific examples

[0293] Use at family dinners

[0294] User: Places smartphone in the center of the table, connects audio assistant via Bluetooth, adjusts equalization and noise cancellation settings through the app, and enables emotion engine.

[0295] Device (smartphone): The microphone collects surrounding conversations and transmits them in real time to the voice assistant. The voice data is sent to the server to obtain text data, which is then displayed on the screen. The emotion engine analyzes the collected voice data, recognizes the user's emotions, and dynamically adjusts voice filtering parameters based on that data.

[0296] Server: Analyzes the voice data, generates text in real time, and sends it back to the smartphone.

[0297] Terminal (smartphone): Saves the text data and sends the correction results to the voice assistant device.

[0298] In this way, the system of the present invention is designed to be easy for users to operate, and the emotion engine optimizes voice data, enabling more effective and natural communication.

[0299] The processing flow will be explained below.

[0300] Step 1:

[0301] User: Downloads and installs the smartphone app. Opens the app and creates a new account.

[0302] Step 2:

[0303] User: Power on the audio assistant and set it to Bluetooth pairing mode.

[0304] Step 3:

[0305] Device (smartphone): Open the Bluetooth settings screen and search for the audio assist device. Select the audio assist device from the search results and send a pairing request. If pairing is successful, establish a connection with the audio assist device.

[0306] Step 4:

[0307] Device (smartphone): Once a Bluetooth connection with the audio assistant is established, obtain microphone permission from the app and begin collecting audio.

[0308] Step 5:

[0309] Device (smartphone): The smartphone's microphone collects surrounding sounds and stores them in a buffer in real time.

[0310] Step 6:

[0311] Terminal (smartphone): Periodically performs digital signal processing on the audio data in the buffer and transmits it to the audio assistant device via Bluetooth.

[0312] Step 7:

[0313] Audio assistant: Decodes the received audio data and adjusts the volume and frequency characteristics to suit the user before playing it back.

[0314] Step 8:

[0315] User: Adjust equalization and noise cancellation parameters in the settings screen of the smartphone app.

[0316] Step 9:

[0317] Device (smartphone): Receives user setting changes and updates the voice data filtering algorithm. Digital filtering is applied to reduce noise before sending the audio data to the voice assistant device.

[0318] Step 10:

[0319] Audio assistant: Reconditions the received filtered audio data to play the best audio for the user.

[0320] Step 11:

[0321] Device (smartphone): Audio data during conversation is sent to the server in real time in streaming format.

[0322] Step 12:

[0323] Server: The received voice data is analyzed using a voice recognition engine and text data is generated.

[0324] Step 13:

[0325] Server: Sends the generated text data back to the smartphone.

[0326] Step 14:

[0327] Device (smartphone): The generated text data is displayed within the app and saved in the internal storage.

[0328] Step 15:

[0329] User: Type into the app to mark parts of the audio that are difficult to hear.

[0330] Step 16:

[0331] Device (smartphone): Sends the marked audio data to the server and requests correction by the generation AI.

[0332] Step 17:

[0333] Server: Generative AI analyzes the voice data and corrects it into natural language.

[0334] Step 18:

[0335] Server: Sends the corrected audio data back to the smartphone.

[0336] Step 19:

[0337] Terminal (smartphone): Notifies the user of the received correction data and plays it on the audio assistant device.

[0338] Step 20:

[0339] Device (smartphone): Sends voice data collected by the microphone to the emotion engine.

[0340] Step 21:

[0341] Emotion engine: Analyzes voice data and recognizes the user's emotions. The recognized emotional information is reflected in filtering parameters and generation AI.

[0342] Step 22:

[0343] Device (smartphone): Optimize voice data settings based on feedback from the emotion engine.

[0344] In this way, the system of the present invention provides a device that users can easily operate and configure, and by optimizing voice data using an emotion engine, it is possible to achieve more effective and natural communication.

[0345] Example 2

[0346] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0347] In recent years, advances in speech recognition and artificial intelligence technologies have improved the convenience of voice-based user interfaces. However, there are problems with degradation of voice data quality due to environmental noise and the user's emotional state. Furthermore, if voice data is not converted to text or corrected in real time, the user experience is significantly impaired. Furthermore, the lack of emotion recognition functionality poses a challenge, making it difficult to optimally process voice according to the user's emotional state.

[0348] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0349] In this invention, the server includes a high-performance voice assistant device, a mobile communication device electrically connected to the voice assistant device, communication means for customizing voice data in the mobile communication device, processing means for performing noise cancellation processing on the voice data, correction means having generative artificial intelligence for correcting the voice data to natural language, display means for converting the voice data into text and displaying it, emotion recognition means for analyzing the voice data and identifying the user's emotion, and optimization means for optimizing other processes based on the emotion information identified by the emotion recognition means, thereby enabling high-quality voice data processing in real time and optimal voice processing according to the user's emotional state.

[0350] A "high-performance audio assist device" is a device that reduces external noise and provides clear audio to the user.

[0351] "Mobile communications device" means a portable electronic device, such as a smartphone or tablet, on which an application for collecting, processing, and displaying voice data is installed.

[0352] "Communication means" refers to a wireless communication interface, such as Bluetooth, for communicating data between the mobile communication device and the audio assistant device.

[0353] "Processing means" refers to the function of performing digital signal processing (DSP) on the audio data collected by the microphone to eliminate noise.

[0354] "Correction means" refers to the function of correcting voice data into natural language using a generative AI model.

[0355] "Display means" refers to the function of converting voice data into text and displaying it on the screen of a smartphone or tablet.

[0356] "Emotion recognition means" refers to an engine that analyzes voice data to identify the user's emotions.

[0357] The "optimization means" refers to a function that optimizes other processes based on the emotion information identified by the emotion recognition means.

[0358] The present invention relates to a system that combines a highly functional voice assistant device, a mobile communication device electrically connected to the device, and a generative AI model and emotion engine. Specifically, the present invention uses the following hardware and software:

[0359] Hardware Configuration

[0360] 1. Audio assistants

[0361] A sophisticated audio assistant is a device that reduces external noise and provides clear audio to users, and can be connected to a mobile communication device via a wireless communication interface such as Bluetooth.

[0362] 2. Mobile Communication Devices

[0363] The mobile communication device is a portable electronic device such as a smartphone or tablet. It is connected to the audio assistant via Bluetooth and has installed an application for collecting, processing, and displaying audio data.

[0364] Software Configuration

[0365] 1. Means of communication

[0366] It is equipped with a wireless communication interface such as Bluetooth for communicating data between the mobile communication device and the audio assistant device.

[0367] 2. Processing means

[0368] This function performs digital signal processing (DSP) on audio data collected by a microphone to eliminate noise.

[0369] 3. Correction means

[0370] This function uses a generative AI model to correct voice data into natural language.

[0371] 4. Display means

[0372] This function converts voice data into text and displays it on the screen of a smartphone or tablet.

[0373] 5. Emotion recognition means

[0374] It has an emotion engine that analyzes voice data to identify the user's emotions.

[0375] 6. Optimization Methods

[0376] This is a function that optimizes other processes based on the emotion information identified by the emotion recognition means.

[0377] Specific examples

[0378] Use at family dinners

[0379] User: Places smartphone in the center of the table, connects audio assistant via Bluetooth, adjusts equalization and noise cancellation settings through the app, and enables emotion engine.

[0380] Device (smartphone): Collects surrounding conversations with a microphone and transmits them to the voice assistant in real time. The voice data is sent to a server to obtain text data, which is then displayed on the screen.

[0381] Emotion Engine: Analyzes voice data and recognizes user emotions. Dynamically adjusts voice filtering parameters based on the recognition.

[0382] Server: Analyzes the voice data, generates text in real time, and sends it back to the smartphone.

[0383] Terminal (smartphone): Saves the text data and sends the correction results to the voice assistant device.

[0384] Prompt Sentence Examples

[0385] An example of a prompt a user can send to a generative AI model is one that includes the instruction, "Please remove the noise from this audio data and correct it into natural language that is easy to listen to."

[0386] The above examples of specific use cases and prompt sentences will help understand the detailed implementation of the system of the present invention.

[0387] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0388] Program processing flow

[0389] Step 1: Bluetooth connection

[0390] The user launches the smartphone app and sets the audio assistant device to Bluetooth pairing mode, which allows the smartphone to recognize the audio assistant device.

[0391] The device (smartphone) searches for Bluetooth devices and displays a list of nearby Bluetooth devices. The user can select an audio assistant device from the list.

[0392] The terminal (smartphone) establishes Bluetooth pairing with the selected device. The input is the user's operation, and the output is the establishment of pairing.

[0393] Step 2: Collect and send audio data

[0394] The device (smartphone) collects ambient sound in real time using a built-in microphone, stores it in a buffer, and then uses digital signal processing (DSP) to reduce noise.

[0395] The terminal (smartphone) transmits the audio data stored in the buffer to the audio assistant device via Bluetooth. The input is the audio data collected by the microphone, and the output is the data in the buffer after DSP processing.

[0396] The audio assistant decodes the received audio data, adjusts the volume and frequency characteristics to suit the audio, and plays it back. The input is the audio data sent from the terminal, and the output is clear audio provided to the user.

[0397] Step 3: Equalizing and noise reduction

[0398] Users can adjust equalization and noise cancellation parameters in the settings screen of the smartphone app, using sliders and presets to customize the sound quality.

[0399] The device (smartphone) updates the filtering algorithm based on the user's settings and applies new parameters. The input is the user's adjusted parameters, and the output is the updated filtering algorithm.

[0400] The terminal (smartphone) transmits the filtered audio data to the audio assistance device via Bluetooth.

[0401] The audio assistant then readjusts the filtered audio data and plays it back optimally, with the input being the filtered audio data and the output being audio optimized for the user.

[0402] Step 4: Transcribe and save the conversation

[0403] The device (smartphone) transmits the voice data of the conversation to the server in real time in streaming format. The input is the voice data collected in real time, and the output is the data transmitted to the server.

[0404] The server analyzes the received voice data and converts it into text data. The input is the voice data sent from the terminal, and the output is text data.

[0405] The device (smartphone) displays the text data returned from the server on the app screen and saves it in its internal storage. The input is the text data from the server, and the output is the displayed and saved text data.

[0406] Step 5: Language correction by generative AI

[0407] The device (smartphone) sends inaudible voice data to the generative AI model. The input is voice data, and the output is data sent to the server.

[0408] The server analyzes the voice data received by the generative AI model and corrects it into natural language. The input is the voice data sent from the device, and the output is the corrected language data.

[0409] The terminal (smartphone) receives the corrected data and plays it on the audio assist device. The input is the corrected data from the server, and the output is the played audio.

[0410] Step 6: Emotion Recognition with the Emotion Engine

[0411] The device (smartphone) transmits voice data collected by a microphone to the emotion engine in real time. The input is the collected voice data, and the output is data sent to the emotion engine.

[0412] The emotion engine analyzes the received voice data and identifies the user's emotion. The input is the voice data, and the output is the identified emotion information.

[0413] The device (smartphone) optimizes the voice data settings based on feedback from the emotion engine. The input is the emotion engine feedback, and the output is the optimized voice data.

[0414] Through the above steps, the system can provide users with comfortable voice interaction.

[0415] (Application example 2)

[0416] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0417] While conventional voice assistance systems can filter voice data and eliminate noise, they lack the ability to recognize the user's emotions and optimize voice data accordingly. This has resulted in issues with not being able to provide appropriate voice feedback in certain situations. Furthermore, they lacked the ability to correct speech to natural language using generative AI and real-time data sharing, making effective security management difficult.

[0418] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server is equipped with an emotion recognition engine and includes means for analyzing voice data in real time, means for correcting voice data using a generation AI, and means for dynamically adjusting filtering parameters. This makes it possible to recognize the user's emotions and optimize voice data based on them. In addition, voice data can be sent to the server in real time and appropriate feedback can be obtained immediately, thereby realizing more effective security management.

[0419] A "high-performance audio assist device" is a device that reduces external noise and provides clear audio.

[0420] A "mobile communication device" is a portable electronic device such as a smartphone or tablet.

[0421] "Communication means" refers to the function of collecting and transmitting voice data and exchanging data between devices.

[0422] The "processing means" is a function that performs digital signal processing such as noise cancellation on audio data.

[0423] The "correction means" is a function that corrects voice data into natural language using generative artificial intelligence.

[0424] The "display means" is a function that converts voice data into text and displays it visually.

[0425] The "emotion recognition means" is a function that recognizes the user's emotions from the voice data and dynamically adjusts the voice filtering parameters accordingly.

[0426] "Means for transmitting voice data in real time" refers to a function for transmitting voice data in real time and for immediate processing and feedback.

[0427] The "equalizing means" is a function that adjusts audio characteristics based on equalizing parameters that can be set by the user.

[0428] The present invention relates to a system that uses a highly functional voice assistant device and a mobile communication device electrically connected to it to effectively process and optimize voice data in a user's environment and apply it to security services. The following describes the specific system configuration and its operation method.

[0429] System configuration

[0430] 1. Hardware

[0431] High-performance audio assist device: Reduces external noise and provides clear audio to users. Connects to mobile communication devices using wireless communication interfaces such as Bluetooth.

[0432] Mobile Communication Device: A portable electronic device, such as a smartphone or tablet, that collects, processes, and displays audio data. It contains a microphone and a Bluetooth module.

[0433] 2. Software

[0434] Bluetooth module: Wirelessly connects the audio assistant device to a smartphone, sending and receiving audio data.

[0435] Digital signal processing (DSP): Performs noise cancellation, equalization, and other processing on collected audio data.

[0436] Emotion engine: Analyzes voice data and recognizes the user's emotions. As a concrete example, it uses a natural language processing library (e.g., Google TensorFlow).

[0437] Generative AI: Generative AI analyzes voice data and corrects it into natural language.

[0438] Display function: Converts voice data into text and displays it visually.

[0439] Operation Overview

[0440] 1. Bluetooth connection

[0441] Server: Pairs the audio assistant device with the mobile communication device via Bluetooth.

[0442] Device: Search for devices, select the audio assistant and establish pairing.

[0443] 2. Collection and transmission of voice data

[0444] Device: Surrounding sounds are collected using the smartphone's microphone and stored in a buffer in real time.

[0445] Terminal: The audio data in the buffer is filtered using digital signal processing and sent to the audio assistant device via Bluetooth.

[0446] Audio assistant: Decodes received audio data and adjusts it to the optimum volume and frequency characteristics.

[0447] 3. Equalization and noise cancellation

[0448] User: Adjust equalization and noise cancellation parameters in the settings screen of the smartphone app.

[0449] Device: Receives user settings, updates the filtering algorithm, and sends it to the audio assistant device.

[0450] 4. Transcribe and save conversations

[0451] Terminal: Sends audio data to the server in real time in streaming format.

[0452] Server: Analyzes the voice data, generates text data, and sends it back to the device.

[0453] Device: Text data is displayed in the app and saved in the internal storage.

[0454] 5. Language Correction by Generative AI

[0455] Device: Sends inaudible voice data to the generating AI.

[0456] Server: The generative AI analyzes the voice data and corrects it into natural language.

[0457] Terminal: Receives the corrected data and plays it on the audio assistant device.

[0458] 6. Emotion Recognition by Emotion Engine

[0459] Terminal: Voice data collected by the microphone is sent to the emotion engine.

[0460] Server: The emotion engine analyzes the voice data and recognizes the user's emotions. This information is reflected in the filtering parameters and generation AI.

[0461] Device: Optimized voice data settings based on feedback from the emotion engine.

[0462] Specific examples

[0463] Office Security Monitoring

[0464] The user acts as a security manager, placing a smartphone in the security room and pairing it with a voice assistant. The system analyzes conversations in the office in real time and can send alerts to managers if it detects tension or stress.

[0465] Prompt Sentence Examples

[0466] "What are the potential emotions that could be present in this scene? Identify these emotions based on the text data. Then update the emotion parameters in the emotion engine and apply the new security settings."

[0467] Thus, the present invention can be implemented in a variety of security environments, and specific hardware and software can be used to achieve effective security management.

[0468] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0469] Step 1:

[0470] Device: Set the audio assistant device to Bluetooth pairing mode, launch the app on the smartphone and search for devices, select the audio assistant device from the search results, and establish pairing. The input is Bluetooth device information, and the output is the status of pairing establishment.

[0471] Step 2:

[0472] Terminal: The smartphone's microphone collects surrounding audio and stores it in a buffer in real time. The input is the surrounding audio data, and the output is the collected audio data in the buffer. Specifically, the smartphone's built-in microphone captures audio, converts the data into digital format, and stores it in the buffer.

[0473] Step 3:

[0474] Terminal: The audio data in the buffer is filtered using a digital signal processing algorithm and sent to the audio assistant via Bluetooth. The input is the audio data in the buffer and the output is the filtered audio data. The audio data is noise-reduced and equalized and sent via the Bluetooth module.

[0475] Step 4:

[0476] Audio assistant: Decodes received audio data and plays it back after adjusting the volume and frequency characteristics to suit the user. The input is filtered audio data, and the output is clear audio. Specifically, the decoding process converts the digital signal into an analog signal, which is then output through a speaker.

[0477] Step 5:

[0478] User: Adjusts equalization and noise cancellation parameters on the settings screen within the smartphone app. The input is the user's setting parameters, and the output is the updated filtering algorithm. The app's UI intuitively manipulates sliders and drop-down menus to change settings.

[0479] Step 6:

[0480] Terminal: Receives user settings, updates the filtering algorithm, and sends the new algorithm to the audio assistant. The input is the updated filtering parameters, and the output is the new filtering algorithm applied to the audio assistant.

[0481] Step 7:

[0482] Terminal: Sends the voice data of the conversation to the server in real time in streaming format. The input is filtered voice data, and the output is streaming data. The data is securely sent to the server via the SSL / TLS protocol.

[0483] Step 8:

[0484] Server: Analyzes the voice data, generates text data, and sends it back to the smartphone. The input is streaming voice data, and the output is the generated text data. The voice recognition engine analyzes the data and converts it into text format.

[0485] Step 9:

[0486] Terminal: Text data is displayed in the app and saved to the internal storage. The input is the text data returned from the server, and the output is the displayed text and saved data. Specifically, it is displayed in the app's text view and saved to the database.

[0487] Step 10:

[0488] Device: Sends inaudible voice data to the generation AI. The input is the collected voice data, and the output is the request data for correction. The API of the generation AI model is called and the voice data is sent.

[0489] Step 11:

[0490] Server: The generative AI analyzes the voice data and corrects it into natural language. The input is the voice data, and the output is the corrected voice data. The generative AI model removes noise and converts it into clear pronunciation.

[0491] Step 12:

[0492] Terminal: Receives the corrected data and plays it back to the audio assistant. The input is the corrected data returned from the server, and the output is the reproduced natural speech.

[0493] Step 13:

[0494] Terminal: Sends voice data collected by a microphone to the emotion engine. The input is the collected voice data, and the output is the emotion recognition request data.

[0495] Step 14:

[0496] Server: The emotion engine analyzes the voice data and recognizes the user's emotions. The input is the voice data, and the output is the recognized emotion information. The emotion data is analyzed through the analytics engine.

[0497] Step 15:

[0498] Terminal: Optimizes voice data settings based on feedback from the emotion engine. The input is the recognized emotion information, and the output is optimized filtering parameters.

[0499] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0500] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0501] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0502] [Second embodiment]

[0503] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0504] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0505] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0506] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0507] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0508] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0509] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0510] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0511] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0512] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0513] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0514] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0515] The system of the present invention is composed of a highly functional voice assistant device and a mobile communication device electrically connected to it. In this system, the mobile communication device (e.g., a smartphone) plays a central role, providing functions such as voice data customization, noise cancellation, correction using generation AI, and transcription.

[0516] System configuration

[0517] 1. Audio assistants

[0518] The advanced audio assistant effectively amplifies surrounding sounds and has filters to reduce noise according to the user's hearing ability, and can be connected to a mobile communication device via a wireless communication interface such as Bluetooth.

[0519] 2. Mobile Communication Devices

[0520] The device is a portable electronic device such as a smartphone or tablet, which connects to the audio assistant via wireless communication such as Bluetooth and has an application installed to collect, process, and display audio data.

[0521] Program processing flow

[0522] The program processing of this system is carried out as follows.

[0523] 1. Bluetooth connection

[0524] User: Launch the smartphone app and set the audio assistant device to Bluetooth pairing mode.

[0525] Device (smartphone): Search for Bluetooth devices, select the audio assistant device, and establish pairing.

[0526] 2. Collection and transmission of voice data

[0527] Device (smartphone): The microphone collects surrounding sounds and stores them in a buffer. The audio data in the buffer is sent in real time to the audio assistant device via wireless communication.

[0528] Audio assistant: Decodes the received audio data and adjusts the volume and frequency characteristics to suit the user before playing it back.

[0529] 3. Equalization and noise cancellation

[0530] User: Adjust equalization and noise cancellation parameters in the settings screen of the smartphone app.

[0531] Device (smartphone): Receives user settings, updates the filtering algorithm, applies a noise cancellation filter to the audio data, and sends it to the audio assistant device.

[0532] Audio assistant: Apply filtered audio data to reduce external noise before playback.

[0533] 4. Transcribe and save conversations

[0534] Terminal (smartphone): Audio data collected by a microphone is sent to a server, where real-time voice recognition is performed.

[0535] Server: Analyzes the voice data, generates text data, and sends it back to the smartphone.

[0536] Device (smartphone): The generated text is displayed on the screen and saved in the internal storage.

[0537] 5. Language Correction by Generative AI

[0538] Device (smartphone): Sends inaudible audio data to the generating AI.

[0539] Server: Generative AI analyzes the voice data and corrects it into natural language.

[0540] Terminal (smartphone): Receives the corrected data and sends it to the audio assistant device for playback.

[0541] Specific examples

[0542] 1. Use at family dinners

[0543] User: Places smartphone in the center of the table and connects to audio assistant via Bluetooth. Adjusts equalization and noise cancellation settings through the app.

[0544] Device (smartphone): The microphone picks up surrounding conversations and transmits them in real time to the voice assistant. The voice data is sent to the server, which then retrieves the text data and displays it on the screen.

[0545] Server: Analyzes the voice data, generates text in real time, and sends it back to the smartphone.

[0546] Terminal (smartphone): Saves the text data and sends the correction results to the voice assistant device.

[0547] Thus, the system of the present invention allows elderly people to participate more naturally in conversations with their families, providing a practical way to compensate for lost communication.

[0548] The processing flow will be explained below.

[0549] Step 1:

[0550] User: Downloads and installs the smartphone app. Opens the app and creates a new account.

[0551] Step 2:

[0552] User: Power on the audio assistant and set it to Bluetooth pairing mode.

[0553] Step 3:

[0554] Device (smartphone): Open the Bluetooth settings screen and search for the audio assist device. Select the audio assist device from the search results and send a pairing request. If pairing is successful, establish a connection with the audio assist device.

[0555] Step 4:

[0556] Device (smartphone): Once a Bluetooth connection with the audio assistant is established, obtain microphone permission from the app and begin collecting audio.

[0557] Step 5:

[0558] Device (smartphone): The smartphone's microphone collects surrounding sounds and stores them in a buffer in real time.

[0559] Step 6:

[0560] Terminal (smartphone): Periodically performs digital signal processing on the audio data in the buffer and transmits it to the audio assistant device via Bluetooth.

[0561] Step 7:

[0562] Audio assistant: Decodes the received audio data and adjusts the volume and frequency characteristics to suit the user before playing it back.

[0563] Step 8:

[0564] User: Adjust equalization and noise cancellation parameters in the settings screen of the smartphone app.

[0565] Step 9:

[0566] Device (smartphone): Receives user setting changes and updates the voice data filtering algorithm. Digital filtering is applied to reduce noise before sending the audio data to the voice assistant device.

[0567] Step 10:

[0568] Audio assistant: Reconditions the received filtered audio data to play the best audio for the user.

[0569] Step 11:

[0570] Device (smartphone): Audio data during conversation is sent to the server in real time in streaming format.

[0571] Step 12:

[0572] Server: The received voice data is analyzed using a voice recognition engine and text data is generated.

[0573] Step 13:

[0574] Server: The generated text data is sent back to the smartphone.

[0575] Step 14:

[0576] Device (smartphone): The generated text data is displayed within the app and saved in the internal storage.

[0577] Step 15:

[0578] User: Tells the app to mark parts of the audio that are difficult to hear.

[0579] Step 16:

[0580] Device (smartphone): Sends the marked audio data to the server and requests correction by the generation AI.

[0581] Step 17:

[0582] Server: The generative AI analyzes the voice data and corrects it into natural language.

[0583] Step 18:

[0584] Server: Sends the corrected audio data back to the smartphone.

[0585] Step 19:

[0586] Terminal (smartphone): Notifies the user of the received correction data and plays it on the audio assistant device.

[0587] In this way, the system of the present invention provides a device that users can easily operate and set up, and not only improves the quality of speech but also has a transcription function that allows users to review past conversations. This system allows users to achieve more natural and comfortable communication.

[0588] Example 1

[0589] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0590] Conventional speech assist devices simply amplify speech, and are not effective enough in noisy environments or when conversations are difficult to hear. Furthermore, they lack the ability to convert speech into text, making it impossible to visually confirm the content of a conversation. This creates challenges for communication, particularly for the elderly and those with hearing impairments.

[0591] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0592] In this invention, the server includes a high-performance audio amplifier, a portable information terminal electrically connected to the audio amplifier, a communication means for customizing audio data on the portable information terminal, a processing means for performing noise reduction processing on the audio data, a correction means having a generative artificial intelligence for correcting the audio data to natural language, a display means for converting the audio data into text and displaying it, a transmission means for storing the audio data in a buffer on the portable information terminal and transmitting the audio data to the audio amplifier in real time via wireless communication, and a setting means for the user to adjust equalization and noise reduction settings on the portable information terminal. This allows for clear audio even in noisy environments and for smooth communication by converting audio data into text and displaying it in real time.

[0593] An "audio amplifier" is a device for audio assistance that has the function of amplifying sound with high precision and reducing noise.

[0594] "Mobile information terminal" is a general term for portable electronic devices such as smartphones and tablets.

[0595] "Communication means" refers to devices or software that have the functionality to collect, send, and receive voice data.

[0596] "Processing means" refers to hardware or software for performing specific processing on audio data.

[0597] "Generative AI" is an AI technology that analyzes voice data and corrects it to make it more natural.

[0598] "Correction means" refers to a device or software that has the function of converting and correcting voice data into natural language.

[0599] "Display means" refers to a device or software that visually presents the text-converted voice data to the user.

[0600] A "buffer" is a storage area for temporarily storing audio data.

[0601] "Transmission means" refers to a device or software that has the function of transmitting audio data to another device via wireless communication.

[0602] "Setting means" refers to a device or software that has a function that allows the user to adjust sound quality and filtering parameters.

[0603] "Equalizing" is a process that emphasizes or attenuates specific frequency bands in audio.

[0604] "Noise reduction" refers to a processing technique for removing unwanted noise from an audio signal.

[0605] "Text conversion" refers to the process of converting voice data into text data.

[0606] "Wireless communication" refers to a communication technology for sending and receiving data without using cables.

[0607] The system of this invention is primarily composed of a high-performance sound amplifier, a mobile information terminal (such as a smartphone or tablet), and a server. This system provides clear audio even in noisy environments, converts audio data into text and displays it in real time, and enables smooth conversations with elderly people and those with hearing impairments.

[0608] Hardware configuration

[0609] Sound amplification equipment:

[0610] It is equipped with high-precision audio amplification and noise reduction filters.

[0611] It can be connected to a mobile information terminal via wireless communication such as Bluetooth.

[0612] Personal digital assistants:

[0613] A smartphone or tablet with a voice assistant application installed.

[0614] Surrounding sounds are picked up through the microphone and stored in a buffer.

[0615] Wireless communication (such as Bluetooth) is used for data communication.

[0616] server:

[0617] It analyzes voice data, transcribes it, and corrects it using AI. It has high-performance processing capabilities and performs secure data communication.

[0618] Software configuration

[0619] Voice-assisted applications:

[0620] Customize audio data, set noise reduction, transcribe audio in real time, and save text.

[0621] It has a settings screen that allows you to adjust equalization and noise reduction parameters.

[0622] Generative AI models:

[0623] It has the ability to analyze voice data and correct it into natural language.

[0624] It runs on the server and returns the correction results to the mobile information terminal.

[0625] Specific examples of operation procedures

[0626] Here, as a concrete example of how the system is used, we will show a scene where the system is used at a family dinner party.

[0627] Example of use at a family dinner:

[0628] User: Places smartphone in the center of the table and connects to sound amplification device via Bluetooth. Adjusts equalization and noise cancellation settings through the app.

[0629] Device (smartphone): The microphone collects the conversations during the dinner party and transmits them in real time to an audio amplifier. The audio data is then sent to a server to obtain text data, which is then displayed on the smartphone screen.

[0630] Server: Analyzes the voice data, generates text in real time, and sends it back to the smartphone. For example, the utterance "Mom, would you like another serving?" is generated as text data.

[0631] Device (smartphone): The generated text data is displayed on the screen so the user can check the content of the conversation. The text data is also saved in the internal storage for later reference. If necessary, language correction is performed on the voice data, and the correction results are sent to the audio amplification device.

[0632] Prompt Sentence Examples

[0633] "How exactly do I use a speech assistive device at a family dinner?"

[0634] This system allows elderly and hearing-impaired people to participate more naturally in conversations with their families and compensate for lost communication.

[0635] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0636] Step 1: Bluetooth connection

[0637] Input: Sound amplification device in pairing mode. User launches smartphone app.

[0638] Operation:

[0639] User: Launch the smartphone app and put the audio amplifier into Bluetooth pairing mode by pressing and holding the pairing button on the audio amplifier.

[0640] Device (smartphone): Start Bluetooth device search and detect the sound amplifier. Select the sound amplifier from the device list and establish pairing. The message "Connection completed" will be displayed.

[0641] Output: "Connected" message upon successful pairing connection.

[0642] Step 2: Collect and send audio data

[0643] Input: Smartphone connected to sound amplification device via Bluetooth.

[0644] Operation:

[0645] Device (smartphone): Uses a microphone to collect surrounding sounds in real time. This data is temporarily stored in a buffer. Consider a dinner party scenario, where the sound of conversations is collected.

[0646] Terminal (smartphone): The audio data stored in the buffer is sent to the audio amplifier via wireless communication (Bluetooth). The data is divided and sent in packet format.

[0647] Audio Amplification Device: Decodes the received audio data and adjusts the volume and frequency characteristics, so that the audio is played back to the user in a clear and crisp manner.

[0648] Output: The audio data with adjusted volume and frequency response reaches the user.

[0649] Step 3: Equalization and noise cancellation settings

[0650] Input: A user is accessing the settings screen within a smartphone app.

[0651] Operation:

[0652] Users: Adjust equalization and noise cancellation parameters in the app's settings, for example, by boosting high frequencies.

[0653] Device (smartphone): Receives user settings and updates the filtering algorithm. When the new settings are reflected in the app, the algorithm processes the audio data accordingly.

[0654] Terminal (smartphone): The filtered audio data is sent back to the audio amplifier. After noise removal, the quality of the audio data is improved.

[0655] Sound amplifier: Reproduces the filtered audio data so that the user hears it with external noise reduced.

[0656] Output: The filtered audio data is delivered to the user.

[0657] Step 4: Transcribe and save the conversation

[0658] Input: Audio data is collected by a smartphone and sent to a server via the internet.

[0659] Operation:

[0660] Device (smartphone): Sends voice data collected by the microphone to the server and starts the voice recognition process.

[0661] Server: The received voice data is converted into text data using a generative AI model. The utterance "Hello, how are you?" is generated as text data.

[0662] Server: Sends the generated text data back to the smartphone.

[0663] Device (smartphone): The text data is displayed on the screen so that the user can check the contents of the conversation, and the text data is saved in the internal storage for later reference.

[0664] Output: The textual audio data is displayed on the screen and saved.

[0665] Step 5: Language correction by generative AI

[0666] Input: Difficult-to-hear audio data collected by a smartphone is sent to the generation AI.

[0667] Operation:

[0668] Device (smartphone): Sends inaudible audio data to the generating AI.

[0669] Server: The generation AI analyzes the voice data and corrects it into natural language. Unclear parts are corrected and clear text data is generated.

[0670] Server: Sends the corrected data back to the smartphone.

[0671] Terminal (smartphone): Receives the corrected data, sends it to an audio amplifier, and plays it back. The corrected audio reaches the user, making it easier to hear.

[0672] Output: The corrected audio data is delivered to the user.

[0673] By using specific actions performed at each step, this system can effectively provide communication support to the elderly and the hearing impaired.

[0674] (Application example 1)

[0675] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0676] Elderly people and those with hearing problems often have difficulty hearing in-store announcements and conversations with store clerks when shopping comfortably in physical stores. This not only degrades the quality of the shopping experience, but can also lead to missing important information. Furthermore, existing voice assistants lack effective noise cancellation and voice correction, which can lead to reduced speech recognition accuracy due to environmental noise.

[0677] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0678] In this invention, the server includes a high-performance voice assistance device, a mobile communication device electrically connected to the voice assistance device, a communication means for customizing voice data on the mobile communication device, a processing means for performing noise cancellation processing on the voice data, a correction means having a generation artificial intelligence for correcting the voice data to natural language, a display means for converting the voice data into text and displaying it, a voice output means for playing back the generated text data, and a server communication means for transmitting the voice data to an external server and retrieving the corrected text data. This allows users who require hearing assistance to easily hear in-store announcements and conversations with store clerks while shopping in a physical store, as well as visually confirm the corrected text data. Furthermore, correcting and playing back the generated text data can provide a natural conversation experience.

[0679] A "high-performance audio assist device" is a device that effectively amplifies surrounding sounds and is equipped with a filter to reduce noise according to the user's hearing ability.

[0680] A "mobile communications device" is a portable electronic device such as a smartphone or tablet.

[0681] The "communication means" is a function including a wireless communication interface for connecting a high-performance voice assistant device and a mobile communication device and customizing voice data.

[0682] "Processing means" refers to algorithms or programs for performing noise cancellation processing on speech data in a mobile communication device.

[0683] "Correction means" refers to a function that uses generative AI to correct voice data into natural language.

[0684] "Display means" refers to a display or screen that converts voice data into text and displays it.

[0685] The "audio output means" refers to a function such as a speaker or earphone for reproducing the generated text data as audio.

[0686] The "server communication means" is a function for transmitting voice data from the mobile communication device to an external server and obtaining corrected text data from the server.

[0687] To implement this invention, the following system configuration and operation method are required. A highly functional audio assistance device is used in combination with a mobile communication device (smartphone or tablet). This enables hearing assistance in a store.

[0688] System Configuration

[0689] 1. Advanced audio assistive devices:

[0690] It effectively amplifies surrounding sounds and has a filter to reduce noise according to the user's hearing ability, and can be connected to a mobile communication device via a wireless communication interface such as Bluetooth.

[0691] 2. Mobile communication devices:

[0692] It is a portable electronic device that connects to the audio assistant via wireless communication such as Bluetooth. This device is equipped with an application that customizes audio data, cancels noise, corrects it using generative AI, transcribes it, and plays it back.

[0693] Hardware and software used

[0694] 1. Hardware:

[0695] Smartphone (iOS or Android)

[0696] High-performance Bluetooth-enabled audio assist device

[0697] Your smartphone's built-in microphone and speaker (or earphones)

[0698] 2. Software:

[0699] Speech recognition libraries (e.g., speech_recognition, Google Speech API)

[0700] Text-to-speech engine (e.g. pyttsx3)

[0701] Server communication library (e.g. requests)

[0702] External generation AI server (API)

[0703] Specific examples

[0704] 1. Scenario 1:

[0705] In a physical store, a user stands in front of a shelf and asks a store clerk where a product is located. The user launches an application on their smartphone and types the question into the microphone.

[0706] Example prompt:

[0707] "Please ask the microphone where the item is."

[0708] 2. Scenario 2:

[0709] The user listens to an announcement about a new campaign in front of the cash register at a physical store. The smartphone is placed in a location where the announcement can be heard, and the speech is converted into text, corrected, and the corrected text is output as voice.

[0710] Example prompt:

[0711] Please check the announcement

[0712] Processing flow

[0713] 1. Device (smartphone):

[0714] The microphone picks up surrounding sounds and transmits the audio data via Bluetooth to a high-performance audio assist device, which simultaneously processes the audio data in real time and transmits it to an external server as needed.

[0715] 2. Server:

[0716] It receives voice data, corrects it into natural language using generative AI, and sends the corrected text data back to the smartphone.

[0717] 3. Device (smartphone):

[0718] The received text data is displayed on a display and output as voice using a text-to-speech engine.

[0719] This allows users who require hearing assistance to easily hear in-store announcements and conversations with store clerks while shopping in a physical store, as well as visually confirm the information. Furthermore, by correcting and playing back the generated text data, a natural conversation experience can be provided.

[0720] The system enables people with hearing problems to more comfortably participate in public activities, reducing social barriers.

[0721] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0722] Step 1:

[0723] User: Launches the smartphone app and sets the audio assistant device to Bluetooth pairing mode. Through the smartphone app interface, the smartphone searches for the audio assistant device via Bluetooth and establishes pairing. Based on this input, the smartphone connects to the audio assistant device via Bluetooth communication. When pairing is successful, the device outputs a notification that the connection is complete.

[0724] Step 2:

[0725] Device (smartphone): The smartphone's microphone collects surrounding sounds in real time and stores them in a buffer. This collected sound data is used as input, and the audio data in the buffer is processed in real time and sent via wireless communication to the audio assistance device. The audio assistance device receives the transmitted audio data, converts it into volume and frequency characteristics adjusted to suit the user's hearing, and plays it back.

[0726] Step 3:

[0727] User: Adjusts equalization and noise cancellation parameters using the settings screen in the smartphone app. Using this setting information as input, the smartphone updates the filtering algorithm and resends the audio data with the noise cancellation filter applied to the audio assistive device. This allows the audio assistive device to reproduce audio with reduced external noise.

[0728] Step 4:

[0729] Terminal (smartphone): The voice data collected by the microphone is sent to the server, which is then requested to perform voice recognition processing. In this process, the server analyzes the voice data as input and generates text data. This generated text data is sent back to the smartphone, and is then displayed on the smartphone screen.

[0730] Step 5:

[0731] Server: The server inputs the voice data sent from the smartphone into a generative AI model and corrects difficult-to-hear parts into natural language. The generative AI corrects the voice data by analyzing it and converting it into appropriate language based on the context. The corrected text data is then sent back from the server to the smartphone.

[0732] Step 6:

[0733] Terminal (smartphone): Receives the corrected text data and displays it on the screen. It also uses a text-to-speech engine to output the text data as voice, allowing the user to confirm the corrected information both visually and audibly.

[0734] Step 7:

[0735] User: In a brick-and-mortar store, the user continues to actively use their smartphone, converting surrounding information into text as needed and receiving natural-language enhanced speech information. For example, if a user asks a store clerk where a product is located, the smartphone microphone collects the question as input and receives enhanced speech that has gone through all the processing steps as output, making the clerk's response easier to understand.

[0736] These processing steps enable users to enjoy sophisticated voice assistance and a natural conversational experience in physical stores.

[0737] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0738] The present invention relates to a system that combines a highly functional voice assistant device with an electrically connected mobile communication device and further includes an emotion engine that recognizes the user's emotions. This system customizes voice data, cancels noise, corrects using generative AI, converts text, and optimizes voice data through emotion recognition.

[0739] System configuration

[0740] 1. Audio assistants

[0741] High-performance audio assistants can reduce external noise and provide clear audio to users, and can connect to mobile communication devices via wireless communication interfaces such as Bluetooth.

[0742] 2. Mobile Communication Devices

[0743] The device is a portable electronic device such as a smartphone or tablet, which is connected to the audio assistant via Bluetooth and has an application installed to collect, process, and display audio data.

[0744] 3. Emotion Engine

[0745] The emotion engine identifies emotions from user speech and input data and uses that information to optimize other processes, such as adjusting filtering parameters for specific emotions in the speech data.

[0746] Program processing flow

[0747] The program processing of this system is carried out as follows.

[0748] 1. Bluetooth connection

[0749] User: Launch the smartphone app and set the audio assistant device to Bluetooth pairing mode.

[0750] Device (smartphone): Search for Bluetooth devices, select the audio assistant device, and establish pairing.

[0751] 2. Collection and transmission of voice data

[0752] Device (smartphone): Collects surrounding audio with a microphone and stores it in a buffer in real time.

[0753] Terminal (smartphone): Digital signal processing is performed on the audio data in the buffer, and the data is sent to the audio assistant device via Bluetooth.

[0754] Audio assistant: Decodes the received audio data and adjusts the volume and frequency characteristics to suit the user before playing it back.

[0755] 3. Equalization and noise cancellation

[0756] User: Adjust equalization and noise cancellation parameters in the settings screen of the smartphone app.

[0757] Device (smartphone): Receives user settings, updates the filtering algorithm, and sends noise-reduced audio data to the audio assistant device.

[0758] Audio assistant: Re-adjusts the filtered audio data for optimal playback.

[0759] 4. Transcribe and save conversations

[0760] Device (smartphone): Audio data during conversation is sent to the server in real time in streaming format.

[0761] Server: Analyzes the voice data, generates text data, and sends it back to the smartphone.

[0762] Device (smartphone): Text data is displayed in the app and saved in the internal storage.

[0763] 5. Language Correction by Generative AI

[0764] Device (smartphone): Sends inaudible audio data to the generating AI.

[0765] Server: Generative AI analyzes the voice data and corrects it into natural language.

[0766] Terminal (smartphone): Receives the corrected data and plays it on the audio assistant device.

[0767] 6. Emotion Recognition by Emotion Engine

[0768] Device (smartphone): Sends voice data collected by the microphone to the emotion engine.

[0769] Emotion engine: Analyzes voice data and recognizes the user's emotions. The recognized emotional information is reflected in filtering parameters and generation AI.

[0770] Device (smartphone): Optimize voice data settings based on feedback from the emotion engine.

[0771] Specific examples

[0772] Use at family dinners

[0773] User: Places smartphone in the center of the table, connects audio assistant via Bluetooth, adjusts equalization and noise cancellation settings through the app, and enables emotion engine.

[0774] Device (smartphone): The microphone collects surrounding conversations and transmits them in real time to the voice assistant. The voice data is sent to the server to obtain text data, which is then displayed on the screen. The emotion engine analyzes the collected voice data, recognizes the user's emotions, and dynamically adjusts voice filtering parameters based on that data.

[0775] Server: Analyzes the voice data, generates text in real time, and sends it back to the smartphone.

[0776] Terminal (smartphone): Saves the text data and sends the correction results to the voice assistant device.

[0777] In this way, the system of the present invention is designed to be easy for users to operate, and the emotion engine optimizes voice data, enabling more effective and natural communication.

[0778] The processing flow will be explained below.

[0779] Step 1:

[0780] User: Downloads and installs the smartphone app. Opens the app and creates a new account.

[0781] Step 2:

[0782] User: Power on the audio assistant and set it to Bluetooth pairing mode.

[0783] Step 3:

[0784] Device (smartphone): Open the Bluetooth settings screen and search for the audio assist device. Select the audio assist device from the search results and send a pairing request. If pairing is successful, establish a connection with the audio assist device.

[0785] Step 4:

[0786] Device (smartphone): Once a Bluetooth connection with the audio assistant is established, obtain microphone permission from the app and begin collecting audio.

[0787] Step 5:

[0788] Device (smartphone): The smartphone's microphone collects surrounding sounds and stores them in a buffer in real time.

[0789] Step 6:

[0790] Terminal (smartphone): Periodically performs digital signal processing on the audio data in the buffer and transmits it to the audio assistant device via Bluetooth.

[0791] Step 7:

[0792] Audio assistant: Decodes the received audio data and adjusts the volume and frequency characteristics to suit the user before playing it back.

[0793] Step 8:

[0794] User: Adjust equalization and noise cancellation parameters in the settings screen of the smartphone app.

[0795] Step 9:

[0796] Device (smartphone): Receives user setting changes and updates the voice data filtering algorithm. Digital filtering is applied to reduce noise before sending the audio data to the voice assistant device.

[0797] Step 10:

[0798] Audio assistant: Reconditions the received filtered audio data to play the best audio for the user.

[0799] Step 11:

[0800] Device (smartphone): Audio data during conversation is sent to the server in real time in streaming format.

[0801] Step 12:

[0802] Server: The received voice data is analyzed using a voice recognition engine and text data is generated.

[0803] Step 13:

[0804] Server: Sends the generated text data back to the smartphone.

[0805] Step 14:

[0806] Device (smartphone): The generated text data is displayed within the app and saved in the internal storage.

[0807] Step 15:

[0808] User: Type into the app to mark parts of the audio that are difficult to hear.

[0809] Step 16:

[0810] Device (smartphone): Sends the marked audio data to the server and requests correction by the generation AI.

[0811] Step 17:

[0812] Server: Generative AI analyzes the voice data and corrects it into natural language.

[0813] Step 18:

[0814] Server: Sends the corrected audio data back to the smartphone.

[0815] Step 19:

[0816] Terminal (smartphone): Notifies the user of the received correction data and plays it on the audio assistant device.

[0817] Step 20:

[0818] Device (smartphone): Sends voice data collected by the microphone to the emotion engine.

[0819] Step 21:

[0820] Emotion engine: Analyzes voice data and recognizes the user's emotions. The recognized emotional information is reflected in filtering parameters and generation AI.

[0821] Step 22:

[0822] Device (smartphone): Optimize voice data settings based on feedback from the emotion engine.

[0823] In this way, the system of the present invention provides a device that users can easily operate and configure, and by optimizing voice data using an emotion engine, it is possible to achieve more effective and natural communication.

[0824] Example 2

[0825] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0826] In recent years, advances in speech recognition and artificial intelligence technologies have improved the convenience of voice-based user interfaces. However, there are problems with degradation of voice data quality due to environmental noise and the user's emotional state. Furthermore, if voice data is not converted to text or corrected in real time, the user experience is significantly impaired. Furthermore, the lack of emotion recognition functionality poses a challenge, making it difficult to optimally process voice according to the user's emotional state.

[0827] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0828] In this invention, the server includes a high-performance voice assistant device, a mobile communication device electrically connected to the voice assistant device, communication means for customizing voice data in the mobile communication device, processing means for performing noise cancellation processing on the voice data, correction means having generative artificial intelligence for correcting the voice data to natural language, display means for converting the voice data into text and displaying it, emotion recognition means for analyzing the voice data and identifying the user's emotion, and optimization means for optimizing other processes based on the emotion information identified by the emotion recognition means, thereby enabling high-quality voice data processing in real time and optimal voice processing according to the user's emotional state.

[0829] A "high-performance audio assist device" is a device that reduces external noise and provides clear audio to the user.

[0830] "Mobile communications device" means a portable electronic device, such as a smartphone or tablet, on which an application for collecting, processing, and displaying voice data is installed.

[0831] "Communication means" refers to a wireless communication interface, such as Bluetooth, for communicating data between the mobile communication device and the audio assistant device.

[0832] "Processing means" refers to the function of performing digital signal processing (DSP) on the audio data collected by the microphone to eliminate noise.

[0833] "Correction means" refers to the function of correcting voice data into natural language using a generative AI model.

[0834] "Display means" refers to the function of converting voice data into text and displaying it on the screen of a smartphone or tablet.

[0835] "Emotion recognition means" refers to an engine that analyzes voice data to identify the user's emotions.

[0836] The "optimization means" refers to a function that optimizes other processes based on the emotion information identified by the emotion recognition means.

[0837] The present invention relates to a system that combines a highly functional voice assistant device, a mobile communication device electrically connected to the device, and a generative AI model and emotion engine. Specifically, the present invention uses the following hardware and software:

[0838] Hardware Configuration

[0839] 1. Audio assistants

[0840] A sophisticated audio assistant is a device that reduces external noise and provides clear audio to users, and can be connected to a mobile communication device via a wireless communication interface such as Bluetooth.

[0841] 2. Mobile Communication Devices

[0842] The mobile communication device is a portable electronic device such as a smartphone or tablet. It is connected to the audio assistant via Bluetooth and has installed an application for collecting, processing, and displaying audio data.

[0843] Software Configuration

[0844] 1. Means of communication

[0845] It is equipped with a wireless communication interface such as Bluetooth for communicating data between the mobile communication device and the audio assistant device.

[0846] 2. Processing means

[0847] This function performs digital signal processing (DSP) on audio data collected by a microphone to eliminate noise.

[0848] 3. Correction means

[0849] This function uses a generative AI model to correct voice data into natural language.

[0850] 4. Display means

[0851] This function converts voice data into text and displays it on the screen of a smartphone or tablet.

[0852] 5. Emotion recognition means

[0853] It has an emotion engine that analyzes voice data to identify the user's emotions.

[0854] 6. Optimization Methods

[0855] This is a function that optimizes other processes based on the emotion information identified by the emotion recognition means.

[0856] Specific examples

[0857] Use at family dinners

[0858] User: Places smartphone in the center of the table, connects audio assistant via Bluetooth, adjusts equalization and noise cancellation settings through the app, and enables emotion engine.

[0859] Device (smartphone): Collects surrounding conversations with a microphone and transmits them to the voice assistant in real time. The voice data is sent to a server to obtain text data, which is then displayed on the screen.

[0860] Emotion Engine: Analyzes voice data and recognizes user emotions. Dynamically adjusts voice filtering parameters based on the recognition.

[0861] Server: Analyzes the voice data, generates text in real time, and sends it back to the smartphone.

[0862] Terminal (smartphone): Saves the text data and sends the correction results to the voice assistant device.

[0863] Prompt Sentence Examples

[0864] An example of a prompt a user can send to a generative AI model is one that includes the instruction, "Please remove the noise from this audio data and correct it into natural language that is easy to listen to."

[0865] The above examples of specific use cases and prompt sentences will help understand the detailed implementation of the system of the present invention.

[0866] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0867] Program processing flow

[0868] Step 1: Bluetooth connection

[0869] The user launches the smartphone app and sets the audio assistant device to Bluetooth pairing mode, which allows the smartphone to recognize the audio assistant device.

[0870] The device (smartphone) searches for Bluetooth devices and displays a list of nearby Bluetooth devices. The user can select an audio assistant device from the list.

[0871] The terminal (smartphone) establishes Bluetooth pairing with the selected device. The input is the user's operation, and the output is the establishment of pairing.

[0872] Step 2: Collect and send audio data

[0873] The device (smartphone) collects ambient sound in real time using a built-in microphone, stores it in a buffer, and then uses digital signal processing (DSP) to reduce noise.

[0874] The terminal (smartphone) transmits the audio data stored in the buffer to the audio assistant device via Bluetooth. The input is the audio data collected by the microphone, and the output is the data in the buffer after DSP processing.

[0875] The audio assistant decodes the received audio data, adjusts the volume and frequency characteristics to suit the audio, and plays it back. The input is the audio data sent from the terminal, and the output is clear audio provided to the user.

[0876] Step 3: Equalizing and noise reduction

[0877] Users can adjust equalization and noise cancellation parameters in the settings screen of the smartphone app, using sliders and presets to customize the sound quality.

[0878] The device (smartphone) updates the filtering algorithm based on the user's settings and applies new parameters. The input is the user's adjusted parameters, and the output is the updated filtering algorithm.

[0879] The terminal (smartphone) transmits the filtered audio data to the audio assistance device via Bluetooth.

[0880] The audio assistant then readjusts the filtered audio data and plays it back optimally, with the input being the filtered audio data and the output being audio optimized for the user.

[0881] Step 4: Transcribe and save the conversation

[0882] The device (smartphone) transmits the voice data of the conversation to the server in real time in streaming format. The input is the voice data collected in real time, and the output is the data transmitted to the server.

[0883] The server analyzes the received voice data and converts it into text data. The input is the voice data sent from the terminal, and the output is text data.

[0884] The device (smartphone) displays the text data returned from the server on the app screen and saves it in its internal storage. The input is the text data from the server, and the output is the displayed and saved text data.

[0885] Step 5: Language correction by generative AI

[0886] The device (smartphone) sends inaudible voice data to the generative AI model. The input is voice data, and the output is data sent to the server.

[0887] The server analyzes the voice data received by the generative AI model and corrects it into natural language. The input is the voice data sent from the device, and the output is the corrected language data.

[0888] The terminal (smartphone) receives the corrected data and plays it on the audio assist device. The input is the corrected data from the server, and the output is the played audio.

[0889] Step 6: Emotion Recognition with the Emotion Engine

[0890] The device (smartphone) transmits voice data collected by a microphone to the emotion engine in real time. The input is the collected voice data, and the output is data sent to the emotion engine.

[0891] The emotion engine analyzes the received voice data and identifies the user's emotion. The input is the voice data, and the output is the identified emotion information.

[0892] The device (smartphone) optimizes the voice data settings based on feedback from the emotion engine. The input is the emotion engine feedback, and the output is the optimized voice data.

[0893] Through the above steps, the system can provide users with comfortable voice interaction.

[0894] (Application example 2)

[0895] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0896] While conventional voice assistance systems can filter voice data and eliminate noise, they lack the ability to recognize the user's emotions and optimize voice data accordingly. This has resulted in issues with not being able to provide appropriate voice feedback in certain situations. Furthermore, they lacked the ability to correct speech to natural language using generative AI and real-time data sharing, making effective security management difficult.

[0897] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server is equipped with an emotion recognition engine and includes means for analyzing voice data in real time, means for correcting voice data using a generation AI, and means for dynamically adjusting filtering parameters. This makes it possible to recognize the user's emotions and optimize voice data based on them. In addition, voice data can be sent to the server in real time, allowing appropriate feedback to be obtained immediately, thereby achieving more effective security management.

[0898] A "high-performance audio assist device" is a device that reduces external noise and provides clear audio.

[0899] A "mobile communication device" is a portable electronic device such as a smartphone or tablet.

[0900] "Communication means" refers to the function of collecting and transmitting voice data and exchanging data between devices.

[0901] The "processing means" is a function that performs digital signal processing such as noise cancellation on audio data.

[0902] The "correction means" is a function that corrects voice data into natural language using generative artificial intelligence.

[0903] The "display means" is a function that converts voice data into text and displays it visually.

[0904] The "emotion recognition means" is a function that recognizes the user's emotions from the voice data and dynamically adjusts the voice filtering parameters accordingly.

[0905] "Means for transmitting voice data in real time" refers to a function for transmitting voice data in real time and for immediate processing and feedback.

[0906] The "equalizing means" is a function that adjusts audio characteristics based on equalizing parameters that can be set by the user.

[0907] The present invention relates to a system that uses a highly functional voice assistant device and a mobile communication device electrically connected to it to effectively process and optimize voice data in a user's environment and apply it to security services. The following describes the specific system configuration and its operation method.

[0908] System configuration

[0909] 1. Hardware

[0910] High-performance audio assist device: Reduces external noise and provides clear audio to users. Connects to mobile communication devices using wireless communication interfaces such as Bluetooth.

[0911] Mobile Communication Device: A portable electronic device, such as a smartphone or tablet, that collects, processes, and displays audio data. It contains a microphone and a Bluetooth module.

[0912] 2. Software

[0913] Bluetooth module: Wirelessly connects the audio assistant device to a smartphone, sending and receiving audio data.

[0914] Digital signal processing (DSP): Performs noise cancellation, equalization, and other processing on collected audio data.

[0915] Emotion engine: Analyzes voice data and recognizes the user's emotions. As a concrete example, it uses a natural language processing library (e.g., Google TensorFlow).

[0916] Generative AI: Generative AI analyzes voice data and corrects it into natural language.

[0917] Display function: Converts voice data into text and displays it visually.

[0918] Operation Overview

[0919] 1. Bluetooth connection

[0920] Server: Pairs the audio assistant device with the mobile communication device via Bluetooth.

[0921] Device: Search for devices, select the audio assistant and establish pairing.

[0922] 2. Collection and transmission of voice data

[0923] Device: Surrounding sounds are collected using the smartphone's microphone and stored in a buffer in real time.

[0924] Terminal: The audio data in the buffer is filtered using digital signal processing and sent to the audio assistant device via Bluetooth.

[0925] Audio assistant: Decodes received audio data and adjusts it to the optimum volume and frequency characteristics.

[0926] 3. Equalization and noise cancellation

[0927] User: Adjust equalization and noise cancellation parameters in the settings screen of the smartphone app.

[0928] Device: Receives user settings, updates the filtering algorithm, and sends it to the audio assistant device.

[0929] 4. Transcribe and save conversations

[0930] Terminal: Sends audio data to the server in real time in streaming format.

[0931] Server: Analyzes the voice data, generates text data, and sends it back to the device.

[0932] Device: Text data is displayed in the app and saved in the internal storage.

[0933] 5. Language Correction by Generative AI

[0934] Device: Sends inaudible voice data to the generating AI.

[0935] Server: The generative AI analyzes the voice data and corrects it into natural language.

[0936] Terminal: Receives the corrected data and plays it on the audio assistant device.

[0937] 6. Emotion Recognition by Emotion Engine

[0938] Terminal: Voice data collected by the microphone is sent to the emotion engine.

[0939] Server: The emotion engine analyzes the voice data and recognizes the user's emotions. This information is reflected in the filtering parameters and generation AI.

[0940] Device: Optimized voice data settings based on feedback from the emotion engine.

[0941] Specific examples

[0942] Office Security Monitoring

[0943] The user acts as a security manager, placing a smartphone in the security room and pairing it with a voice assistant. The system analyzes conversations in the office in real time and can send alerts to managers if it detects tension or stress.

[0944] Prompt Sentence Examples

[0945] "What are the potential emotions that could be present in this scene? Identify these emotions based on the text data. Then update the emotion parameters in the emotion engine and apply the new security settings."

[0946] Thus, the present invention can be implemented in a variety of security environments, and specific hardware and software can be used to achieve effective security management.

[0947] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0948] Step 1:

[0949] Device: Set the audio assistant device to Bluetooth pairing mode, launch the app on the smartphone and search for devices, select the audio assistant device from the search results, and establish pairing. The input is Bluetooth device information, and the output is the status of pairing establishment.

[0950] Step 2:

[0951] Terminal: The smartphone's microphone collects surrounding audio and stores it in a buffer in real time. The input is the surrounding audio data, and the output is the collected audio data in the buffer. Specifically, the smartphone's built-in microphone captures audio, converts the data into digital format, and stores it in the buffer.

[0952] Step 3:

[0953] Terminal: The audio data in the buffer is filtered using a digital signal processing algorithm and sent to the audio assistant via Bluetooth. The input is the audio data in the buffer and the output is the filtered audio data. The audio data is noise-reduced and equalized and sent via the Bluetooth module.

[0954] Step 4:

[0955] Audio assistant: Decodes received audio data and plays it back after adjusting the volume and frequency characteristics to suit the user. The input is filtered audio data, and the output is clear audio. Specifically, the decoding process converts the digital signal into an analog signal, which is then output through a speaker.

[0956] Step 5:

[0957] User: Adjusts equalization and noise cancellation parameters on the settings screen within the smartphone app. The input is the user's setting parameters, and the output is the updated filtering algorithm. The app's UI intuitively manipulates sliders and drop-down menus to change settings.

[0958] Step 6:

[0959] Terminal: Receives user settings, updates the filtering algorithm, and sends the new algorithm to the audio assistant. The input is the updated filtering parameters, and the output is the new filtering algorithm applied to the audio assistant.

[0960] Step 7:

[0961] Terminal: Sends the voice data of the conversation to the server in real time in streaming format. The input is filtered voice data, and the output is streaming data. The data is securely sent to the server via the SSL / TLS protocol.

[0962] Step 8:

[0963] Server: Analyzes the voice data, generates text data, and sends it back to the smartphone. The input is streaming voice data, and the output is the generated text data. The voice recognition engine analyzes the data and converts it into text format.

[0964] Step 9:

[0965] Terminal: Text data is displayed in the app and saved to the internal storage. The input is the text data returned from the server, and the output is the displayed text and saved data. Specifically, it is displayed in the app's text view and saved to the database.

[0966] Step 10:

[0967] Device: Sends inaudible voice data to the generation AI. The input is the collected voice data, and the output is the request data for correction. The API of the generation AI model is called and the voice data is sent.

[0968] Step 11:

[0969] Server: The generative AI analyzes the voice data and corrects it into natural language. The input is the voice data, and the output is the corrected voice data. The generative AI model removes noise and converts it into clear pronunciation.

[0970] Step 12:

[0971] Terminal: Receives the corrected data and plays it back to the audio assistant. The input is the corrected data returned from the server, and the output is the reproduced natural speech.

[0972] Step 13:

[0973] Terminal: Sends voice data collected by a microphone to the emotion engine. The input is the collected voice data, and the output is the emotion recognition request data.

[0974] Step 14:

[0975] Server: The emotion engine analyzes the voice data and recognizes the user's emotions. The input is the voice data, and the output is the recognized emotion information. The emotion data is analyzed through the analytics engine.

[0976] Step 15:

[0977] Terminal: Optimizes voice data settings based on feedback from the emotion engine. The input is the recognized emotion information, and the output is optimized filtering parameters.

[0978] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0979] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0980] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0981] [Third embodiment]

[0982] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0983] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[0984] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0985] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0986] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0987] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0988] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0989] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0990] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0991] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0992] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0993] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0994] The system of the present invention is composed of a highly functional voice assistant device and a mobile communication device electrically connected to it. In this system, the mobile communication device (e.g., a smartphone) plays a central role, providing functions such as voice data customization, noise cancellation, correction using generation AI, and transcription.

[0995] System configuration

[0996] 1. Audio assistants

[0997] The advanced audio assist device effectively amplifies surrounding sounds and has filters to reduce noise according to the user's hearing ability, and can be connected to a mobile communication device via a wireless communication interface such as Bluetooth.

[0998] 2. Mobile Communication Devices

[0999] The device is a portable electronic device such as a smartphone or tablet. It is connected to the audio assistant via wireless communication such as Bluetooth and has an application installed to collect, process, and display audio data.

[1000] Program processing flow

[1001] The program processing of this system is carried out as follows.

[1002] 1. Bluetooth connection

[1003] User: Launch the smartphone app and set the audio assistant device to Bluetooth pairing mode.

[1004] Device (smartphone): Search for Bluetooth devices, select the audio assistant device, and establish pairing.

[1005] 2. Collection and transmission of voice data

[1006] Device (smartphone): The microphone collects surrounding sounds and stores them in a buffer. The audio data in the buffer is then transmitted in real time to the audio assistant via wireless communication.

[1007] Audio assistant: Decodes the received audio data and adjusts the volume and frequency characteristics to suit the user before playing it back.

[1008] 3. Equalization and noise cancellation

[1009] User: Adjust equalization and noise cancellation parameters in the settings screen of the smartphone app.

[1010] Device (smartphone): Receives user settings, updates the filtering algorithm, applies a noise cancellation filter to the audio data, and sends it to the audio assistant device.

[1011] Audio assistant: Apply filtered audio data to reduce external noise before playback.

[1012] 4. Transcribe and save conversations

[1013] Terminal (smartphone): Audio data collected by a microphone is sent to a server, where real-time voice recognition is performed.

[1014] Server: Analyzes the voice data, generates text data, and sends it back to the smartphone.

[1015] Device (smartphone): The generated text is displayed on the screen and saved in the internal storage.

[1016] 5. Language Correction by Generative AI

[1017] Device (smartphone): Sends inaudible audio data to the generating AI.

[1018] Server: Generative AI analyzes the voice data and corrects it into natural language.

[1019] Terminal (smartphone): Receives the corrected data and sends it to the audio assistant device for playback.

[1020] Specific examples

[1021] 1. Use at family dinners

[1022] User: Places smartphone in the center of the table and connects to audio assistant via Bluetooth. Adjusts equalization and noise cancellation settings through the app.

[1023] Device (smartphone): The microphone picks up surrounding conversations and transmits them in real time to the voice assistant. The voice data is sent to the server, which then retrieves the text data and displays it on the screen.

[1024] Server: Analyzes the voice data, generates text in real time, and sends it back to the smartphone.

[1025] Terminal (smartphone): Saves the text data and sends the correction results to the voice assistant device.

[1026] Thus, the system of the present invention allows elderly people to participate more naturally in conversations with their families, providing a practical way to compensate for lost communication.

[1027] The processing flow will be explained below.

[1028] Step 1:

[1029] User: Downloads and installs the smartphone app. Opens the app and creates a new account.

[1030] Step 2:

[1031] User: Power on the audio assistant and set it to Bluetooth pairing mode.

[1032] Step 3:

[1033] Device (smartphone): Open the Bluetooth settings screen and search for the audio assist device. Select the audio assist device from the search results and send a pairing request. If pairing is successful, establish a connection with the audio assist device.

[1034] Step 4:

[1035] Device (smartphone): Once a Bluetooth connection with the audio assistant is established, obtain microphone permission from the app and begin collecting audio.

[1036] Step 5:

[1037] Device (smartphone): The smartphone's microphone collects surrounding sounds and stores them in a buffer in real time.

[1038] Step 6:

[1039] Terminal (smartphone): Periodically performs digital signal processing on the audio data in the buffer and transmits it to the audio assistant device via Bluetooth.

[1040] Step 7:

[1041] Audio assistant: Decodes the received audio data and adjusts the volume and frequency characteristics to suit the user before playing it back.

[1042] Step 8:

[1043] User: Adjust equalization and noise cancellation parameters in the settings screen of the smartphone app.

[1044] Step 9:

[1045] Device (smartphone): Receives user setting changes and updates the voice data filtering algorithm. Digital filtering is applied to reduce noise before sending the audio data to the voice assistant device.

[1046] Step 10:

[1047] Audio assistant: Reconditions the received filtered audio data to play the best audio for the user.

[1048] Step 11:

[1049] Device (smartphone): Audio data during conversation is sent to the server in real time in streaming format.

[1050] Step 12:

[1051] Server: The received voice data is analyzed using a voice recognition engine and text data is generated.

[1052] Step 13:

[1053] Server: Sends the generated text data back to the smartphone.

[1054] Step 14:

[1055] Device (smartphone): The generated text data is displayed within the app and saved in the internal storage.

[1056] Step 15:

[1057] User: Tells the app to mark parts of the audio that are difficult to hear.

[1058] Step 16:

[1059] Device (smartphone): Sends the marked audio data to the server and requests correction by the generation AI.

[1060] Step 17:

[1061] Server: Generative AI analyzes the voice data and corrects it into natural language.

[1062] Step 18:

[1063] Server: Sends the corrected audio data back to the smartphone.

[1064] Step 19:

[1065] Terminal (smartphone): Notifies the user of the received correction data and plays it on the audio assistant device.

[1066] In this way, the system of the present invention provides a device that users can easily operate and set up, and not only improves the quality of speech but also has a transcription function that allows users to review past conversations. This system allows users to achieve more natural and comfortable communication.

[1067] Example 1

[1068] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1069] Conventional speech assist devices simply amplify speech, and are not effective enough in noisy environments or when conversations are difficult to hear. Furthermore, they lack the ability to convert speech into text, making it impossible to visually confirm the content of a conversation. This creates challenges for communication, particularly for the elderly and those with hearing impairments.

[1070] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1071] In this invention, the server includes a high-performance audio amplifier, a portable information terminal electrically connected to the audio amplifier, a communication means for customizing audio data on the portable information terminal, a processing means for performing noise reduction processing on the audio data, a correction means having a generative artificial intelligence for correcting the audio data to natural language, a display means for converting the audio data into text and displaying it, a transmission means for storing the audio data in a buffer on the portable information terminal and transmitting the audio data to the audio amplifier in real time via wireless communication, and a setting means for the user to adjust equalization and noise reduction settings on the portable information terminal. This allows for clear audio even in noisy environments and for smooth communication by converting audio data into text and displaying it in real time.

[1072] An "audio amplifier" is a device for audio assistance that has the function of amplifying sound with high precision and reducing noise.

[1073] "Mobile information terminal" is a general term for portable electronic devices such as smartphones and tablets.

[1074] "Communication means" refers to devices or software that have the functionality to collect, send, and receive voice data.

[1075] "Processing means" refers to hardware or software for performing specific processing on audio data.

[1076] "Generative AI" is an AI technology that analyzes voice data and corrects it to make it more natural.

[1077] "Correction means" refers to a device or software that has the function of converting and correcting voice data into natural language.

[1078] "Display means" refers to a device or software that visually presents the text-converted voice data to the user.

[1079] A "buffer" is a storage area for temporarily storing audio data.

[1080] "Transmission means" refers to a device or software that has the function of transmitting audio data to another device via wireless communication.

[1081] "Setting means" refers to a device or software that has a function that allows the user to adjust sound quality and filtering parameters.

[1082] "Equalizing" is a process that emphasizes or attenuates specific frequency bands in audio.

[1083] "Noise reduction" refers to a processing technique for removing unwanted noise from an audio signal.

[1084] "Text conversion" refers to the process of converting voice data into text data.

[1085] "Wireless communication" refers to a communication technology for sending and receiving data without using cables.

[1086] The system of this invention is primarily composed of a high-performance sound amplifier, a mobile information terminal (such as a smartphone or tablet), and a server. This system provides clear audio even in noisy environments, converts audio data into text and displays it in real time, and enables smooth conversations with elderly people and those with hearing impairments.

[1087] Hardware configuration

[1088] Sound amplification equipment:

[1089] It is equipped with high-precision audio amplification and noise reduction filters.

[1090] It can be connected to a mobile information terminal via wireless communication such as Bluetooth.

[1091] Mobile devices:

[1092] A smartphone or tablet with a voice assistant application installed.

[1093] Surrounding sounds are picked up through the microphone and stored in a buffer.

[1094] Wireless communication (such as Bluetooth) is used for data communication.

[1095] server:

[1096] It analyzes voice data, transcribes it, and corrects it using AI. It has high-performance processing capabilities and performs secure data communication.

[1097] Software configuration

[1098] Voice-assisted applications:

[1099] Customize audio data, set noise reduction, transcribe audio in real time, and save text.

[1100] It has a settings screen that allows you to adjust equalization and noise reduction parameters.

[1101] Generative AI models:

[1102] It has the ability to analyze voice data and correct it into natural language.

[1103] It runs on the server and returns the correction results to the mobile information terminal.

[1104] Specific examples of operation procedures

[1105] Here, as a concrete example of how the system is used, we will show a scene where the system is used at a family dinner party.

[1106] Example of use at a family dinner:

[1107] User: Places smartphone in the center of the table and connects to sound amplification device via Bluetooth. Adjusts equalization and noise cancellation settings through the app.

[1108] Device (smartphone): The microphone collects the conversations during the dinner party and transmits them in real time to an audio amplifier. The audio data is then sent to a server to obtain text data, which is then displayed on the smartphone screen.

[1109] Server: Analyzes the voice data, generates text in real time, and sends it back to the smartphone. For example, the utterance "Mom, would you like another serving?" is generated as text data.

[1110] Device (smartphone): The generated text data is displayed on the screen so the user can check the content of the conversation. The text data is also saved in the internal storage for later reference. If necessary, language correction is performed on the voice data, and the correction results are sent to the audio amplification device.

[1111] Prompt Sentence Examples

[1112] "How exactly do I use a speech assistive device at a family dinner?"

[1113] This system allows elderly and hearing-impaired people to participate more naturally in conversations with their families and compensate for lost communication.

[1114] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1115] Step 1: Bluetooth connection

[1116] Input: Sound amplification device in pairing mode. User launches smartphone app.

[1117] Operation:

[1118] User: Launch the smartphone app and put the audio amplifier into Bluetooth pairing mode by pressing and holding the pairing button on the audio amplifier.

[1119] Device (smartphone): Start Bluetooth device search and detect the sound amplifier. Select the sound amplifier from the device list and establish pairing. The message "Connection completed" will be displayed.

[1120] Output: "Connected" message upon successful pairing connection.

[1121] Step 2: Collect and send audio data

[1122] Input: Smartphone connected to sound amplification device via Bluetooth.

[1123] Operation:

[1124] Device (smartphone): Uses a microphone to collect surrounding sounds in real time. This data is temporarily stored in a buffer. Consider a dinner party scenario, where the sound of conversations is collected.

[1125] Terminal (smartphone): The audio data stored in the buffer is sent to the audio amplifier via wireless communication (Bluetooth). The data is divided and sent in packet format.

[1126] Audio Amplification Device: Decodes the received audio data and adjusts the volume and frequency characteristics, so that the audio is played back to the user in a clear and crisp manner.

[1127] Output: The audio data with adjusted volume and frequency response reaches the user.

[1128] Step 3: Equalization and noise cancellation settings

[1129] Input: A user is accessing the settings screen within a smartphone app.

[1130] Operation:

[1131] Users: Adjust equalization and noise cancellation parameters in the app's settings, for example, by boosting high frequencies.

[1132] Device (smartphone): Receives user settings and updates the filtering algorithm. When the new settings are reflected in the app, the algorithm processes the audio data accordingly.

[1133] Terminal (smartphone): The filtered audio data is sent back to the audio amplifier. After noise removal, the quality of the audio data is improved.

[1134] Sound amplifier: Reproduces the filtered audio data so that the user hears it with external noise reduced.

[1135] Output: The filtered audio data is delivered to the user.

[1136] Step 4: Transcribe and save the conversation

[1137] Input: Audio data is collected by a smartphone and sent to a server via the internet.

[1138] Operation:

[1139] Device (smartphone): Sends voice data collected by the microphone to the server and starts the voice recognition process.

[1140] Server: The received voice data is converted into text data using a generative AI model. The utterance "Hello, how are you?" is generated as text data.

[1141] Server: Sends the generated text data back to the smartphone.

[1142] Device (smartphone): The text data is displayed on the screen so that the user can check the contents of the conversation, and the text data is saved in the internal storage for later reference.

[1143] Output: The textual audio data is displayed on the screen and saved.

[1144] Step 5: Language correction by generative AI

[1145] Input: Difficult-to-hear audio data collected by a smartphone is sent to the generation AI.

[1146] Operation:

[1147] Device (smartphone): Sends inaudible audio data to the generating AI.

[1148] Server: The generation AI analyzes the voice data and corrects it into natural language. Unclear parts are corrected and clear text data is generated.

[1149] Server: Sends the corrected data back to the smartphone.

[1150] Terminal (smartphone): Receives the corrected data, sends it to an audio amplifier, and plays it back. The corrected audio reaches the user, making it easier to hear.

[1151] Output: The corrected audio data is delivered to the user.

[1152] By using specific actions performed at each step, this system can effectively provide communication support to the elderly and the hearing impaired.

[1153] (Application example 1)

[1154] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1155] Elderly people and those with hearing problems often have difficulty hearing in-store announcements and conversations with store clerks when shopping comfortably in physical stores. This not only degrades the quality of the shopping experience, but can also lead to missing important information. Furthermore, existing voice assistants lack effective noise cancellation and voice correction, which can lead to reduced speech recognition accuracy due to environmental noise.

[1156] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1157] In this invention, the server includes a high-performance voice assistance device, a mobile communication device electrically connected to the voice assistance device, a communication means for customizing voice data on the mobile communication device, a processing means for performing noise cancellation processing on the voice data, a correction means having a generation artificial intelligence for correcting the voice data to natural language, a display means for converting the voice data into text and displaying it, a voice output means for playing back the generated text data, and a server communication means for transmitting the voice data to an external server and retrieving the corrected text data. This allows users who require hearing assistance to easily hear in-store announcements and conversations with store clerks while shopping in a physical store, as well as visually confirm the corrected text data. Furthermore, correcting and playing back the generated text data can provide a natural conversation experience.

[1158] A "high-performance audio assist device" is a device that effectively amplifies surrounding sounds and is equipped with a filter to reduce noise according to the user's hearing ability.

[1159] A "mobile communications device" is a portable electronic device such as a smartphone or tablet.

[1160] The "communication means" is a function including a wireless communication interface for connecting a high-performance voice assistant device and a mobile communication device and customizing voice data.

[1161] "Processing means" refers to algorithms or programs for performing noise cancellation processing on speech data in a mobile communication device.

[1162] "Correction means" refers to a function that uses generative AI to correct voice data into natural language.

[1163] "Display means" refers to a display or screen that converts voice data into text and displays it.

[1164] The "audio output means" refers to a function such as a speaker or earphone for reproducing the generated text data as audio.

[1165] The "server communication means" is a function for transmitting voice data from the mobile communication device to an external server and obtaining corrected text data from the server.

[1166] To implement this invention, the following system configuration and operation method are required. A highly functional audio assistance device is used in combination with a mobile communication device (smartphone or tablet). This enables hearing assistance in a store.

[1167] System Configuration

[1168] 1. Advanced audio assistive devices:

[1169] It effectively amplifies surrounding sounds and has a filter to reduce noise according to the user's hearing ability, and can be connected to a mobile communication device via a wireless communication interface such as Bluetooth.

[1170] 2. Mobile communication devices:

[1171] It is a portable electronic device that connects to the audio assistant via wireless communication such as Bluetooth. This device is equipped with an application that customizes audio data, cancels noise, corrects it using generative AI, transcribes it, and plays it back.

[1172] Hardware and software used

[1173] 1. Hardware:

[1174] Smartphone (iOS or Android)

[1175] High-performance Bluetooth-enabled audio assist device

[1176] Your smartphone's built-in microphone and speaker (or earphones)

[1177] 2. Software:

[1178] Speech recognition libraries (e.g., speech_recognition, Google Speech API)

[1179] Text-to-speech engine (e.g. pyttsx3)

[1180] Server communication library (e.g. requests)

[1181] External generation AI server (API)

[1182] Specific examples

[1183] 1. Scenario 1:

[1184] In a physical store, a user stands in front of a shelf and asks a store clerk where a product is located. The user launches an application on their smartphone and types the question into the microphone.

[1185] Example prompt:

[1186] "Please ask the microphone where the item is."

[1187] 2. Scenario 2:

[1188] The user listens to an announcement about a new campaign in front of the cash register at a physical store. The smartphone is placed in a location where the announcement can be heard, and the speech is converted into text, corrected, and the corrected text is output as voice.

[1189] Example prompt:

[1190] Please check the announcement

[1191] Processing flow

[1192] 1. Device (smartphone):

[1193] The microphone picks up surrounding sounds and transmits the audio data via Bluetooth to a high-performance audio assist device, which simultaneously processes the audio data in real time and transmits it to an external server as needed.

[1194] 2. Server:

[1195] It receives voice data, corrects it into natural language using generative AI, and sends the corrected text data back to the smartphone.

[1196] 3. Device (smartphone):

[1197] The received text data is displayed on a display and output as voice using a text-to-speech engine.

[1198] This allows users who require hearing assistance to easily hear in-store announcements and conversations with store clerks while shopping in a physical store, as well as visually confirm the information. Furthermore, by correcting and playing back the generated text data, a natural conversation experience can be provided.

[1199] The system enables people with hearing problems to more comfortably participate in public activities, reducing social barriers.

[1200] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1201] Step 1:

[1202] User: Launches the smartphone app and sets the audio assistant device to Bluetooth pairing mode. Through the smartphone app interface, the smartphone searches for the audio assistant device via Bluetooth and establishes pairing. Based on this input, the smartphone connects to the audio assistant device via Bluetooth communication. When pairing is successful, the device outputs a notification that the connection is complete.

[1203] Step 2:

[1204] Device (smartphone): The smartphone's microphone collects surrounding sounds in real time and stores them in a buffer. This collected sound data is used as input, and the audio data in the buffer is processed in real time and sent via wireless communication to the audio assistance device. The audio assistance device receives the transmitted audio data, converts it into volume and frequency characteristics adjusted to suit the user's hearing, and plays it back.

[1205] Step 3:

[1206] User: Adjusts equalization and noise cancellation parameters using the settings screen in the smartphone app. Using this setting information as input, the smartphone updates the filtering algorithm and resends the audio data with the noise cancellation filter applied to the audio assistive device. This allows the audio assistive device to reproduce audio with reduced external noise.

[1207] Step 4:

[1208] Terminal (smartphone): The voice data collected by the microphone is sent to the server, which is then requested to perform voice recognition processing. In this process, the server analyzes the voice data as input and generates text data. This generated text data is sent back to the smartphone, and is then displayed on the smartphone screen.

[1209] Step 5:

[1210] Server: The server inputs the voice data sent from the smartphone into a generative AI model and corrects difficult-to-hear parts into natural language. The generative AI corrects the voice data by analyzing it and converting it into appropriate language based on the context. The corrected text data is then sent back from the server to the smartphone.

[1211] Step 6:

[1212] Terminal (smartphone): Receives the corrected text data and displays it on the screen. It also uses a text-to-speech engine to output the text data as voice, allowing the user to confirm the corrected information both visually and audibly.

[1213] Step 7:

[1214] User: In a brick-and-mortar store, the user continues to actively use their smartphone, converting surrounding information into text as needed and receiving natural-language enhanced speech information. For example, if a user asks a store clerk where a product is located, the smartphone microphone collects the question as input and receives enhanced speech that has gone through all the processing steps as output, making the clerk's response easier to understand.

[1215] These processing steps enable users to enjoy sophisticated voice assistance and a natural conversational experience in physical stores.

[1216] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1217] The present invention relates to a system that combines a highly functional voice assistant device with an electrically connected mobile communication device and further includes an emotion engine that recognizes the user's emotions. This system customizes voice data, cancels noise, corrects using generative AI, converts text, and optimizes voice data through emotion recognition.

[1218] System configuration

[1219] 1. Audio assistants

[1220] High-performance audio assistants reduce external noise and provide clear audio to users, and can be connected to mobile communication devices via wireless communication interfaces such as Bluetooth.

[1221] 2. Mobile Communication Devices

[1222] The device is a portable electronic device such as a smartphone or tablet, which is connected to the audio assistant via Bluetooth and has an application installed to collect, process, and display audio data.

[1223] 3. Emotion Engine

[1224] The emotion engine identifies emotions from user speech and input data and uses that information to optimize other processes, such as adjusting filtering parameters for specific emotions in the speech data.

[1225] Program processing flow

[1226] The program processing of this system is carried out as follows.

[1227] 1. Bluetooth connection

[1228] User: Launch the smartphone app and set the audio assistant device to Bluetooth pairing mode.

[1229] Device (smartphone): Search for Bluetooth devices, select the audio assistant device, and establish pairing.

[1230] 2. Collection and transmission of voice data

[1231] Device (smartphone): Collects surrounding audio with a microphone and stores it in a buffer in real time.

[1232] Terminal (smartphone): Digital signal processing is performed on the audio data in the buffer, and the data is sent to the audio assistant device via Bluetooth.

[1233] Audio assistant: Decodes the received audio data and adjusts the volume and frequency characteristics to suit the user before playing it back.

[1234] 3. Equalization and noise cancellation

[1235] User: Adjust equalization and noise cancellation parameters in the settings screen of the smartphone app.

[1236] Device (smartphone): Receives user settings, updates the filtering algorithm, and sends noise-reduced audio data to the audio assistant device.

[1237] Audio assistant: Re-adjusts the filtered audio data for optimal playback.

[1238] 4. Transcribe and save conversations

[1239] Device (smartphone): Audio data during conversation is sent to the server in real time in streaming format.

[1240] Server: Analyzes the voice data, generates text data, and sends it back to the smartphone.

[1241] Device (smartphone): Text data is displayed in the app and saved in the internal storage.

[1242] 5. Language Correction by Generative AI

[1243] Device (smartphone): Sends inaudible audio data to the generating AI.

[1244] Server: Generative AI analyzes the voice data and corrects it into natural language.

[1245] Terminal (smartphone): Receives the corrected data and plays it on the audio assistant device.

[1246] 6. Emotion Recognition by Emotion Engine

[1247] Device (smartphone): Sends voice data collected by the microphone to the emotion engine.

[1248] Emotion engine: Analyzes voice data and recognizes the user's emotions. The recognized emotional information is reflected in filtering parameters and generation AI.

[1249] Device (smartphone): Optimize voice data settings based on feedback from the emotion engine.

[1250] Specific examples

[1251] Use at family dinners

[1252] User: Places smartphone in the center of the table, connects audio assistant via Bluetooth, adjusts equalization and noise cancellation settings through the app, and enables emotion engine.

[1253] Device (smartphone): The microphone collects surrounding conversations and transmits them in real time to the voice assistant. The voice data is sent to the server to obtain text data, which is then displayed on the screen. The emotion engine analyzes the collected voice data, recognizes the user's emotions, and dynamically adjusts voice filtering parameters based on that data.

[1254] Server: Analyzes the voice data, generates text in real time, and sends it back to the smartphone.

[1255] Terminal (smartphone): Saves the text data and sends the correction results to the voice assistant device.

[1256] In this way, the system of the present invention is designed to be easy for users to operate, and the emotion engine optimizes voice data, enabling more effective and natural communication.

[1257] The processing flow will be explained below.

[1258] Step 1:

[1259] User: Downloads and installs the smartphone app. Opens the app and creates a new account.

[1260] Step 2:

[1261] User: Power on the audio assistant and set it to Bluetooth pairing mode.

[1262] Step 3:

[1263] Device (smartphone): Open the Bluetooth settings screen and search for the audio assist device. Select the audio assist device from the search results and send a pairing request. If pairing is successful, establish a connection with the audio assist device.

[1264] Step 4:

[1265] Device (smartphone): Once a Bluetooth connection with the audio assistant is established, obtain microphone permission from the app and begin collecting audio.

[1266] Step 5:

[1267] Device (smartphone): The smartphone's microphone collects surrounding sounds and stores them in a buffer in real time.

[1268] Step 6:

[1269] Terminal (smartphone): Periodically performs digital signal processing on the audio data in the buffer and transmits it to the audio assistant device via Bluetooth.

[1270] Step 7:

[1271] Audio assistant: Decodes the received audio data and adjusts the volume and frequency characteristics to suit the user before playing it back.

[1272] Step 8:

[1273] User: Adjust equalization and noise cancellation parameters in the settings screen of the smartphone app.

[1274] Step 9:

[1275] Device (smartphone): Receives user setting changes and updates the voice data filtering algorithm. Digital filtering is applied to reduce noise before sending the audio data to the voice assistant device.

[1276] Step 10:

[1277] Audio assistant: Reconditions the received filtered audio data to play the best audio for the user.

[1278] Step 11:

[1279] Device (smartphone): Audio data during conversation is sent to the server in real time in streaming format.

[1280] Step 12:

[1281] Server: The received voice data is analyzed using a voice recognition engine and text data is generated.

[1282] Step 13:

[1283] Server: The generated text data is sent back to the smartphone.

[1284] Step 14:

[1285] Device (smartphone): The generated text data is displayed within the app and saved in the internal storage.

[1286] Step 15:

[1287] User: Type into the app to mark parts of the audio that are difficult to hear.

[1288] Step 16:

[1289] Device (smartphone): Sends the marked audio data to the server and requests correction by the generation AI.

[1290] Step 17:

[1291] Server: The generative AI analyzes the voice data and corrects it into natural language.

[1292] Step 18:

[1293] Server: Sends the corrected audio data back to the smartphone.

[1294] Step 19:

[1295] Terminal (smartphone): Notifies the user of the received correction data and plays it on the audio assistant device.

[1296] Step 20:

[1297] Device (smartphone): Sends voice data collected by the microphone to the emotion engine.

[1298] Step 21:

[1299] Emotion engine: Analyzes voice data and recognizes the user's emotions. The recognized emotional information is reflected in filtering parameters and generation AI.

[1300] Step 22:

[1301] Device (smartphone): Optimize voice data settings based on feedback from the emotion engine.

[1302] In this way, the system of the present invention provides a device that users can easily operate and configure, and by optimizing voice data using an emotion engine, it is possible to achieve more effective and natural communication.

[1303] Example 2

[1304] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1305] In recent years, advances in speech recognition and artificial intelligence technologies have improved the convenience of voice-based user interfaces. However, there are problems with degradation of voice data quality due to environmental noise and the user's emotional state. Furthermore, if voice data is not converted to text or corrected in real time, the user experience is significantly impaired. Furthermore, the lack of emotion recognition functionality poses a challenge, making it difficult to optimally process voice according to the user's emotional state.

[1306] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1307] In this invention, the server includes a high-performance voice assistant device, a mobile communication device electrically connected to the voice assistant device, communication means for customizing voice data in the mobile communication device, processing means for performing noise cancellation processing on the voice data, correction means having generative artificial intelligence for correcting the voice data to natural language, display means for converting the voice data into text and displaying it, emotion recognition means for analyzing the voice data and identifying the user's emotion, and optimization means for optimizing other processes based on the emotion information identified by the emotion recognition means, thereby enabling high-quality voice data processing in real time and optimal voice processing according to the user's emotional state.

[1308] A "high-performance audio assist device" is a device that reduces external noise and provides clear audio to the user.

[1309] "Mobile communications device" means a portable electronic device, such as a smartphone or tablet, on which an application for collecting, processing, and displaying voice data is installed.

[1310] "Communication means" refers to a wireless communication interface, such as Bluetooth, for communicating data between the mobile communication device and the audio assistant device.

[1311] "Processing means" refers to the function of performing digital signal processing (DSP) on the audio data collected by the microphone to eliminate noise.

[1312] "Correction means" refers to the function of correcting voice data into natural language using a generative AI model.

[1313] "Display means" refers to the function of converting voice data into text and displaying it on the screen of a smartphone or tablet.

[1314] "Emotion recognition means" refers to an engine that analyzes voice data to identify the user's emotions.

[1315] The "optimization means" refers to a function that optimizes other processes based on the emotion information identified by the emotion recognition means.

[1316] The present invention relates to a system that combines a highly functional voice assistant device, a mobile communication device electrically connected to the device, and a generative AI model and emotion engine. Specifically, the present invention uses the following hardware and software:

[1317] Hardware Configuration

[1318] 1. Audio assistants

[1319] A sophisticated audio assistant is a device that reduces external noise and provides clear audio to users, and can be connected to a mobile communication device via a wireless communication interface such as Bluetooth.

[1320] 2. Mobile Communication Devices

[1321] The mobile communication device is a portable electronic device such as a smartphone or tablet. It is connected to the audio assistant via Bluetooth and has installed an application for collecting, processing, and displaying audio data.

[1322] Software Configuration

[1323] 1. Means of communication

[1324] It is equipped with a wireless communication interface such as Bluetooth for communicating data between the mobile communication device and the audio assistant device.

[1325] 2. Processing means

[1326] This function performs digital signal processing (DSP) on audio data collected by a microphone to eliminate noise.

[1327] 3. Correction means

[1328] This function uses a generative AI model to correct voice data into natural language.

[1329] 4. Display means

[1330] This function converts voice data into text and displays it on the screen of a smartphone or tablet.

[1331] 5. Emotion recognition means

[1332] It has an emotion engine that analyzes voice data to identify the user's emotions.

[1333] 6. Optimization Methods

[1334] This is a function that optimizes other processes based on the emotion information identified by the emotion recognition means.

[1335] Specific examples

[1336] Use at family dinners

[1337] User: Places smartphone in the center of the table, connects audio assistant via Bluetooth, adjusts equalization and noise cancellation settings through the app, and enables emotion engine.

[1338] Device (smartphone): The microphone collects surrounding conversations and transmits them to the voice assistant in real time. The voice data is sent to the server to obtain text data, which is then displayed on the screen.

[1339] Emotion Engine: Analyzes voice data and recognizes user emotions. Dynamically adjusts voice filtering parameters based on the recognition.

[1340] Server: Analyzes the voice data, generates text in real time, and sends it back to the smartphone.

[1341] Terminal (smartphone): Saves the text data and sends the correction results to the voice assistant device.

[1342] Prompt Sentence Examples

[1343] An example of a prompt a user can send to a generative AI model is one that includes the instruction, "Please remove the noise from this audio data and correct it into natural language that is easy to listen to."

[1344] The above examples of specific use cases and prompt sentences will help understand the detailed implementation of the system of the present invention.

[1345] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1346] Program processing flow

[1347] Step 1: Bluetooth connection

[1348] The user launches the smartphone app and sets the audio assistant device to Bluetooth pairing mode, which allows the smartphone to recognize the audio assistant device.

[1349] The device (smartphone) searches for Bluetooth devices and displays a list of nearby Bluetooth devices. The user can select an audio assistant device from the list.

[1350] The terminal (smartphone) establishes Bluetooth pairing with the selected device. The input is the user's operation, and the output is the establishment of pairing.

[1351] Step 2: Collect and send audio data

[1352] The device (smartphone) collects ambient sound in real time using a built-in microphone, stores it in a buffer, and then uses digital signal processing (DSP) to reduce noise.

[1353] The terminal (smartphone) transmits the audio data stored in the buffer to the audio assistant device via Bluetooth. The input is the audio data collected by the microphone, and the output is the data in the buffer after DSP processing.

[1354] The audio assistant decodes the received audio data, adjusts the volume and frequency characteristics to suit the audio, and plays it back. The input is the audio data sent from the terminal, and the output is clear audio provided to the user.

[1355] Step 3: Equalizing and noise reduction

[1356] Users can adjust equalization and noise cancellation parameters in the settings screen of the smartphone app, using sliders and presets to customize the sound quality.

[1357] The device (smartphone) updates the filtering algorithm based on the user's settings and applies new parameters. The input is the user's adjusted parameters, and the output is the updated filtering algorithm.

[1358] The terminal (smartphone) transmits the filtered audio data to the audio assistance device via Bluetooth.

[1359] The audio assistant then readjusts the filtered audio data and plays it back optimally, with the input being the filtered audio data and the output being audio optimized for the user.

[1360] Step 4: Transcribe and save the conversation

[1361] The device (smartphone) transmits the voice data of the conversation to the server in real time in streaming format. The input is the voice data collected in real time, and the output is the data transmitted to the server.

[1362] The server analyzes the received voice data and converts it into text data. The input is the voice data sent from the terminal, and the output is text data.

[1363] The device (smartphone) displays the text data returned from the server on the app screen and saves it in its internal storage. The input is the text data from the server, and the output is the displayed and saved text data.

[1364] Step 5: Language correction by generative AI

[1365] The device (smartphone) sends inaudible voice data to the generative AI model. The input is voice data, and the output is data sent to the server.

[1366] The server analyzes the voice data received by the generative AI model and corrects it into natural language. The input is the voice data sent from the device, and the output is the corrected language data.

[1367] The terminal (smartphone) receives the corrected data and plays it on the audio assist device. The input is the corrected data from the server, and the output is the played audio.

[1368] Step 6: Emotion Recognition with the Emotion Engine

[1369] The device (smartphone) transmits voice data collected by a microphone to the emotion engine in real time. The input is the collected voice data, and the output is data sent to the emotion engine.

[1370] The emotion engine analyzes the received voice data and identifies the user's emotion. The input is the voice data, and the output is the identified emotion information.

[1371] The device (smartphone) optimizes the voice data settings based on feedback from the emotion engine. The input is the emotion engine feedback, and the output is the optimized voice data.

[1372] Through the above steps, the system can provide users with comfortable voice interaction.

[1373] (Application example 2)

[1374] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1375] While conventional voice assistance systems can filter voice data and eliminate noise, they lack the ability to recognize the user's emotions and optimize voice data accordingly. This has resulted in issues with not being able to provide appropriate voice feedback in certain situations. Furthermore, they lacked the ability to correct speech to natural language using generative AI and real-time data sharing, making effective security management difficult.

[1376] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server is equipped with an emotion recognition engine and includes means for analyzing voice data in real time, means for correcting voice data using a generation AI, and means for dynamically adjusting filtering parameters. This makes it possible to recognize the user's emotions and optimize voice data based on them. In addition, voice data can be sent to the server in real time, allowing appropriate feedback to be obtained immediately, thereby achieving more effective security management.

[1377] A "high-performance audio assist device" is a device that reduces external noise and provides clear audio.

[1378] A "mobile communication device" is a portable electronic device such as a smartphone or tablet.

[1379] "Communication means" refers to the function of collecting and transmitting voice data and exchanging data between devices.

[1380] The "processing means" is a function that performs digital signal processing such as noise cancellation on audio data.

[1381] The "correction means" is a function that corrects voice data into natural language using generative artificial intelligence.

[1382] The "display means" is a function that converts voice data into text and displays it visually.

[1383] The "emotion recognition means" is a function that recognizes the user's emotions from the voice data and dynamically adjusts the voice filtering parameters accordingly.

[1384] "Means for transmitting voice data in real time" refers to a function for transmitting voice data in real time and for immediate processing and feedback.

[1385] The "equalizing means" is a function that adjusts audio characteristics based on equalizing parameters that can be set by the user.

[1386] The present invention relates to a system that uses a highly functional voice assistant device and a mobile communication device electrically connected to it to effectively process and optimize voice data in a user's environment and apply it to security services. The following describes the specific system configuration and its operation method.

[1387] System configuration

[1388] 1. Hardware

[1389] High-performance audio assist device: Reduces external noise and provides clear audio to users. Connects to mobile communication devices using wireless communication interfaces such as Bluetooth.

[1390] Mobile Communication Device: A portable electronic device, such as a smartphone or tablet, that collects, processes, and displays audio data. It contains a microphone and a Bluetooth module.

[1391] 2. Software

[1392] Bluetooth module: Wirelessly connects the audio assistant device to a smartphone, sending and receiving audio data.

[1393] Digital signal processing (DSP): Performs noise cancellation, equalization, and other processing on collected audio data.

[1394] Emotion engine: Analyzes voice data and recognizes the user's emotions. As a concrete example, it uses a natural language processing library (e.g., Google TensorFlow).

[1395] Generative AI: Generative AI analyzes voice data and corrects it into natural language.

[1396] Display function: Converts voice data into text and displays it visually.

[1397] Operation Overview

[1398] 1. Bluetooth connection

[1399] Server: Pairs the audio assistant device with the mobile communication device via Bluetooth.

[1400] Device: Search for devices, select the audio assistant and establish pairing.

[1401] 2. Collection and transmission of voice data

[1402] Device: Surrounding sounds are collected using the smartphone's microphone and stored in a buffer in real time.

[1403] Terminal: The audio data in the buffer is filtered using digital signal processing and sent to the audio assistant device via Bluetooth.

[1404] Audio assistant: Decodes received audio data and adjusts it to the optimum volume and frequency characteristics.

[1405] 3. Equalization and noise cancellation

[1406] User: Adjust equalization and noise cancellation parameters in the settings screen of the smartphone app.

[1407] Device: Receives user settings, updates the filtering algorithm, and sends it to the audio assistant device.

[1408] 4. Transcribe and save conversations

[1409] Terminal: Sends audio data to the server in real time in streaming format.

[1410] Server: Analyzes the voice data, generates text data, and sends it back to the device.

[1411] Device: Text data is displayed in the app and saved in the internal storage.

[1412] 5. Language Correction by Generative AI

[1413] Device: Sends inaudible voice data to the generating AI.

[1414] Server: The generative AI analyzes the voice data and corrects it into natural language.

[1415] Terminal: Receives the corrected data and plays it on the audio assistant device.

[1416] 6. Emotion Recognition by Emotion Engine

[1417] Terminal: Voice data collected by the microphone is sent to the emotion engine.

[1418] Server: The emotion engine analyzes the voice data and recognizes the user's emotions. This information is reflected in the filtering parameters and generation AI.

[1419] Device: Optimized voice data settings based on feedback from the emotion engine.

[1420] Specific examples

[1421] Office Security Monitoring

[1422] The user acts as a security manager, placing a smartphone in the security room and pairing it with a voice assistant. The system analyzes conversations in the office in real time and can send alerts to managers if it detects tension or stress.

[1423] Prompt Sentence Examples

[1424] "What are the potential emotions that could be present in this scene? Identify these emotions based on the text data. Then update the emotion parameters in the emotion engine and apply the new security settings."

[1425] Thus, the present invention can be implemented in a variety of security environments, and specific hardware and software can be used to achieve effective security management.

[1426] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1427] Step 1:

[1428] Device: Set the audio assistant device to Bluetooth pairing mode, launch the app on the smartphone and search for devices, select the audio assistant device from the search results, and establish pairing. The input is Bluetooth device information, and the output is the status of pairing establishment.

[1429] Step 2:

[1430] Terminal: The smartphone's microphone collects surrounding audio and stores it in a buffer in real time. The input is the surrounding audio data, and the output is the collected audio data in the buffer. Specifically, the smartphone's built-in microphone captures audio, converts the data into digital format, and stores it in the buffer.

[1431] Step 3:

[1432] Terminal: The audio data in the buffer is filtered using a digital signal processing algorithm and sent to the audio assistant via Bluetooth. The input is the audio data in the buffer and the output is the filtered audio data. The audio data is noise-reduced and equalized and sent via the Bluetooth module.

[1433] Step 4:

[1434] Audio assistant: Decodes received audio data and plays it back after adjusting the volume and frequency characteristics to suit the user. The input is filtered audio data, and the output is clear audio. Specifically, the decoding process converts the digital signal into an analog signal, which is then output through a speaker.

[1435] Step 5:

[1436] User: Adjusts equalization and noise cancellation parameters on the settings screen within the smartphone app. The input is the user's setting parameters, and the output is the updated filtering algorithm. The app's UI intuitively manipulates sliders and drop-down menus to change settings.

[1437] Step 6:

[1438] Terminal: Receives user settings, updates the filtering algorithm, and sends the new algorithm to the audio assistant. The input is the updated filtering parameters, and the output is the new filtering algorithm applied to the audio assistant.

[1439] Step 7:

[1440] Terminal: Sends the voice data of the conversation to the server in real time in streaming format. The input is filtered voice data, and the output is streaming data. The data is securely sent to the server via the SSL / TLS protocol.

[1441] Step 8:

[1442] Server: Analyzes the voice data, generates text data, and sends it back to the smartphone. The input is streaming voice data, and the output is the generated text data. The voice recognition engine analyzes the data and converts it into text format.

[1443] Step 9:

[1444] Terminal: Text data is displayed in the app and saved to the internal storage. The input is the text data returned from the server, and the output is the displayed text and saved data. Specifically, it is displayed in the app's text view and saved to the database.

[1445] Step 10:

[1446] Device: Sends inaudible voice data to the generation AI. The input is the collected voice data, and the output is the request data for correction. The API of the generation AI model is called and the voice data is sent.

[1447] Step 11:

[1448] Server: The generative AI analyzes the voice data and corrects it into natural language. The input is the voice data, and the output is the corrected voice data. The generative AI model removes noise and converts it into clear pronunciation.

[1449] Step 12:

[1450] Terminal: Receives the corrected data and plays it back to the audio assistant. The input is the corrected data returned from the server, and the output is the reproduced natural speech.

[1451] Step 13:

[1452] Terminal: Sends voice data collected by a microphone to the emotion engine. The input is the collected voice data, and the output is the emotion recognition request data.

[1453] Step 14:

[1454] Server: The emotion engine analyzes the voice data and recognizes the user's emotions. The input is the voice data, and the output is the recognized emotion information. The emotion data is analyzed through the analytics engine.

[1455] Step 15:

[1456] Terminal: Optimizes voice data settings based on feedback from the emotion engine. The input is the recognized emotion information, and the output is optimized filtering parameters.

[1457] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1458] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1459] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1460] [Fourth embodiment]

[1461] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1462] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1463] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1464] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1465] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1466] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1467] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1468] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1469] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1470] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1471] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1472] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1473] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1474] The system of the present invention is composed of a highly functional voice assistant device and a mobile communication device electrically connected to it. In this system, the mobile communication device (e.g., a smartphone) plays a central role, providing functions such as voice data customization, noise cancellation, correction using generation AI, and transcription.

[1475] System configuration

[1476] 1. Audio assistants

[1477] The advanced audio assistant effectively amplifies surrounding sounds and has filters to reduce noise according to the user's hearing ability, and can be connected to a mobile communication device via a wireless communication interface such as Bluetooth.

[1478] 2. Mobile Communication Devices

[1479] The device is a portable electronic device such as a smartphone or tablet, which connects to the audio assistant via wireless communication such as Bluetooth and has an application installed to collect, process, and display audio data.

[1480] Program processing flow

[1481] The program processing of this system is carried out as follows.

[1482] 1. Bluetooth connection

[1483] User: Launch the smartphone app and set the audio assistant device to Bluetooth pairing mode.

[1484] Device (smartphone): Search for Bluetooth devices, select the audio assistant device, and establish pairing.

[1485] 2. Collection and transmission of voice data

[1486] Device (smartphone): The microphone collects surrounding sounds and stores them in a buffer. The audio data in the buffer is then transmitted in real time to the audio assistant via wireless communication.

[1487] Audio assistant: Decodes the received audio data and adjusts the volume and frequency characteristics to suit the user before playing it back.

[1488] 3. Equalization and noise cancellation

[1489] User: Adjust equalization and noise cancellation parameters in the settings screen of the smartphone app.

[1490] Device (smartphone): Receives user settings, updates the filtering algorithm, applies a noise cancellation filter to the audio data, and sends it to the audio assistant device.

[1491] Audio assistant: Apply filtered audio data to reduce external noise before playback.

[1492] 4. Transcribe and save conversations

[1493] Terminal (smartphone): Audio data collected by a microphone is sent to a server, where real-time voice recognition is performed.

[1494] Server: Analyzes the voice data, generates text data, and sends it back to the smartphone.

[1495] Device (smartphone): The generated text is displayed on the screen and saved in the internal storage.

[1496] 5. Language Correction by Generative AI

[1497] Device (smartphone): Sends inaudible audio data to the generating AI.

[1498] Server: The generative AI analyzes the voice data and corrects it into natural language.

[1499] Terminal (smartphone): Receives the corrected data and sends it to the audio assistant device for playback.

[1500] Specific examples

[1501] 1. Use at family dinners

[1502] User: Places smartphone in the center of the table and connects to audio assistant via Bluetooth. Adjusts equalization and noise cancellation settings through the app.

[1503] Device (smartphone): The microphone picks up surrounding conversations and transmits them in real time to the voice assistant. The voice data is sent to the server, which then retrieves the text data and displays it on the screen.

[1504] Server: Analyzes the voice data, generates text in real time, and sends it back to the smartphone.

[1505] Terminal (smartphone): Saves the text data and sends the correction results to the voice assistant device.

[1506] Thus, the system of the present invention allows elderly people to participate more naturally in conversations with their families, providing a practical way to compensate for lost communication.

[1507] The processing flow will be explained below.

[1508] Step 1:

[1509] User: Downloads and installs the smartphone app. Opens the app and creates a new account.

[1510] Step 2:

[1511] User: Power on the audio assistant and set it to Bluetooth pairing mode.

[1512] Step 3:

[1513] Device (smartphone): Open the Bluetooth settings screen and search for the audio assist device. Select the audio assist device from the search results and send a pairing request. If pairing is successful, establish a connection with the audio assist device.

[1514] Step 4:

[1515] Device (smartphone): Once a Bluetooth connection with the audio assistant is established, obtain microphone permission from the app and begin collecting audio.

[1516] Step 5:

[1517] Device (smartphone): The smartphone's microphone collects surrounding sounds and stores them in a buffer in real time.

[1518] Step 6:

[1519] Terminal (smartphone): Periodically performs digital signal processing on the audio data in the buffer and transmits it to the audio assistant device via Bluetooth.

[1520] Step 7:

[1521] Audio assistant: Decodes the received audio data and adjusts the volume and frequency characteristics to suit the user before playing it back.

[1522] Step 8:

[1523] User: Adjust equalization and noise cancellation parameters in the settings screen of the smartphone app.

[1524] Step 9:

[1525] Device (smartphone): Receives user setting changes and updates the voice data filtering algorithm. Digital filtering is applied to reduce noise before sending the audio data to the voice assistant device.

[1526] Step 10:

[1527] Audio assistant: Reconditions the received filtered audio data to play the best audio for the user.

[1528] Step 11:

[1529] Device (smartphone): Audio data during conversation is sent to the server in real time in streaming format.

[1530] Step 12:

[1531] Server: The received voice data is analyzed using a voice recognition engine and text data is generated.

[1532] Step 13:

[1533] Server: Sends the generated text data back to the smartphone.

[1534] Step 14:

[1535] Device (smartphone): The generated text data is displayed within the app and saved in the internal storage.

[1536] Step 15:

[1537] User: Tells the app to mark parts of the audio that are difficult to hear.

[1538] Step 16:

[1539] Device (smartphone): Sends the marked audio data to the server and requests correction by the generation AI.

[1540] Step 17:

[1541] Server: The generative AI analyzes the voice data and corrects it into natural language.

[1542] Step 18:

[1543] Server: Sends the corrected audio data back to the smartphone.

[1544] Step 19:

[1545] Terminal (smartphone): Notifies the user of the received correction data and plays it on the audio assistant device.

[1546] In this way, the system of the present invention provides a device that users can easily operate and set up, and not only improves the quality of speech but also has a transcription function that allows users to review past conversations. This system allows users to achieve more natural and comfortable communication.

[1547] Example 1

[1548] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1549] Conventional speech assist devices simply amplify speech, and are not effective enough in noisy environments or when conversations are difficult to hear. Furthermore, they lack the ability to convert speech into text, making it impossible to visually confirm the content of a conversation. This creates challenges for communication, particularly for the elderly and those with hearing impairments.

[1550] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1551] In this invention, the server includes a high-performance audio amplifier, a portable information terminal electrically connected to the audio amplifier, a communication means for customizing audio data on the portable information terminal, a processing means for performing noise reduction processing on the audio data, a correction means having a generative artificial intelligence for correcting the audio data to natural language, a display means for converting the audio data into text and displaying it, a transmission means for storing the audio data in a buffer on the portable information terminal and transmitting the audio data to the audio amplifier in real time via wireless communication, and a setting means for the user to adjust equalization and noise reduction settings on the portable information terminal. This allows for clear audio even in noisy environments and for smooth communication by converting audio data into text and displaying it in real time.

[1552] An "audio amplifier" is a device for audio assistance that has the function of amplifying sound with high precision and reducing noise.

[1553] "Mobile information terminal" is a general term for portable electronic devices such as smartphones and tablets.

[1554] "Communication means" refers to devices or software that have the functionality to collect, send, and receive voice data.

[1555] "Processing means" refers to hardware or software for performing specific processing on audio data.

[1556] "Generative AI" is an AI technology that analyzes voice data and corrects it to make it more natural.

[1557] "Correction means" refers to a device or software that has the function of converting and correcting voice data into natural language.

[1558] "Display means" refers to a device or software that visually presents the text-converted voice data to the user.

[1559] A "buffer" is a storage area for temporarily storing audio data.

[1560] "Transmission means" refers to a device or software that has the function of transmitting audio data to another device via wireless communication.

[1561] "Setting means" refers to a device or software that has a function that allows the user to adjust sound quality and filtering parameters.

[1562] "Equalizing" is a process that emphasizes or attenuates specific frequency bands in audio.

[1563] "Noise reduction" refers to a processing technique for removing unwanted noise from an audio signal.

[1564] "Text conversion" refers to the process of converting voice data into text data.

[1565] "Wireless communication" refers to a communication technology for sending and receiving data without using cables.

[1566] The system of this invention is primarily composed of a high-performance sound amplifier, a mobile information terminal (such as a smartphone or tablet), and a server. This system provides clear audio even in noisy environments, converts audio data into text and displays it in real time, and enables smooth conversations with elderly people and those with hearing impairments.

[1567] Hardware configuration

[1568] Sound amplification equipment:

[1569] It is equipped with high-precision audio amplification and noise reduction filters.

[1570] It can be connected to a mobile information terminal via wireless communication such as Bluetooth.

[1571] Personal digital assistants:

[1572] A smartphone or tablet with a voice assistant application installed.

[1573] Surrounding sounds are picked up through the microphone and stored in a buffer.

[1574] Wireless communication (such as Bluetooth) is used for data communication.

[1575] server:

[1576] It analyzes voice data, transcribes it, and corrects it using AI. It has high-performance processing capabilities and performs secure data communication.

[1577] Software configuration

[1578] Voice-assisted applications:

[1579] Customize audio data, set noise reduction, transcribe audio in real time, and save text.

[1580] It has a settings screen that allows you to adjust equalization and noise reduction parameters.

[1581] Generative AI models:

[1582] It has the ability to analyze voice data and correct it into natural language.

[1583] It runs on the server and returns the correction results to the mobile information terminal.

[1584] Specific examples of operation procedures

[1585] Here, as a concrete example of how the system is used, we will show a scene where the system is used at a family dinner party.

[1586] Example of use at a family dinner:

[1587] User: Places smartphone in the center of the table and connects to sound amplification device via Bluetooth. Adjusts equalization and noise cancellation settings through the app.

[1588] Device (smartphone): The microphone collects the conversations during the dinner party and transmits them in real time to an audio amplifier. The audio data is then sent to a server to obtain text data, which is then displayed on the smartphone screen.

[1589] Server: Analyzes the voice data, generates text in real time, and sends it back to the smartphone. For example, the utterance "Mom, would you like another serving?" is generated as text data.

[1590] Device (smartphone): The generated text data is displayed on the screen so the user can check the content of the conversation. The text data is also saved in the internal storage for later reference. If necessary, language correction is performed on the voice data, and the correction results are sent to the audio amplification device.

[1591] Prompt Sentence Examples

[1592] "How exactly do I use a speech assistive device at a family dinner?"

[1593] This system allows elderly and hearing-impaired people to participate more naturally in conversations with their families and compensate for lost communication.

[1594] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1595] Step 1: Bluetooth connection

[1596] Input: Sound amplification device in pairing mode. User launches smartphone app.

[1597] Operation:

[1598] User: Launch the smartphone app and put the audio amplifier into Bluetooth pairing mode by pressing and holding the pairing button on the audio amplifier.

[1599] Device (smartphone): Start Bluetooth device search and detect the sound amplifier. Select the sound amplifier from the device list and establish pairing. The message "Connection completed" will be displayed.

[1600] Output: "Connected" message upon successful pairing connection.

[1601] Step 2: Collect and send audio data

[1602] Input: Smartphone connected to sound amplification device via Bluetooth.

[1603] Operation:

[1604] Device (smartphone): Uses a microphone to collect surrounding sounds in real time. This data is temporarily stored in a buffer. Consider a dinner party scenario, where the sound of conversations is collected.

[1605] Terminal (smartphone): The audio data stored in the buffer is sent to the audio amplifier via wireless communication (Bluetooth). The data is divided and sent in packet format.

[1606] Audio Amplification Device: Decodes the received audio data and adjusts the volume and frequency characteristics, so that the audio is played back to the user in a clear and crisp manner.

[1607] Output: The audio data with adjusted volume and frequency response reaches the user.

[1608] Step 3: Equalization and noise cancellation settings

[1609] Input: A user is accessing the settings screen within a smartphone app.

[1610] Operation:

[1611] Users: Adjust equalization and noise cancellation parameters in the app's settings, for example, by boosting high frequencies.

[1612] Device (smartphone): Receives user settings and updates the filtering algorithm. When the new settings are reflected in the app, the algorithm processes the audio data accordingly.

[1613] Terminal (smartphone): The filtered audio data is sent back to the audio amplifier. After noise removal, the quality of the audio data is improved.

[1614] Sound amplifier: Reproduces the filtered audio data so that the user hears it with external noise reduced.

[1615] Output: The filtered audio data is delivered to the user.

[1616] Step 4: Transcribe and save the conversation

[1617] Input: Audio data is collected by a smartphone and sent to a server via the internet.

[1618] Operation:

[1619] Device (smartphone): Sends voice data collected by the microphone to the server and starts the voice recognition process.

[1620] Server: The received voice data is converted into text data using a generative AI model. The utterance "Hello, how are you?" is generated as text data.

[1621] Server: Sends the generated text data back to the smartphone.

[1622] Device (smartphone): The text data is displayed on the screen so that the user can check the contents of the conversation, and the text data is saved in the internal storage for later reference.

[1623] Output: The textual audio data is displayed on the screen and saved.

[1624] Step 5: Language correction by generative AI

[1625] Input: Difficult-to-hear audio data collected by a smartphone is sent to the generation AI.

[1626] Operation:

[1627] Device (smartphone): Sends inaudible audio data to the generating AI.

[1628] Server: The generation AI analyzes the voice data and corrects it into natural language. Unclear parts are corrected and clear text data is generated.

[1629] Server: Sends the corrected data back to the smartphone.

[1630] Terminal (smartphone): Receives the corrected data, sends it to an audio amplifier, and plays it back. The corrected audio reaches the user, making it easier to hear.

[1631] Output: The corrected audio data is delivered to the user.

[1632] By using specific actions performed at each step, this system can effectively provide communication support to the elderly and the hearing impaired.

[1633] (Application example 1)

[1634] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1635] Elderly people and those with hearing problems often have difficulty hearing in-store announcements and conversations with store clerks when shopping comfortably in physical stores. This not only degrades the quality of the shopping experience, but can also lead to missing important information. Furthermore, existing voice assistants lack effective noise cancellation and voice correction, which can lead to reduced speech recognition accuracy due to environmental noise.

[1636] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1637] In this invention, the server includes a high-performance voice assistance device, a mobile communication device electrically connected to the voice assistance device, a communication means for customizing voice data on the mobile communication device, a processing means for performing noise cancellation processing on the voice data, a correction means having a generation artificial intelligence for correcting the voice data to natural language, a display means for converting the voice data into text and displaying it, a voice output means for playing back the generated text data, and a server communication means for transmitting the voice data to an external server and retrieving the corrected text data. This allows users who require hearing assistance to easily hear in-store announcements and conversations with store clerks while shopping in a physical store, as well as visually confirm the corrected text data. Furthermore, correcting and playing back the generated text data can provide a natural conversation experience.

[1638] A "high-performance audio assist device" is a device that effectively amplifies surrounding sounds and is equipped with a filter to reduce noise according to the user's hearing ability.

[1639] A "mobile communications device" is a portable electronic device such as a smartphone or tablet.

[1640] The "communication means" is a function including a wireless communication interface for connecting a high-performance voice assistant device and a mobile communication device and customizing voice data.

[1641] "Processing means" refers to algorithms or programs for performing noise cancellation processing on speech data in a mobile communication device.

[1642] "Correction means" refers to a function that uses generative AI to correct voice data into natural language.

[1643] "Display means" refers to a display or screen that converts voice data into text and displays it.

[1644] The "audio output means" refers to a function such as a speaker or earphone for reproducing the generated text data as audio.

[1645] The "server communication means" is a function for transmitting voice data from the mobile communication device to an external server and obtaining corrected text data from the server.

[1646] To implement this invention, the following system configuration and operation method are required. A highly functional audio assistance device is used in combination with a mobile communication device (smartphone or tablet). This enables hearing assistance in a store.

[1647] System Configuration

[1648] 1. Advanced audio assistive devices:

[1649] It effectively amplifies surrounding sounds and has a filter to reduce noise according to the user's hearing ability, and can be connected to a mobile communication device via a wireless communication interface such as Bluetooth.

[1650] 2. Mobile communication devices:

[1651] It is a portable electronic device that connects to the audio assistant via wireless communication such as Bluetooth. This device is equipped with an application that customizes audio data, cancels noise, corrects it using generative AI, transcribes it, and plays it back.

[1652] Hardware and software used

[1653] 1. Hardware:

[1654] Smartphone (iOS or Android)

[1655] High-performance Bluetooth-enabled audio assist device

[1656] Your smartphone's built-in microphone and speaker (or earphones)

[1657] 2. Software:

[1658] Speech recognition libraries (e.g., speech_recognition, Google Speech API)

[1659] Text-to-speech engine (e.g. pyttsx3)

[1660] Server communication library (e.g. requests)

[1661] External generation AI server (API)

[1662] Specific examples

[1663] 1. Scenario 1:

[1664] In a physical store, a user stands in front of a shelf and asks a store clerk where a product is located. The user launches an application on their smartphone and types the question into the microphone.

[1665] Example prompt:

[1666] "Please ask the microphone where the item is."

[1667] 2. Scenario 2:

[1668] The user listens to an announcement about a new campaign in front of the cash register at a physical store. The smartphone is placed in a location where the announcement can be heard, and the speech is converted into text, corrected, and the corrected text is output as voice.

[1669] Example prompt:

[1670] Please check the announcement

[1671] Processing flow

[1672] 1. Device (smartphone):

[1673] The microphone picks up surrounding sounds and transmits the audio data via Bluetooth to a high-performance audio assist device, which simultaneously processes the audio data in real time and transmits it to an external server as needed.

[1674] 2. Server:

[1675] It receives voice data, corrects it into natural language using generative AI, and sends the corrected text data back to the smartphone.

[1676] 3. Device (smartphone):

[1677] The received text data is displayed on a display and output as voice using a text-to-speech engine.

[1678] This allows users who require hearing assistance to easily hear in-store announcements and conversations with store clerks while shopping in a physical store, as well as visually confirm the information. Furthermore, by correcting and playing back the generated text data, a natural conversation experience can be provided.

[1679] The system enables people with hearing problems to more comfortably participate in public activities, reducing social barriers.

[1680] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1681] Step 1:

[1682] User: Launches the smartphone app and sets the audio assistant device to Bluetooth pairing mode. Through the smartphone app interface, the smartphone searches for the audio assistant device via Bluetooth and establishes pairing. Based on this input, the smartphone connects to the audio assistant device via Bluetooth communication. When pairing is successful, the device outputs a notification that the connection is complete.

[1683] Step 2:

[1684] Device (smartphone): The smartphone's microphone collects surrounding sounds in real time and stores them in a buffer. This collected sound data is used as input, and the audio data in the buffer is processed in real time and sent via wireless communication to the audio assistance device. The audio assistance device receives the transmitted audio data, converts it into volume and frequency characteristics adjusted to suit the user's hearing, and plays it back.

[1685] Step 3:

[1686] User: Adjusts equalization and noise cancellation parameters using the settings screen in the smartphone app. Using this setting information as input, the smartphone updates the filtering algorithm and resends the audio data with the noise cancellation filter applied to the audio assistive device. This allows the audio assistive device to reproduce audio with reduced external noise.

[1687] Step 4:

[1688] Terminal (smartphone): The voice data collected by the microphone is sent to the server, which is then requested to perform voice recognition processing. In this process, the server analyzes the voice data as input and generates text data. This generated text data is sent back to the smartphone, and is then displayed on the smartphone screen.

[1689] Step 5:

[1690] Server: The server inputs the voice data sent from the smartphone into a generative AI model and corrects difficult-to-hear parts into natural language. The generative AI corrects the voice data by analyzing it and converting it into appropriate language based on the context. The corrected text data is then sent back from the server to the smartphone.

[1691] Step 6:

[1692] Terminal (smartphone): Receives the corrected text data and displays it on the screen. It also uses a text-to-speech engine to output the text data as voice, allowing the user to confirm the corrected information both visually and audibly.

[1693] Step 7:

[1694] User: In a brick-and-mortar store, the user continues to actively use their smartphone, converting surrounding information into text as needed and receiving natural-language enhanced speech information. For example, if a user asks a store clerk where a product is located, the smartphone microphone collects the question as input and receives enhanced speech that has gone through all the processing steps as output, making the clerk's response easier to understand.

[1695] These processing steps enable users to enjoy sophisticated voice assistance and a natural conversational experience in physical stores.

[1696] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1697] The present invention relates to a system that combines a highly functional voice assistant device with an electrically connected mobile communication device and further includes an emotion engine that recognizes the user's emotions. This system customizes voice data, cancels noise, corrects using generative AI, converts text, and optimizes voice data through emotion recognition.

[1698] System configuration

[1699] 1. Audio assistants

[1700] High-performance audio assistants reduce external noise and provide clear audio to users, and can be connected to mobile communication devices via wireless communication interfaces such as Bluetooth.

[1701] 2. Mobile Communication Devices

[1702] The device is a portable electronic device such as a smartphone or tablet, which is connected to the audio assistant via Bluetooth and has an application installed to collect, process, and display audio data.

[1703] 3. Emotion Engine

[1704] The emotion engine identifies emotions from user speech and input data and uses that information to optimize other processes, such as adjusting filtering parameters for specific emotions in the speech data.

[1705] Program processing flow

[1706] The program processing of this system is carried out as follows.

[1707] 1. Bluetooth connection

[1708] User: Launch the smartphone app and set the audio assistant device to Bluetooth pairing mode.

[1709] Device (smartphone): Search for Bluetooth devices, select the audio assistant device, and establish pairing.

[1710] 2. Collection and transmission of voice data

[1711] Device (smartphone): Collects surrounding audio with a microphone and stores it in a buffer in real time.

[1712] Terminal (smartphone): Digital signal processing is performed on the audio data in the buffer, and the data is sent to the audio assistant device via Bluetooth.

[1713] Audio assistant: Decodes the received audio data and adjusts the volume and frequency characteristics to suit the user before playing it back.

[1714] 3. Equalization and noise cancellation

[1715] User: Adjust equalization and noise cancellation parameters in the settings screen of the smartphone app.

[1716] Device (smartphone): Receives user settings, updates the filtering algorithm, and sends noise-reduced audio data to the audio assistant device.

[1717] Audio assistant: Re-adjusts the filtered audio data for optimal playback.

[1718] 4. Transcribe and save conversations

[1719] Device (smartphone): Audio data during conversation is sent to the server in real time in streaming format.

[1720] Server: Analyzes the voice data, generates text data, and sends it back to the smartphone.

[1721] Device (smartphone): Text data is displayed in the app and saved in the internal storage.

[1722] 5. Language Correction by Generative AI

[1723] Device (smartphone): Sends inaudible audio data to the generating AI.

[1724] Server: The generative AI analyzes the voice data and corrects it into natural language.

[1725] Terminal (smartphone): Receives the corrected data and plays it on the audio assistant device.

[1726] 6. Emotion Recognition by Emotion Engine

[1727] Device (smartphone): Sends voice data collected by the microphone to the emotion engine.

[1728] Emotion engine: Analyzes voice data and recognizes the user's emotions. The recognized emotional information is reflected in filtering parameters and generation AI.

[1729] Device (smartphone): Optimize voice data settings based on feedback from the emotion engine.

[1730] Specific examples

[1731] Use at family dinners

[1732] User: Places smartphone in the center of the table, connects audio assistant via Bluetooth, adjusts equalization and noise cancellation settings through the app, and enables emotion engine.

[1733] Device (smartphone): The microphone collects surrounding conversations and transmits them in real time to the voice assistant. The voice data is sent to the server to obtain text data, which is then displayed on the screen. The emotion engine analyzes the collected voice data, recognizes the user's emotions, and dynamically adjusts voice filtering parameters based on that data.

[1734] Server: Analyzes the voice data, generates text in real time, and sends it back to the smartphone.

[1735] Terminal (smartphone): Saves the text data and sends the correction results to the voice assistant device.

[1736] In this way, the system of the present invention is designed to be easy for users to operate, and the emotion engine optimizes voice data, enabling more effective and natural communication.

[1737] The processing flow will be explained below.

[1738] Step 1:

[1739] User: Downloads and installs the smartphone app. Opens the app and creates a new account.

[1740] Step 2:

[1741] User: Power on the audio assistant and set it to Bluetooth pairing mode.

[1742] Step 3:

[1743] Device (smartphone): Open the Bluetooth settings screen and search for the audio assist device. Select the audio assist device from the search results and send a pairing request. If pairing is successful, establish a connection with the audio assist device.

[1744] Step 4:

[1745] Device (smartphone): Once a Bluetooth connection with the audio assistant is established, obtain microphone permission from the app and begin collecting audio.

[1746] Step 5:

[1747] Device (smartphone): The smartphone's microphone collects surrounding sounds and stores them in a buffer in real time.

[1748] Step 6:

[1749] Terminal (smartphone): Periodically performs digital signal processing on the audio data in the buffer and transmits it to the audio assistant device via Bluetooth.

[1750] Step 7:

[1751] Audio assistant: Decodes the received audio data and adjusts the volume and frequency characteristics to suit the user before playing it back.

[1752] Step 8:

[1753] User: Adjust equalization and noise cancellation parameters in the settings screen of the smartphone app.

[1754] Step 9:

[1755] Device (smartphone): Receives user setting changes and updates the voice data filtering algorithm. Digital filtering is applied to reduce noise before sending the audio data to the voice assistant device.

[1756] Step 10:

[1757] Audio assistant: Reconditions the received filtered audio data to play the best audio for the user.

[1758] Step 11:

[1759] Device (smartphone): Audio data during conversation is sent to the server in real time in streaming format.

[1760] Step 12:

[1761] Server: The received voice data is analyzed using a voice recognition engine and text data is generated.

[1762] Step 13:

[1763] Server: Sends the generated text data back to the smartphone.

[1764] Step 14:

[1765] Device (smartphone): The generated text data is displayed within the app and saved in the internal storage.

[1766] Step 15:

[1767] User: Type into the app to mark parts of the audio that are difficult to hear.

[1768] Step 16:

[1769] Device (smartphone): Sends the marked audio data to the server and requests correction by the generation AI.

[1770] Step 17:

[1771] Server: Generative AI analyzes the voice data and corrects it into natural language.

[1772] Step 18:

[1773] Server: Sends the corrected audio data back to the smartphone.

[1774] Step 19:

[1775] Terminal (smartphone): Notifies the user of the received correction data and plays it on the audio assistant device.

[1776] Step 20:

[1777] Device (smartphone): Sends voice data collected by the microphone to the emotion engine.

[1778] Step 21:

[1779] Emotion engine: Analyzes voice data and recognizes the user's emotions. The recognized emotional information is reflected in filtering parameters and generation AI.

[1780] Step 22:

[1781] Device (smartphone): Optimize voice data settings based on feedback from the emotion engine.

[1782] In this way, the system of the present invention provides a device that users can easily operate and configure, and by optimizing voice data using an emotion engine, it is possible to achieve more effective and natural communication.

[1783] Example 2

[1784] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1785] In recent years, advances in speech recognition and artificial intelligence technologies have improved the convenience of voice-based user interfaces. However, there are problems with degradation of voice data quality due to environmental noise and the user's emotional state. Furthermore, if voice data is not converted to text or corrected in real time, the user experience is significantly impaired. Furthermore, the lack of emotion recognition functionality poses a challenge, making it difficult to optimally process voice according to the user's emotional state.

[1786] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1787] In this invention, the server includes a high-performance voice assistant device, a mobile communication device electrically connected to the voice assistant device, communication means for customizing voice data in the mobile communication device, processing means for performing noise cancellation processing on the voice data, correction means having generative artificial intelligence for correcting the voice data to natural language, display means for converting the voice data into text and displaying it, emotion recognition means for analyzing the voice data and identifying the user's emotion, and optimization means for optimizing other processes based on the emotion information identified by the emotion recognition means, thereby enabling high-quality voice data processing in real time and optimal voice processing according to the user's emotional state.

[1788] A "high-performance audio assist device" is a device that reduces external noise and provides clear audio to the user.

[1789] "Mobile communications device" means a portable electronic device, such as a smartphone or tablet, on which an application for collecting, processing, and displaying voice data is installed.

[1790] "Communication means" refers to a wireless communication interface, such as Bluetooth, for communicating data between the mobile communication device and the audio assistant device.

[1791] "Processing means" refers to the function of performing digital signal processing (DSP) on the audio data collected by the microphone to eliminate noise.

[1792] "Correction means" refers to the function of correcting voice data into natural language using a generative AI model.

[1793] "Display means" refers to the function of converting voice data into text and displaying it on the screen of a smartphone or tablet.

[1794] "Emotion recognition means" refers to an engine that analyzes voice data to identify the user's emotions.

[1795] The "optimization means" refers to a function that optimizes other processes based on the emotion information identified by the emotion recognition means.

[1796] The present invention relates to a system that combines a highly functional voice assistant device, a mobile communication device electrically connected to the device, and a generative AI model and emotion engine. Specifically, the present invention uses the following hardware and software:

[1797] Hardware Configuration

[1798] 1. Audio assistants

[1799] A sophisticated audio assistant is a device that reduces external noise and provides clear audio to users, and can be connected to a mobile communication device via a wireless communication interface such as Bluetooth.

[1800] 2. Mobile Communication Devices

[1801] The mobile communication device is a portable electronic device such as a smartphone or tablet. It is connected to the audio assistant via Bluetooth and has installed an application for collecting, processing, and displaying audio data.

[1802] Software Configuration

[1803] 1. Means of communication

[1804] It is equipped with a wireless communication interface such as Bluetooth for communicating data between the mobile communication device and the audio assistant device.

[1805] 2. Processing means

[1806] This function performs digital signal processing (DSP) on audio data collected by a microphone to eliminate noise.

[1807] 3. Correction means

[1808] This function uses a generative AI model to correct voice data into natural language.

[1809] 4. Display means

[1810] This function converts voice data into text and displays it on the screen of a smartphone or tablet.

[1811] 5. Emotion recognition means

[1812] It has an emotion engine that analyzes voice data to identify the user's emotions.

[1813] 6. Optimization Methods

[1814] This is a function that optimizes other processes based on the emotion information identified by the emotion recognition means.

[1815] Specific examples

[1816] Use at family dinners

[1817] User: Places smartphone in the center of the table, connects audio assistant via Bluetooth, adjusts equalization and noise cancellation settings through the app, and enables emotion engine.

[1818] Device (smartphone): Collects surrounding conversations with a microphone and transmits them to the voice assistant in real time. The voice data is sent to a server to obtain text data, which is then displayed on the screen.

[1819] Emotion Engine: Analyzes voice data and recognizes user emotions. Dynamically adjusts voice filtering parameters based on the recognition.

[1820] Server: Analyzes the voice data, generates text in real time, and sends it back to the smartphone.

[1821] Terminal (smartphone): Saves the text data and sends the correction results to the voice assistant device.

[1822] Prompt Sentence Examples

[1823] An example of a prompt a user can send to a generative AI model is one that includes the instruction, "Please remove the noise from this audio data and correct it into natural language that is easy to listen to."

[1824] The above examples of specific use cases and prompt sentences will help understand the detailed implementation of the system of the present invention.

[1825] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1826] Program processing flow

[1827] Step 1: Bluetooth connection

[1828] The user launches the smartphone app and sets the audio assistant device to Bluetooth pairing mode, which allows the smartphone to recognize the audio assistant device.

[1829] The device (smartphone) searches for Bluetooth devices and displays a list of nearby Bluetooth devices. The user can select an audio assistant device from the list.

[1830] The terminal (smartphone) establishes Bluetooth pairing with the selected device. The input is the user's operation, and the output is the establishment of pairing.

[1831] Step 2: Collect and send audio data

[1832] The device (smartphone) collects ambient sound in real time using a built-in microphone, stores it in a buffer, and then uses digital signal processing (DSP) to reduce noise.

[1833] The terminal (smartphone) transmits the audio data stored in the buffer to the audio assistant device via Bluetooth. The input is the audio data collected by the microphone, and the output is the data in the buffer after DSP processing.

[1834] The audio assistant decodes the received audio data, adjusts the volume and frequency characteristics to suit the audio, and plays it back. The input is the audio data sent from the terminal, and the output is clear audio provided to the user.

[1835] Step 3: Equalizing and noise reduction

[1836] Users can adjust equalization and noise cancellation parameters in the settings screen of the smartphone app, using sliders and presets to customize the sound quality.

[1837] The device (smartphone) updates the filtering algorithm based on the user's settings and applies new parameters. The input is the user's adjusted parameters, and the output is the updated filtering algorithm.

[1838] The terminal (smartphone) transmits the filtered audio data to the audio assistance device via Bluetooth.

[1839] The audio assistant then readjusts the filtered audio data and plays it back optimally, with the input being the filtered audio data and the output being audio optimized for the user.

[1840] Step 4: Transcribe and save the conversation

[1841] The device (smartphone) transmits the voice data of the conversation to the server in real time in streaming format. The input is the voice data collected in real time, and the output is the data transmitted to the server.

[1842] The server analyzes the received voice data and converts it into text data. The input is the voice data sent from the terminal, and the output is text data.

[1843] The device (smartphone) displays the text data returned from the server on the app screen and saves it in its internal storage. The input is the text data from the server, and the output is the displayed and saved text data.

[1844] Step 5: Language correction by generative AI

[1845] The device (smartphone) sends inaudible voice data to the generative AI model. The input is voice data, and the output is data sent to the server.

[1846] The server analyzes the voice data received by the generative AI model and corrects it into natural language. The input is the voice data sent from the device, and the output is the corrected language data.

[1847] The terminal (smartphone) receives the corrected data and plays it on the audio assist device. The input is the corrected data from the server, and the output is the played audio.

[1848] Step 6: Emotion Recognition with the Emotion Engine

[1849] The device (smartphone) transmits voice data collected by a microphone to the emotion engine in real time. The input is the collected voice data, and the output is data sent to the emotion engine.

[1850] The emotion engine analyzes the received voice data and identifies the user's emotion. The input is the voice data, and the output is the identified emotion information.

[1851] The device (smartphone) optimizes the voice data settings based on feedback from the emotion engine. The input is the emotion engine feedback, and the output is the optimized voice data.

[1852] Through the above steps, the system can provide users with comfortable voice interaction.

[1853] (Application example 2)

[1854] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1855] While conventional voice assistance systems can filter voice data and eliminate noise, they lack the ability to recognize the user's emotions and optimize voice data accordingly. This has resulted in issues with not being able to provide appropriate voice feedback in certain situations. Furthermore, they lacked the ability to correct speech to natural language using generative AI and real-time data sharing, making effective security management difficult.

[1856] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server is equipped with an emotion recognition engine and includes means for analyzing voice data in real time, means for correcting voice data using a generation AI, and means for dynamically adjusting filtering parameters. This makes it possible to recognize the user's emotions and optimize voice data based on them. In addition, voice data can be sent to the server in real time, allowing appropriate feedback to be obtained immediately, thereby achieving more effective security management.

[1857] A "high-performance audio assist device" is a device that reduces external noise and provides clear audio.

[1858] A "mobile communication device" is a portable electronic device such as a smartphone or tablet.

[1859] "Communication means" refers to the function of collecting and transmitting voice data and exchanging data between devices.

[1860] The "processing means" is a function that performs digital signal processing such as noise cancellation on audio data.

[1861] The "correction means" is a function that corrects voice data into natural language using generative artificial intelligence.

[1862] The "display means" is a function that converts voice data into text and displays it visually.

[1863] The "emotion recognition means" is a function that recognizes the user's emotions from the voice data and dynamically adjusts the voice filtering parameters accordingly.

[1864] "Means for transmitting voice data in real time" refers to a function for transmitting voice data in real time and for immediate processing and feedback.

[1865] The "equalizing means" is a function that adjusts audio characteristics based on equalizing parameters that can be set by the user.

[1866] The present invention relates to a system that uses a highly functional voice assistant device and a mobile communication device electrically connected to it to effectively process and optimize voice data in a user's environment and apply it to security services. The following describes the specific system configuration and its operation method.

[1867] System configuration

[1868] 1. Hardware

[1869] High-performance audio assist device: Reduces external noise and provides clear audio to users. Connects to mobile communication devices using wireless communication interfaces such as Bluetooth.

[1870] Mobile Communication Device: A portable electronic device, such as a smartphone or tablet, that collects, processes, and displays audio data. It contains a microphone and a Bluetooth module.

[1871] 2. Software

[1872] Bluetooth module: Wirelessly connects the audio assistant device to a smartphone, sending and receiving audio data.

[1873] Digital signal processing (DSP): Performs noise cancellation, equalization, and other processing on collected audio data.

[1874] Emotion engine: Analyzes voice data and recognizes the user's emotions. As a concrete example, it uses a natural language processing library (e.g., Google TensorFlow).

[1875] Generative AI: Generative AI analyzes voice data and corrects it into natural language.

[1876] Display function: Converts voice data into text and displays it visually.

[1877] Operation Overview

[1878] 1. Bluetooth connection

[1879] Server: Pairs the audio assistant device with the mobile communication device via Bluetooth.

[1880] Device: Search for devices, select the audio assistant and establish pairing.

[1881] 2. Collection and transmission of voice data

[1882] Device: Surrounding sounds are collected using the smartphone's microphone and stored in a buffer in real time.

[1883] Terminal: The audio data in the buffer is filtered using digital signal processing and sent to the audio assistant device via Bluetooth.

[1884] Audio assistant: Decodes received audio data and adjusts it to the optimum volume and frequency characteristics.

[1885] 3. Equalization and noise cancellation

[1886] User: Adjust equalization and noise cancellation parameters in the settings screen of the smartphone app.

[1887] Device: Receives user settings, updates the filtering algorithm, and sends it to the audio assistant device.

[1888] 4. Transcribe and save conversations

[1889] Terminal: Sends audio data to the server in real time in streaming format.

[1890] Server: Analyzes the voice data, generates text data, and sends it back to the device.

[1891] Device: Text data is displayed in the app and saved in the internal storage.

[1892] 5. Language Correction by Generative AI

[1893] Device: Sends inaudible voice data to the generating AI.

[1894] Server: The generative AI analyzes the voice data and corrects it into natural language.

[1895] Terminal: Receives the corrected data and plays it on the audio assistant device.

[1896] 6. Emotion Recognition by Emotion Engine

[1897] Terminal: Voice data collected by the microphone is sent to the emotion engine.

[1898] Server: The emotion engine analyzes the voice data and recognizes the user's emotions. This information is reflected in the filtering parameters and generation AI.

[1899] Device: Optimized voice data settings based on feedback from the emotion engine.

[1900] Specific examples

[1901] Office Security Monitoring

[1902] The user acts as a security manager, placing a smartphone in the security room and pairing it with a voice assistant. The system analyzes conversations in the office in real time and can send alerts to managers if it detects tension or stress.

[1903] Prompt Sentence Examples

[1904] "What are the potential emotions that could be present in this scene? Identify these emotions based on the text data. Then update the emotion parameters in the emotion engine and apply the new security settings."

[1905] Thus, the present invention can be implemented in a variety of security environments, and specific hardware and software can be used to achieve effective security management.

[1906] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1907] Step 1:

[1908] Device: Set the audio assistant device to Bluetooth pairing mode, launch the app on the smartphone and search for devices, select the audio assistant device from the search results, and establish pairing. The input is Bluetooth device information, and the output is the status of pairing establishment.

[1909] Step 2:

[1910] Terminal: The smartphone's microphone collects surrounding audio and stores it in a buffer in real time. The input is the surrounding audio data, and the output is the collected audio data in the buffer. Specifically, the smartphone's built-in microphone captures audio, converts the data into digital format, and stores it in the buffer.

[1911] Step 3:

[1912] Terminal: The audio data in the buffer is filtered using a digital signal processing algorithm and sent to the audio assistant via Bluetooth. The input is the audio data in the buffer and the output is the filtered audio data. The audio data is noise-reduced and equalized and sent via the Bluetooth module.

[1913] Step 4:

[1914] Audio assistant: Decodes received audio data and plays it back after adjusting the volume and frequency characteristics to suit the user. The input is filtered audio data, and the output is clear audio. Specifically, the decoding process converts the digital signal into an analog signal, which is then output through a speaker.

[1915] Step 5:

[1916] User: Adjusts equalization and noise cancellation parameters on the settings screen within the smartphone app. The input is the user's setting parameters, and the output is the updated filtering algorithm. The app's UI intuitively manipulates sliders and drop-down menus to change settings.

[1917] Step 6:

[1918] Terminal: Receives user settings, updates the filtering algorithm, and sends the new algorithm to the audio assistant. The input is the updated filtering parameters, and the output is the new filtering algorithm applied to the audio assistant.

[1919] Step 7:

[1920] Terminal: Sends the voice data of the conversation to the server in real time in streaming format. The input is filtered voice data, and the output is streaming data. The data is securely sent to the server via the SSL / TLS protocol.

[1921] Step 8:

[1922] Server: Analyzes the voice data, generates text data, and sends it back to the smartphone. The input is streaming voice data, and the output is the generated text data. The voice recognition engine analyzes the data and converts it into text format.

[1923] Step 9:

[1924] Terminal: Text data is displayed in the app and saved to the internal storage. The input is the text data returned from the server, and the output is the displayed text and saved data. Specifically, it is displayed in the app's text view and saved to the database.

[1925] Step 10:

[1926] Device: Sends inaudible voice data to the generation AI. The input is the collected voice data, and the output is the request data for correction. The API of the generation AI model is called and the voice data is sent.

[1927] Step 11:

[1928] Server: The generative AI analyzes the voice data and corrects it into natural language. The input is the voice data, and the output is the corrected voice data. The generative AI model removes noise and converts it into clear pronunciation.

[1929] Step 12:

[1930] Terminal: Receives the corrected data and plays it back to the audio assistant. The input is the corrected data returned from the server, and the output is the reproduced natural speech.

[1931] Step 13:

[1932] Terminal: Sends voice data collected by a microphone to the emotion engine. The input is the collected voice data, and the output is the emotion recognition request data.

[1933] Step 14:

[1934] Server: The emotion engine analyzes the voice data and recognizes the user's emotions. The input is the voice data, and the output is the recognized emotion information. The emotion data is analyzed through the analytics engine.

[1935] Step 15:

[1936] Terminal: Optimizes voice data settings based on feedback from the emotion engine. The input is the recognized emotion information, and the output is optimized filtering parameters.

[1937] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1938] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1939] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1940] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1941] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1942] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1943] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1944] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1945] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1946] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1947] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1948] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1949] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1950] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1951] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1952] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1953] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1954] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1955] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1956] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1957] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1958] The following is further disclosed regarding the above embodiment.

[1959] (Claim 1)

[1960] High-performance audio assist devices and

[1961] a mobile communication device electrically connected to the audio assistance device;

[1962] a communication means for customizing voice data in the mobile communication device;

[1963] a processing means for performing noise cancellation processing on the audio data;

[1964] a correcting means having a generating artificial intelligence for correcting the voice data into natural language;

[1965] a display means for converting the voice data into text and displaying it;

[1966] A system including:

[1967] (Claim 2)

[1968] 2. The system of claim 1, further comprising a communication means for transmitting the audio data in real time from the mobile communication device to the audio assistant device.

[1969] (Claim 3)

[1970] 10. The system of claim 1, further comprising user-configurable equalization means in the mobile communication device.

[1971] "Example 1"

[1972] (Claim 1)

[1973] High-performance sound amplification equipment,

[1974] a portable information terminal electrically connected to the sound amplifier;

[1975] a communication means for customizing voice data in the portable information terminal;

[1976] a processing means for performing noise removal processing on the audio data;

[1977] a correcting means having a generating artificial intelligence for correcting the voice data into natural language;

[1978] a display means for converting the voice data into text and displaying the text;

[1979] the portable information terminal stores the audio data in a buffer;

[1980] a transmitting means for transmitting the audio data to the audio amplifier in real time via wireless communication;

[1981] setting means for allowing the user to adjust equalization and noise reduction settings on the mobile information terminal;

[1982] A system including:

[1983] (Claim 2)

[1984] 2. The system according to claim 1, further comprising a communication means for transmitting audio data from said portable information terminal to said sound amplifier in real time via said wireless communication.

[1985] (Claim 3)

[1986] 2. The system according to claim 1, further comprising an equalizing means in said portable information terminal, the setting of which can be changed by a user.

[1987] "Application Example 1"

[1988] (Claim 1)

[1989] High-performance audio assist devices and

[1990] a mobile communication device electrically connected to the audio assistance device;

[1991] a communication means for customizing voice data in the mobile communication device;

[1992] a processing means for performing noise cancellation processing on the audio data;

[1993] a correcting means having a generating artificial intelligence for correcting the voice data into natural language;

[1994] a display means for converting the voice data into text and displaying it;

[1995] a voice output means for reproducing the generated text data;

[1996] a server communication means for transmitting the voice data to an external server and acquiring corrected text data;

[1997] A system including:

[1998] (Claim 2)

[1999] 2. The system of claim 1, further comprising a communication means for transmitting the audio data in real time from the mobile communication device to the audio assistant device.

[2000] (Claim 3)

[2001] 10. The system of claim 1, further comprising a user-configurable equalizing means and a means for playing back generated text data in the mobile communication device.

[2002] "Example 2: Combining Emotion Engines"

[2003] (Claim 1)

[2004] High-performance audio assist devices and

[2005] a mobile communication device electrically connected to the audio assistance device;

[2006] a communication means for customizing voice data in the mobile communication device;

[2007] a processing means for performing noise cancellation processing on the audio data;

[2008] a correcting means having a generating artificial intelligence for correcting the voice data into natural language;

[2009] a display means for converting the voice data into text and displaying it;

[2010] emotion recognition means for analyzing the voice data to identify the emotion of the user;

[2011] an optimization means for optimizing other processes based on the emotion information identified by the emotion recognition means;

[2012] A system including:

[2013] (Claim 2)

[2014] 2. The system of claim 1, further comprising a communication means for transmitting the audio data in real time from the mobile communication device to the audio assistant device.

[2015] (Claim 3)

[2016] 10. The system of claim 1, further comprising user-configurable equalization means in the mobile communication device.

[2017] "Application example 2 when combining emotion engines"

[2018] (Claim 1)

[2019] High-performance audio assist devices and

[2020] a mobile communication device electrically connected to the audio assistance device;

[2021] a communication means for customizing voice data in the mobile communication device;

[2022] a processing means for performing noise cancellation processing on the audio data;

[2023] a correcting means having a generating artificial intelligence for correcting the voice data into natural language;

[2024] a display means for converting the voice data into text and displaying it;

[2025] emotion recognition means for recognizing emotions from the voice data and dynamically adjusting voice filtering parameters based on the recognized emotion information;

[2026] means for transmitting voice data to a server in real time in the mobile communication device and displaying and storing data returned from the server;

[2027] A system including:

[2028] (Claim 2)

[2029] 2. The system of claim 1, further comprising a communication means for transmitting the audio data in real time from the mobile communication device to the audio assistant device.

[2030] (Claim 3)

[2031] 10. The system of claim 1, further comprising user-configurable equalization means in the mobile communication device. [Explanation of symbols]

[2032] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. High-performance audio assist devices and a mobile communication device electrically connected to the audio assistance device; a communication means for customizing voice data in the mobile communication device; a processing means for performing noise cancellation processing on the audio data; a correcting means having a generating artificial intelligence for correcting the voice data into natural language; a display means for converting the voice data into text and displaying it; A system including:

2. 2. The system of claim 1, further comprising communication means for transmitting said voice data in real time from said mobile communication device to said voice assistant device.

3. 10. The system of claim 1, further comprising user-configurable equalizing means in said mobile communication device.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A