System

A system using real-time digital signal processing and generative AI to distinguish and cancel abnormal sounds during online meetings addresses the stress caused by sneezing, hiccups, and coughs, ensuring a peaceful meeting environment.

JP2026014939APending Publication Date: 2026-01-29SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024116413
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-19
Publication Date
2026-01-29

AI Technical Summary

Technical Problem

Physiological phenomena such as sneezing, hiccups, and coughs during online meetings can be stressful for individuals with conditions like hay fever or Tourette's syndrome, and existing technologies do not effectively eliminate these unwanted sounds in real-time.

Method used

A system that uses a microphone to capture voice signals, applies digital signal processing to extract features, employs generative AI to distinguish between normal and abnormal sounds, generates an inverted audio signal to cancel out unpleasant sounds, and outputs the processed audio in real-time.

Benefits of technology

Enables participants to engage in online meetings without others hearing unpleasant sounds, maintaining a comfortable communication environment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026014939000001_ABST
    Figure 2026014939000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system, comprising: means for acquiring a speech signal; generative artificial intelligence means for analyzing the acquired speech signal and distinguishing between normal speech and abnormal speech; means for generating an antiphase speech signal if abnormal speech is detected; means for mixing the original speech signal and the antiphase speech signal to cancel unpleasant speech; and means for outputting the speech signal.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] In today's world where online meetings have become commonplace, physiological phenomena such as sneezing, hiccups, and coughs that occur during meetings have become a problem. For people with conditions such as hay fever, asthma, or Tourette's syndrome, hearing unwanted sounds can be particularly stressful and can be a barrier to participating in meetings. To solve this problem, technology is needed that can automatically eliminate these unpleasant sounds and provide an environment where people can participate in online meetings with peace of mind. [Means for solving the problem]

[0005] To solve the above problems, the present invention provides the following means. A system is provided that includes a means for acquiring a voice signal, a generating artificial intelligence means for analyzing the acquired voice signal and distinguishing between normal voice and abnormal voice, a means for generating an inverted voice signal when an abnormal voice is detected, a means for mixing the original voice signal with the inverted voice signal to cancel out the unpleasant voice, and a means for outputting the voice signal. This system allows users to participate in online meetings with peace of mind without exposing other participants to unpleasant voices. Furthermore, this system can process voice signals in real time, constantly detecting abnormal voices and generating inverted voices, eliminating the need for user operation. Furthermore, this system extracts voice features using Mel-frequency cepstrum coefficients, enabling more accurate detection of abnormal voices.

[0006] An "audio signal" is audio data expressed as an electrical signal.

[0007] The "means for acquiring" is a mechanism for capturing audio signals using a microphone, a sensor, or the like.

[0008] "Generative artificial intelligence means" refers to technology that uses machine learning and deep learning to analyze voice data and identify specific patterns and features.

[0009] "Analyzing" means interpreting the acquired voice data through digital signal processing and feature extraction.

[0010] "Normal voice" is everyday voice such as speech or conversation, and is a voice signal containing intended content.

[0011] "Abnormal voice" refers to a voice signal generated by an unintentional physiological phenomenon such as sneezing, coughing, or hiccuping.

[0012] "Distinguishing" means analyzing the acquired audio signal and identifying whether it is a normal audio signal or an abnormal audio signal.

[0013] "Detect" means to recognize that an abnormal sound has occurred.

[0014] An "out-of-phase audio signal" is an audio signal that is 180 degrees out of phase with the original audio signal, and is used to cancel the original audio signal.

[0015] "Generating means" refers to a function or mechanism that generates a new audio signal or an inverted audio signal through necessary processing.

[0016] "Mixing" refers to the process of synthesizing multiple audio signals to create one audio signal.

[0017] "Cancelling" means reducing or eliminating unpleasant sound components.

[0018] The "means for outputting" is a mechanism for transmitting the processed audio signal to the outside via a speaker or a communication device.

[0019] "Real-time processing means" refers to methods or techniques that process audio signals immediately and produce results without delay.

[0020] "Mel-frequency cepstrum coefficients" are specific statistics that represent the characteristics of a speech signal, and are features that are often used in speech recognition and abnormal speech detection.

[0021] A "feature" is data that numerically represents the characteristics of an audio signal and is used for classification and analysis. [Brief explanation of the drawings]

[0022] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3]FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0023] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0024] First, the terms used in the following description will be explained.

[0025] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0026] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0027] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0028] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0029] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0030] [First embodiment]

[0031] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0032] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0033] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0034] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0035] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0036] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0037] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0038] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0039] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0040] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0041] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0042] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0043] This invention relates to a system that detects physiological phenomena such as sneezing, coughing, and hiccups that occur during online meetings and automatically mutes them. This system acquires audio signals in real time, uses generation AI to distinguish between normal and abnormal sounds, and if an abnormal sound is detected, generates and mixes an opposite-phase audio signal to cancel out the unpleasant sound.

[0044] System configuration

[0045] 1. Audio acquisition method

[0046] Terminal: This system is equipped with a microphone to capture the user's voice in real time. The microphone continuously captures the voice signal.

[0047] 2. Voice Analysis Methods

[0048] Terminal: The captured audio signal is analyzed using digital signal processing techniques, which apply noise reduction algorithms, divide it into frames, and perform a Fourier transform to extract the audio frequency components.

[0049] 3. Feature Extraction Method

[0050] Terminal: Calculate features such as Mel-Frequency Cepstrum Coefficients (MFCC) from the preprocessed speech data and generate a feature vector, which represents the characteristics of the speech as numerical data.

[0051] 4. Abnormal sound detection method

[0052] Terminal: The feature vector is input into the generation AI to distinguish between normal and abnormal voices. In particular, when abnormal voices such as sneezing, coughing, and hiccups are detected, the start and end times are recorded.

[0053] 5. Antiphase speech generation means

[0054] Terminal: When an abnormal audio is detected, it generates an audio signal with the opposite phase to the abnormal audio portion. The opposite phase audio signal has the phase information necessary to cancel out the original audio signal.

[0055] 6. Audio Mixing Methods

[0056] Terminal: Mix the original audio signal with an inverted audio signal to cancel out any unwanted sounds. This operation ensures that the final audio is output smoothly.

[0057] 7. Audio output means

[0058] Terminal: The processed audio signal is sent to the online meeting application, which prevents other participants from hearing the disturbing audio.

[0059] Specific examples

[0060] Situation:

[0061] User A suddenly sneezes during an online meeting.

[0062] 1. Audio capture

[0063] Device: The microphone captures user A's voice in real time and stores it in a buffer.

[0064] 2. Voice Analysis

[0065] Terminal: Noise reduction is applied to the acquired audio data, it is divided into frames, and a Fourier transform is performed.

[0066] 3. Feature Extraction

[0067] Terminal: Calculate Mel-frequency cepstrum coefficients from the audio data and generate a feature vector.

[0068] 4. Abnormal Audio Detection

[0069] Terminal: The feature vector is input into the generative AI, which detects that the sneeze occurred between 1.5 and 1.8 seconds.

[0070] 5. Antiphase speech generation

[0071] Terminal: Generates an audio signal in anti-phase with a sneeze sound between 1.5 and 1.8 seconds.

[0072] 6. Audio Mixing

[0073] Terminal: Mixes the original audio signal with an out-of-phase audio signal to cancel the sneeze sound.

[0074] 7. Audio Output

[0075] Device: Sends the processed audio signal to the online meeting application, where other participants hear the audio without hearing any sneeze sounds.

[0076] Using this system, users can participate in online meetings with peace of mind and prevent unintended audio from being heard by others.

[0077] The processing flow will be explained below.

[0078] Step 1: Getting voice input

[0079] On the device: The microphone is activated and the user's voice is captured in real time. The captured voice data is stored in a buffer.

[0080] Step 2: Preprocessing the audio data

[0081] Terminal: A noise reduction algorithm is applied to the acquired audio data. The audio data is then divided into short time frames, and frequency components are extracted by performing a Fourier transform on each frame.

[0082] Step 3: Extract audio features

[0083] Terminal: To extract features from the preprocessed speech data, Mel-Frequency Cepstral Coefficients (MFCCs) are calculated, which generates a feature vector that quantifies the speech features.

[0084] Step 4: Classify normal and abnormal voices

[0085] Terminal: The generated feature vector is input into a generative artificial intelligence (generative AI) to classify normal and abnormal speech. This AI identifies patterns in the speech signal based on past training data and detects abnormal speech, such as sneezing or coughing.

[0086] Step 5: Detecting Audio Anomalies

[0087] Terminal: When the artificial intelligence detects an abnormal sound, it records its start and end time. For example, if a certain part of the audio signal is recognized as a sneeze, that time frame will be marked.

[0088] Step 6: Generate an anti-phase audio signal

[0089] Terminal: Generates an audio signal with an inverse phase to the detected abnormal audio portion. The inverse phase audio signal is an audio signal with a phase difference of 180 degrees to cancel the original abnormal audio signal.

[0090] Step 7: Mixing the audio signals

[0091] Terminal: Mix the original audio signal with the generated out-of-phase audio signal. This mixing cancels out the unpleasant abnormal sounds.

[0092] Step 8: Output the audio signal

[0093] Terminal: Sends the processed audio data to the audio stream of the online meeting, so that other participants hear the audio with the abnormal audio canceled.

[0094] Step 9: Real-time monitoring and adjustment

[0095] Terminal: All processes from audio signal acquisition to output are repeated in real time, continuously monitoring for abnormal audio and canceling it as appropriate.

[0096] In this way, the system of the present invention can automatically cancel unpleasant sounds such as sneezing and coughing that occur during an online meeting in real time, allowing users to maintain smooth conversations.

[0097] Example 1

[0098] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0099] During online meetings, physiological phenomena such as sneezing, coughing, and hiccups can interrupt conversations and cause discomfort to other participants. Because these sounds are not automatically muted, users may unintentionally let other participants hear the unpleasant sounds.

[0100] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0101] In this invention, the server includes means for acquiring voice data, means for preprocessing the acquired voice data using digital signal processing technology, means for extracting voice features from the preprocessed voice data, means for inputting the extracted voice features into a generative artificial intelligence model to distinguish between normal voice and abnormal voice, means for generating out-of-phase voice data when abnormal voice is detected, means for mixing the original voice data with the out-of-phase voice data to cancel unpleasant voice, and means for outputting the processed voice data. This prevents unpleasant sounds caused by physiological phenomena such as sneezing, coughing, and hiccups from reaching other participants during online meetings, enabling comfortable communication.

[0102] A "means for acquiring voice data" is a device or method for digitally capturing a user's speech or voice in real time using hardware such as a microphone.

[0103] "Digital signal processing" is the process of converting analog audio signals into digital form and using techniques such as noise reduction and Fourier transform to improve the quality and characteristics of the audio.

[0104] "Speech features" are numerical representations of the characteristics of a speech signal, and are typically extracted using techniques such as Mel-Frequency Cepstrum Coefficients (MFCC).

[0105] A "generative artificial intelligence model" is an algorithm or network that is trained to use machine learning or deep learning techniques to analyze speech data and distinguish between normal and abnormal speech.

[0106] "Out-of-phase audio data" is audio data that has the opposite phase to the original audio signal and is generated to cancel or reduce audio by mixing it with the original audio signal.

[0107] The "means for outputting audio data" refers to a device or method for transferring the processed audio data to a user's terminal or an online meeting application.

[0108] "Abnormal voices" are sounds caused by noises or physiological phenomena that are not part of normal conversation, such as sneezing, coughing, or hiccuping, and are detected by the generative artificial intelligence model.

[0109] This invention relates to a system that detects physiological phenomena such as sneezing, coughing, and hiccups that occur during online meetings and automatically mutes them. This system acquires audio signals in real time, uses a generative AI model to distinguish between normal and abnormal sounds, and when abnormal sounds are detected, generates and mixes an opposite-phase audio signal to cancel out the unpleasant sounds.

[0110] The system includes the following components:

[0111] 1. How to obtain audio data

[0112] Terminal: This system is equipped with a microphone to capture the user's voice in real time. Voice data is continuously acquired from the microphone.

[0113] 2. Preprocessing using digital signal processing

[0114] Terminal: The captured audio data is pre-processed using digital signal processing techniques, which apply noise reduction algorithms, divide the audio data into frames, perform Fourier transforms, and extract frequency components.

[0115] 3. Methods for extracting speech features

[0116] Terminal: Calculate Mel-Frequency Cepstral Coefficients (MFCC) from the preprocessed speech data to generate a feature vector, which represents the characteristics of the speech as numerical data.

[0117] 4. Method for detecting abnormal voice using generative AI model

[0118] On the device: The feature vector is input into a generative AI model to distinguish between normal and abnormal speech. In particular, when abnormal speech such as sneezing, coughing, or hiccuping is detected, the start and end times are recorded.

[0119] 5. Means for generating out-of-phase audio data

[0120] Terminal: When an abnormal audio is detected, it generates audio data with an inverse phase to the abnormal audio portion. The inverse phase audio data has the phase information necessary to cancel out the original audio data.

[0121] 6. Means of mixing audio data

[0122] Terminal: Mix the original audio data with the opposite phase audio data to cancel out the unpleasant sounds. This operation makes the final audio output smooth.

[0123] 7. Means of outputting audio data

[0124] Terminal: The processed audio data is sent to the online meeting application, which prevents other participants from hearing the disturbing audio.

[0125] Specific examples are shown below.

[0126] If user A suddenly sneezes during an online meeting, the following occurs:

[0127] 1. How to obtain audio data

[0128] Device: The microphone captures user A's voice in real time and stores it in a buffer.

[0129] 2. Preprocessing using digital signal processing

[0130] Terminal: Apply noise reduction to the acquired audio data, divide it into frames, and perform a Fourier transform.

[0131] 3. Methods for extracting speech features

[0132] Terminal: Calculate Mel-Frequency Cepstrum Coefficients (MFCC) from the audio data and generate a feature vector.

[0133] 4. Method for detecting abnormal voice using generative AI model

[0134] On the device: The feature vector is fed into a generative AI model to detect that the sneeze occurred between 1.5 and 1.8 seconds.

[0135] 5. Means for generating out-of-phase audio data

[0136] Device: Generates audio data in antiphase to a sneeze sound between 1.5 and 1.8 seconds.

[0137] 6. Means of mixing audio data

[0138] Terminal: Mix the original audio data with out-of-phase audio data to cancel the sneeze sound.

[0139] 7. Means of outputting audio data

[0140] Device: Sends the processed audio data to the online meeting application, where other participants hear the audio without hearing the sneeze.

[0141] This system allows users to participate in online meetings with peace of mind and prevents unintended audio from being heard by others.

[0142] Prompt Sentence Examples

[0143] "Please explain the system that automatically mutes you when you sneeze during an online meeting."

[0144] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0145] Step 1:

[0146] Acquiring audio data

[0147] Device: Uses a microphone to capture the user's voice data in real time.

[0148] Input: User speech and surrounding audio.

[0149] Output: Digital audio data from the microphone.

[0150] What it does: The microphone is always on, capturing audio data in real time and storing it in a fixed buffer.

[0151] Step 2:

[0152] Audio signal preprocessing

[0153] Terminal: Performs digital signal processing on the captured audio data, applying noise reduction algorithms, dividing the audio data into frames, and performing a Fourier transform to extract frequency components.

[0154] Input: Raw audio data.

[0155] Output: Denoised per-frame frequency spectrum.

[0156] What it does: It filters the audio signal to remove noise, then divides it into small time segments (frames), then performs a Fourier transform on each frame to calculate its frequency spectrum.

[0157] Step 3:

[0158] Feature extraction

[0159] Terminal: Calculate the Mel-Frequency Cepstral Coefficients (MFCC) from each frame and generate a feature vector.

[0160] Input: Frequency spectrum obtained by Fourier transform.

[0161] Output: MFCC-based feature vectors representing audio features.

[0162] Specific operation: Calculate MFCC using frequency spectrum as input, and save the calculation result as a feature vector in list format.

[0163] Step 4:

[0164] Abnormal audio detection

[0165] On the device: The feature vector is fed into a generative AI model to distinguish between normal and abnormal speech. If abnormal speech is detected, the start and end times of the speech are recorded.

[0166] Input: MFCC-based feature vector.

[0167] Output: Whether or not there is an abnormal sound, and its start and end times.

[0168] Specific operation: The generative AI model detects abnormal sounds based on the input feature vector and saves the period during which the abnormal sounds occurred as a timestamp.

[0169] Step 5:

[0170] Antiphase audio data generation

[0171] Terminal: Generates audio data in the opposite phase to the time range of the detected abnormal audio.

[0172] Input: Start and end times of the abnormal audio, and the original audio data.

[0173] Output: Out-of-phase audio data.

[0174] Specific operation: Based on the time range of the detected abnormal sound, calculate and save the opposite phase audio waveform.

[0175] Step 6:

[0176] Mixing of audio data

[0177] Terminal: Mixes the original audio data with out-of-phase audio data to cancel abnormal sounds.

[0178] Input: Original audio data, inverted audio data.

[0179] Output: Audio data with abnormal sounds canceled.

[0180] Specific operation: The original audio data is added to the out-of-phase audio data, and a process is performed to cancel out abnormal sounds.

[0181] Step 7:

[0182] Sending audio data

[0183] Terminal: Sends the processed audio data to the online meeting application.

[0184] Input: Mixed audio data.

[0185] Output: Final audio data (audio heard by other participants in the online meeting).

[0186] What it does: Stream the processed audio data in real time to an online meeting application.

[0187] (Application example 1)

[0188] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0189] In conventional online meetings and autonomous vehicle environments, sudden physiological events such as sneezing or coughing can cause unpleasant sounds to those around you and temporarily worsen the air quality, resulting in a decrease in voice quality, hindering conversations and the transmission of instructions, as well as reducing the comfort of the vehicle's interior.

[0190] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0191] In this invention, the server includes means for acquiring an audio signal, artificial intelligence means for analyzing the acquired audio signal and distinguishing between normal and abnormal audio, means for generating an opposite-phase audio signal when an abnormal audio signal is detected, means for mixing the original audio signal with the opposite-phase audio signal to cancel the unpleasant audio, means for outputting the audio signal, and means for temporarily strengthening related systems when canceling a new physiological phenomenon. This not only enables immediate detection and control of abnormal audio, but also temporarily improves the air environment, making it possible to maintain a comfortable and healthy environment.

[0192] plaintext

[0193] The "means for acquiring an audio signal" refers to a microphone and associated hardware for converting sound into an electronic signal.

[0194] The "generative artificial intelligence means" is an algorithm using artificial intelligence to analyze the acquired voice signal and distinguish between normal voice and abnormal voice.

[0195] The "means for generating an inverted phase audio signal" refers to a device or algorithm for generating an audio signal with the phase of an abnormal audio signal inverted when the abnormal audio signal is detected.

[0196] "Means for canceling unpleasant sounds" is a technology that synthesizes an audio signal that is out of phase with the original audio signal to cancel out unpleasant sounds.

[0197] The "means for outputting an audio signal" refers to a device such as a speaker or earphone for reproducing the processed audio signal.

[0198] The "means for temporarily enhancing related systems" is a control device for temporarily enhancing the air filter system and other related systems in the vehicle when a physiological phenomenon is detected.

[0199] plaintext

[0200] This invention is a system primarily designed to improve the riding experience in self-driving vehicles. It detects sudden physiological phenomena such as sneezing and coughing, cancels the unpleasant sounds that occur during these events, and temporarily improves the in-car environment.

[0201] The elements of the server, terminal, and user are as follows:

[0202] server

[0203] The vehicle is equipped with a microphone that captures audio signals and a computer system that includes an artificial intelligence algorithm for analyzing the audio. Audio signals are captured in real time and analyzed using digital signal processing technology. If an abnormal sound is detected, an opposite-phase audio signal is generated, thereby canceling out the unpleasant sound. This detection also temporarily strengthens the cabin air filter system.

[0204] Terminal

[0205] The terminal is a component of the in-vehicle system that acquires and analyzes audio signals in real time. Specifically, it performs noise reduction on the audio signal, frequency analysis using Fourier transform, and calculation of Mel Frequency Cepstrum Coefficients (MFCC). It also uses generative artificial intelligence to distinguish between normal and abnormal audio, and generates and mixes an inverse audio signal for detected abnormal audio to cancel it. It also includes a device that outputs the processed audio signal.

[0206] User

[0207] As passengers in self-driving vehicles, users can benefit from the ability to avoid physiological events such as sneezing or coughing while in an online meeting or using the navigation system, causing discomfort to other passengers and the system.

[0208] Specific examples

[0209] For example, the following prompt sentences can be used to input to a generative AI model to recreate a situation in which the present invention is implemented:

[0210] "Detect the sound of a sneeze in the car, mute the audio system at that moment, and strengthen the air filtration system."

[0211] This allows the in-car environment to be controlled quickly and automatically, providing a comfortable riding experience.

[0212] Hardware and Software

[0213] The system utilizes the following hardware and software:

[0214] Hardware: microphones, on-board computers, speakers, air filtration systems.

[0215] Software: Noise reduction algorithm, Fourier transform algorithm, Mel-Frequency Cepstral Coefficient (MFCC) extraction algorithm, generative AI model (TensorFlow, PyTorch, etc.).

[0216] By using the above prompt sentences, the generative AI model can immediately detect abnormal sounds and improve the air quality inside the car while canceling out unpleasant sounds.

[0217] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0218] plaintext

[0219] Step 1:

[0220] The server captures audio signals in real time from microphones inside the vehicle. The captured analog audio signals are converted to digital signals and passed to the next processing step. This input includes various sounds, such as ambient sounds inside the vehicle and conversation sounds.

[0221] Step 2:

[0222] The terminal performs noise reduction filtering on the captured voice signal, specifically using digital signal processing techniques to remove unwanted background noise, resulting in a clean voice signal being output.

[0223] Step 3:

[0224] The terminal divides the noise-reduced speech signal into frames and performs a Fourier transform to extract frequency components. The input is the speech signal after noise reduction, and the output is the frequency spectrum.

[0225] Step 4:

[0226] The device extracts speech features by calculating Mel-Frequency Cepstrum Coefficients (MFCC) from the frequency spectrum. This method expresses speech characteristics as numerical data. The input is the frequency spectrum, and the output is an MFCC feature vector.

[0227] Step 5:

[0228] The device inputs the MFCC feature vectors into a generative AI model (e.g., a model using TensorFlow or PyTorch) to distinguish between normal and abnormal speech. If abnormal speech (such as sneezing or coughing) is detected in this step, the device outputs the start and end times of the speech.

[0229] Step 6:

[0230] The terminal generates an audio signal with an opposite phase to the detected abnormal audio. To generate the opposite phase audio signal, an operation is performed to invert the phase of the original audio signal. The input is the audio signal of the abnormal audio, and the output is the audio signal with the opposite phase.

[0231] Step 7:

[0232] The terminal mixes the original audio signal with the generated out-of-phase audio signal, thereby performing a calculation to cancel the abnormal audio. The input is the original audio signal and the out-of-phase audio signal, and the output is the cancelled audio signal.

[0233] Step 8:

[0234] The device outputs the processed audio signal through the vehicle's speakers, allowing other passengers to hear the audio with the unpleasant sounds canceled out.

[0235] Step 9:

[0236] When an abnormal sound is detected, the server triggers a temporary strengthening of the air filter system inside the vehicle. Specifically, it increases the airflow of the air filter to quickly remove allergens and viruses in the air. This operation improves the air environment inside the vehicle.

[0237] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0238] This invention relates to a system that detects physiological phenomena such as sneezing, coughing, and hiccups that occur during online meetings and automatically mutes them. This system acquires audio signals in real time and uses generative AI to distinguish between normal and abnormal sounds. If abnormal sounds are detected, it generates and mixes an opposite-phase audio signal to cancel out the unpleasant sounds. Furthermore, by combining it with an emotion engine that recognizes the user's emotions, it has the ability to identify the user's emotional state and provide appropriate feedback in addition to processing the audio signals.

[0239] System configuration

[0240] 1. Audio acquisition method

[0241] Device: Equipped with a microphone for capturing audio signals in real time. Audio signals are continuously captured from this microphone.

[0242] 2. Voice Analysis Methods

[0243] Terminal: The captured audio data is analyzed using digital signal processing techniques, which apply a noise reduction algorithm, divide it into frames, and perform a Fourier transform to extract the frequency components of the audio.

[0244] 3. Feature Extraction Method

[0245] Terminal: To extract features from the preprocessed speech data, Mel-Frequency Cepstral Coefficients (MFCCs) are calculated, which generates a feature vector that quantifies the speech features.

[0246] 4. Classification of normal and abnormal voices

[0247] Terminal: The generated feature vector is input into the generative artificial intelligence (generative AI) to classify normal and abnormal voices. The generative AI identifies patterns in the voice signal based on past training data and detects abnormal voices such as sneezing and coughing.

[0248] 5. Abnormal Audio Detection

[0249] Terminal: When an abnormal sound is detected, its start and end time are recorded. For example, if a certain part of the sound signal is recognized as a sneeze, that time frame is marked.

[0250] 6. Generation of anti-phase audio signals

[0251] Terminal: Generates an audio signal with an inverse phase to the detected abnormal audio portion. The inverse phase audio signal is an audio signal with a phase difference of 180 degrees to cancel the original abnormal audio signal.

[0252] 7. Mixing of audio signals

[0253] Terminal: Mix the original audio signal with the generated out-of-phase audio signal. This mixing cancels out the unpleasant abnormal sounds.

[0254] 8. Audio signal output

[0255] Terminal: Sends the processed audio data to the audio stream of the online meeting, so that other participants hear the audio with the abnormal audio canceled.

[0256] 9. Emotion Engine

[0257] Device: The emotion engine has algorithms to identify the emotional state from the user's voice signal. The emotion engine analyzes voice characteristics such as tone, pitch, and rate to identify the emotion the user is feeling.

[0258] Server: Provides different feedback depending on the emotional state identified by the emotion engine. For example, if the server determines that the user is feeling stressed, it displays a message on the screen suggesting that the user relax.

[0259] Specific examples

[0260] Situation:

[0261] User B suddenly sneezes during an online meeting and feels stressed.

[0262] 1. Audio capture

[0263] Device: The microphone captures user B's voice in real time and stores it in a buffer.

[0264] 2. Voice Analysis

[0265] Terminal: Noise reduction is applied to the acquired audio data, it is divided into frames, and a Fourier transform is performed.

[0266] 3. Feature Extraction

[0267] Terminal: Calculate Mel-frequency cepstrum coefficients from the audio data and generate a feature vector.

[0268] 4. Classification of normal and abnormal voices

[0269] On the device: The feature vector is fed into a generative AI to detect that a sneeze occurred within a specific time frame.

[0270] 5. Abnormal Audio Detection

[0271] Device: Detects that the sneeze occurred between 1.5 and 1.8 seconds and marks that time frame.

[0272] 6. Generation of anti-phase audio signals

[0273] Device: Generates an audio signal in antiphase to a sneeze sound lasting 1.5 to 1.8 seconds.

[0274] 7. Mixing of audio signals

[0275] Terminal: Mixes the original audio signal with an out-of-phase audio signal to cancel the sneeze sound.

[0276] 8. Audio signal output

[0277] Terminal: Sends the processed audio signal to the online meeting application. Other participants receive the audio without hearing the sneeze.

[0278] 9. Emotion recognition

[0279] Device: The emotion engine identifies User B's emotional state from audio signals including the sneeze sound.

[0280] 10. Emotional Feedback

[0281] Server: If the emotion engine identifies User B's emotional state as "stressed," a message such as "Relax" will be displayed on the screen of the online meeting application.

[0282] Using this system, users can not only automatically cancel unpleasant sounds such as sneezing and coughing during online meetings in real time, but also receive support based on their emotional state using an emotion engine, allowing users to participate in online meetings with greater peace of mind.

[0283] The processing flow will be explained below.

[0284] Step 1: Getting voice input

[0285] On the device: The microphone is activated and the user's voice is captured in real time. The captured voice data is stored in a buffer.

[0286] Step 2: Preprocessing the audio data

[0287] Terminal: A noise reduction algorithm is applied to the acquired audio data to obtain a clear audio signal. The audio data is then divided into short time frames, and a Fourier transform is performed on each frame to extract the frequency components of the audio.

[0288] Step 3: Extract audio features

[0289] Terminal: To extract features from the preprocessed speech data, Mel-Frequency Cepstral Coefficients (MFCCs) are calculated, which generates a feature vector that quantifies the speech features.

[0290] Step 4: Classify normal and abnormal voices

[0291] Terminal: The generated feature vector is input into the generative artificial intelligence (generative AI) to classify normal and abnormal voices. The generative AI identifies patterns in the voice signal based on past training data and detects abnormal voices such as sneezing and coughing.

[0292] Step 5: Detecting Audio Anomalies

[0293] Terminal: When abnormal sounds are detected by the AI ​​generator, their start and end times are recorded. For example, if the audio signal is recognized as a sneeze from 1.5 seconds to 1.8 seconds, that time frame is marked.

[0294] Step 6: Generate an anti-phase audio signal

[0295] Terminal: Generates an audio signal with an inverse phase to the detected abnormal audio portion. The inverse phase audio signal is an audio signal with a phase difference of 180 degrees to cancel the original abnormal audio signal.

[0296] Step 7: Mixing the audio signals

[0297] Terminal: Mix the original audio signal with the generated out-of-phase audio signal. This mixing cancels out the unpleasant abnormal sounds.

[0298] Step 8: Output the audio signal

[0299] Terminal: Sends the processed audio data to the audio stream of the online meeting, so that other participants hear the audio with the abnormal audio canceled.

[0300] Step 9: Emotion Recognition Processing

[0301] On the device: The voice signal is processed by an emotion engine that analyzes characteristics such as tone, pitch, and rate to identify the user's emotional state.

[0302] On the device: The emotion engine identifies the user's emotion based on the analyzed features. For example, if the voice is spoken with a strong tone, high pitch, and rapid speed, it will be identified as stressed or angry.

[0303] Step 10: Emotional State Feedback

[0304] Server: Provides appropriate feedback based on the emotional state identified by the emotion engine. For example, if the server determines that the user is feeling stressed, it displays a message such as "Relax" on the screen of an online meeting application.

[0305] Server: Based on the user's emotional state, it provides a feedback function that notifies other participants of the user's state, depending on the settings.

[0306] In this way, the entire system performs a series of processes in real time, from acquiring voice signals to canceling unpleasant sounds, and then identifying the user's emotions and providing appropriate feedback. This allows users to continue participating in online meetings with peace of mind, even when experiencing physiological or emotional changes.

[0307] Example 2

[0308] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0309] Physiological phenomena such as sneezing, coughing, and hiccups that occur during online meetings can be unpleasant for other participants and disrupt the flow of conversation. Furthermore, users often feel stressed if they become too aware of these physiological phenomena. Furthermore, it is difficult to grasp users' emotional state in real time, and there is a lack of means to provide appropriate feedback. This can reduce the quality of online meetings and impair participants' concentration and comfort.

[0310] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0311] In this invention, the terminal includes means for acquiring an audio signal, artificial intelligence means for analyzing the acquired audio signal and distinguishing between normal audio and abnormal audio, means for generating an opposite-phase audio signal when abnormal audio is detected, means for mixing the original audio signal with the opposite-phase audio signal and canceling unpleasant audio, means for outputting an audio signal, and emotion recognition means for recognizing the user's emotional state and providing feedback. This makes it possible to automatically cancel unpleasant physiological sounds during online meetings and provide support according to the user's emotional state.

[0312] The "audio signal" is a user's speech converted into an electrical signal, and is acquired through an input device such as a microphone.

[0313] "Means for acquiring" refers to a device or method for acquiring audio signals in real time using a microphone, sensor device, etc.

[0314] "Analyzing" is the process of analyzing an audio signal using digital signal processing techniques to extract various features.

[0315] "Normal voice" refers to voice that is not an abnormal voice, such as a normal conversation or speech.

[0316] "Abnormal sounds" refers to physiological sounds that interfere with normal conversation, such as sneezing, coughing, and hiccups.

[0317] "Artificial intelligence tools" are algorithms or models that use historical training data to identify speech patterns and categorize or differentiate between speech sounds.

[0318] "Detecting abnormal sounds" means identifying abnormal sounds from the acquired audio signals and recognizing the time of occurrence and their content.

[0319] An "opposite phase audio signal" is an audio signal generated by reversing the phase of the original abnormal audio signal by 180 degrees, and is used to cancel the abnormal audio.

[0320] "Mixing" is a process of synthesizing an audio signal that is out of phase with the original audio signal to cancel out abnormal audio.

[0321] "Unpleasant sounds" are sounds generated by physiological phenomena such as sneezing, coughing, and hiccuping, which are unpleasant to other participants and disrupt conversation.

[0322] "Cancellation means" refers to a method for filtering out abnormal sounds through a mixed audio signal, allowing other sounds to be heard normally.

[0323] "Means for outputting audio signals" refers to devices or methods for converting the processed audio data back into an audio stream for transmission to other participants.

[0324] "Emotion recognition means" refers to an algorithm or system that grasps the user's emotional state from their voice signal and analyzes it in real time.

[0325] A "means for providing feedback" is a system or method for displaying appropriate advice or suggestions to a user depending on the perceived emotional state.

[0326] This invention is a system that detects unpleasant physiological sounds (sneezing, coughing, hiccups, etc.) that occur during online meetings in real time and automatically cancels them. This system also aims to improve the quality of online meetings by recognizing the user's emotional state in real time and providing appropriate feedback.

[0327] System Configuration

[0328] Audio acquisition means

[0329] The device has a built-in microphone that captures the user's voice signal in real time, which is then stored in a buffer and used in subsequent processing steps.

[0330] Voice analysis methods

[0331] The captured audio signal is pre-processed using digital signal processing techniques, which involves applying a noise reduction algorithm to remove background noise, dividing the signal into frames, and then performing a Fourier transform to extract the audio frequency components.

[0332] Feature extraction method

[0333] Mel-frequency cepstral coefficients (MFCCs) are calculated from the preprocessed speech data, and a feature vector is generated that quantifies the speech features, enabling detailed analysis of the speech signal.

[0334] Classification of normal and abnormal voices

[0335] The device inputs the generated feature vector into a generative AI model to distinguish between normal and abnormal voices. The generative AI model uses past training data to identify patterns in the voice signal and detect abnormal sounds (e.g., sneezing, coughing).

[0336] Abnormal voice detection

[0337] If the generative AI model detects an abnormal sound, it records the time it occurred. For example, if a sneeze occurs between 1.5 and 1.8 seconds, that time frame will be recorded.

[0338] Generation of out-of-phase speech

[0339] An audio signal with an opposite phase is generated for the detected abnormal sound portion. The opposite phase audio signal is an audio signal with a phase difference of 180 degrees to cancel the original abnormal sound signal.

[0340] Audio signal mixing

[0341] By mixing the original audio signal with the generated out-of-phase audio signal, the unpleasant abnormal sounds are cancelled out.

[0342] Audio signal output

[0343] The processed audio data is sent to the audio stream of the online meeting application, so that other participants can hear the audio with the abnormal sounds canceled out.

[0344] emotion recognition means

[0345] The device is equipped with an emotion recognition engine that recognizes the user's emotional state from their voice signal. The engine analyzes the tone, pitch, and speed of the voice to determine the user's current emotional state.

[0346] Feedback methods

[0347] The server then provides appropriate feedback based on the emotional state identified by the emotion recognition engine. For example, if the server determines that the user is feeling stressed, a message such as "Relax" will be displayed on the screen of an online meeting application.

[0348] Specific situations and prompt sentence examples

[0349] Situation

[0350] User B sneezes during an online meeting and feels stressed.

[0351] 1. Audio capture

[0352] The device uses a microphone to capture user B's voice in real time and stores it in a buffer.

[0353] 2. Audio analysis

[0354] Noise reduction is applied to the voice data acquired by the device, and frequency components are extracted.

[0355] 3. Feature Extraction

[0356] The terminal calculates the Mel-frequency cepstral coefficients and generates a feature vector.

[0357] 4. Audio Classification

[0358] The device inputs the feature vector into a generative AI model to identify the time frame of the sneeze sound.

[0359] 5. Abnormal Audio Detection

[0360] The device records the time frame in which the sneeze occurred.

[0361] 6. Generation of out-of-phase speech

[0362] The terminal generates an audio signal that is in the opposite phase to the sneeze sound.

[0363] 7. Mixing of audio signals

[0364] The device mixes the original audio with an out-of-phase audio to cancel out the sneeze sound.

[0365] 8. Audio signal output

[0366] The terminal sends the processed audio signal to the online meeting.

[0367] 9. Emotion recognition

[0368] The device uses an emotion engine to recognize the emotional state of user B.

[0369] 10. Emotional Feedback

[0370] The server displays the message "Relax" based on User B's emotional state.

[0371] Prompt Sentence Examples

[0372] "When a user sneezes during an online meeting, create a program that analyzes the audio signal in real time, generates and mixes an opposite-phase audio signal, and cancels out the unpleasant sound. Also, use an emotion engine to identify the user's emotional state and provide feedback."

[0373] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0374] Step 1: Acquire audio

[0375] The device uses a microphone to capture the user's voice in real time. The input is the user's speech, and the output is the audio signal stored in the buffer. Specifically, the microphone is always on, and the user's voice is continuously sampled and stored in the buffer as digital data.

[0376] Step 2: Preprocessing the audio signal

[0377] The terminal performs digital signal processing (DSP) on the captured audio signal. The input is the raw audio signal stored in the buffer, and the output is the pre-processed audio signal. Specifically, it first applies a noise reduction algorithm to remove background noise. Then it divides the audio signal into frames and applies a Fourier transform to each frame to extract the frequency spectrum.

[0378] Step 3: Feature extraction

[0379] The device calculates Mel-Frequency Cepstral Coefficients (MFCCs) from the preprocessed audio signal. The input is the preprocessed audio signal, and the output is a feature vector. Specifically, the device calculates MFCCs for each frame and combines these MFCC values ​​to generate a single feature vector.

[0380] Step 4: Classify the audio

[0381] The device inputs the generated feature vector into a generative AI model to classify normal and abnormal speech. The input is the feature vector, and the output is the speech classification result. Specifically, the feature vector is input into the AI ​​model, and the AI ​​executes a process to determine whether the speech is normal or abnormal.

[0382] Step 5: Detecting Audio Anomalies

[0383] The device records the time when the abnormal sound occurred based on the classification result. The input is the sound classification result, and the output is the time information of the abnormal sound occurrence. Specifically, it identifies the time frame in which the AI ​​judged the sound to be abnormal and saves that information as a timestamp.

[0384] Step 6: Generate out-of-phase audio

[0385] The terminal generates an audio signal with an inverse phase to the detected abnormal audio portion. The input is the time frame of the abnormal audio and the original audio signal, and the output is an audio signal with an inverse phase. Specifically, it generates an inverse phase signal by inverting the frequency components of the abnormal audio and places that signal in the specified time frame.

[0386] Step 7: Mixing the audio signals

[0387] The terminal mixes the original audio signal with the generated out-of-phase audio signal. The input is the original audio signal and the out-of-phase audio signal, and the output is the mixed audio signal. Specifically, the original audio signal and the out-of-phase audio signal are added together to generate an audio signal that cancels the abnormal audio.

[0388] Step 8: Output the audio signal

[0389] The terminal transmits the processed audio data to the audio stream of the online meeting application. The input is the mixed audio signal, and the output is the processed audio that is transmitted to the participants of the online meeting. In specific operations, the terminal sends the processed audio signal to the transmission channel of the meeting application, allowing other participants to receive the audio with the abnormal sound canceled.

[0390] Step 9: Emotion Recognition

[0391] The device uses an emotion recognition engine to identify the user's emotional state from the voice signal. The input is the real-time voice signal, and the output is the user's emotional state. Specifically, the device analyzes the tone, pitch, and speed of the voice signal to determine emotions such as stress, joy, and sadness.

[0392] Step 10: Emotional Feedback

[0393] The server provides appropriate feedback according to the emotional state identified by the emotion recognition engine. The input is the user's emotional state, and the output is a feedback message. Specifically, the server generates a message that matches the user's emotional state and displays appropriate feedback, such as "Relax," on the screen of the online meeting application.

[0394] (Application example 2)

[0395] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0396] Robots in factories are required to accurately receive voice instructions from operators without misrecognition. However, if abnormal sounds such as sneezing or coughing are included in the voice, the risk of the robot malfunctioning increases. Furthermore, if the operator is stressed, it will affect work efficiency and safety. To solve these problems, a system is needed that combines the detection and removal of abnormal sounds and the recognition of the operator's emotional state.

[0397] The identification processing by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for acquiring a voice signal, artificial intelligence means for analyzing the acquired voice signal and distinguishing between normal voice and abnormal voice, means for generating an opposite-phase voice signal when an abnormal voice is detected, means for mixing the original voice signal with the opposite-phase voice signal to cancel out unpleasant voice, means for outputting a voice signal, emotion engine means for extracting features from the voice signal and identifying the emotional state, and means for providing feedback according to the emotional state. This allows the robot to accurately receive the operator's voice instructions, is not affected by abnormal sounds (such as sneezing or coughing), and can provide appropriate feedback according to the operator's emotional state.

[0398] The "means for acquiring an audio signal" refers to a means for obtaining audio data by a microphone or other audio input device.

[0399] "Generative artificial intelligence means for analyzing acquired voice signals and distinguishing between normal voices and abnormal voices" refers to a means for analyzing acquired voice data and using a generative AI to distinguish between normal voices and abnormal voices such as sneezing and coughing.

[0400] The "means for generating an audio signal of an opposite phase when an abnormal sound is detected" refers to a means for generating an audio signal of a phase shift of 180 degrees when an abnormal sound is detected in order to cancel out the sound.

[0401] The "means for mixing the original audio signal with an out-of-phase audio signal to cancel unpleasant sounds" is a means for canceling abnormal sounds, specifically unpleasant sounds, by combining the original audio with the generated out-of-phase audio.

[0402] "Means for outputting audio signals" refers to means for transmitting processed audio data to other systems or devices.

[0403] "Emotion engine means for extracting features from voice signals and identifying emotional states" refers to a means for extracting emotion-related features such as tone and pitch from voice data and analyzing and identifying the emotional state of the operator.

[0404] The "means for providing feedback according to emotional state" is a means for providing appropriate feedback or a message based on the emotional state of the operator identified by the emotion engine.

[0405] This invention is a system that enables robots in factories to accurately receive voice instructions from operators without misrecognition. In particular, it eliminates the effects of abnormal sounds such as sneezing and coughing, and provides feedback according to the operator's emotional state.

[0406] The system includes many means, but the main parts are as follows:

[0407] 1. Acquisition of audio signals:

[0408] The server uses a microphone or other audio input device to capture the operator's voice signals in real time, which consist of instructions and other environmental sounds in the factory.

[0409] 2. Means of analyzing the audio signal:

[0410] The server analyzes the captured audio signal using digital signal processing techniques. Specifically, it applies a noise reduction algorithm, divides the audio signal into frames, performs a Fourier transform, and extracts the audio frequency components. This process is performed using the feature.mfcc function from librosa.

[0411] 3. Generative AI means:

[0412] The server inputs the features extracted from the voice signal into a generative AI model, which distinguishes between normal voices and abnormal voices such as sneezing and coughing. This generative AI model uses a pre-trained convolutional neural network (CNN).

[0413] 4. Means for generating anti-phase audio signals:

[0414] When an abnormal sound is detected, the server generates an audio signal that is 180 degrees out of phase with the detected sound to cancel it out, effectively canceling the abnormal sound.

[0415] 5. Audio signal mixing means:

[0416] The server cancels out abnormal sounds by mixing the original audio signal with the generated out-of-phase audio, so that only normal instruction audio reaches the robot.

[0417] 6. Audio output means:

[0418] The server transmits the processed voice data to the robot so that the robot receives the voice instructions accurately.

[0419] 7. Emotion Engine Means:

[0420] The server analyzes the operator's emotional state using features extracted from the voice signal, using algorithms that use voice features such as tone and pitch.

[0421] 8. Feedback methods:

[0422] The server provides appropriate feedback based on the operator's emotional state as identified by the emotion engine. For example, if stress is detected, the robot will display a message encouraging the operator to take a break.

[0423] Specific examples

[0424] In a real-world scenario, an operator sneezes multiple times while working and then becomes extremely stressed. Detecting this state, the robot pauses the work and displays a message encouraging the operator to take a break. The system's processing steps are seamlessly handled by the server, and appropriate actions are automatically taken.

[0425] Prompt Sentence Examples

[0426] Build a system where if an operator sneezes multiple times and then becomes very stressed, the robot will display a message encouraging them to take a break.

[0427] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0428] Step 1:

[0429] The server uses a microphone to capture the operator's voice signal in real time. The voice signal obtained from this microphone includes environmental sounds in the factory and the operator's voice instructions. The input is an analog voice signal, and the output is digitized voice data. The server uses a buffer memory to store this data.

[0430] Step 2:

[0431] The server applies digital signal processing technology to the acquired audio signal, performs noise reduction, divides the audio signal into frames, performs a Fourier transform, and extracts the frequency components of the audio. The input is the digitized audio data obtained in step 1, and the output is frame data converted into frequency components.

[0432] Step 3:

[0433] The server calculates Mel-Frequency Cepstrum Coefficients (MFCC) from the frame data and extracts speech features. The input is the frequency component data obtained in step 2, and the output is a feature vector obtained using MFCC.

[0434] Step 4:

[0435] The server inputs the extracted feature vector into a generative AI model to classify normal and abnormal speech. The generative AI model identifies speech signal patterns based on past training data. The input is the feature vector obtained in step 3, and the output is a label: "normal speech" or "abnormal speech."

[0436] Step 5:

[0437] When an abnormal sound (sneezing or coughing) is detected, the server records its start and end times. The input is the classification label obtained in step 4, and the output is the time frame information of the abnormal sound.

[0438] Step 6:

[0439] The server generates an audio signal with an inverse phase to the detected abnormal audio. This inverse phase audio signal is an audio signal with a phase difference of 180 degrees to cancel out the original abnormal audio. The input is the time frame information of the abnormal audio obtained in step 5, and the output is the inverse phase audio signal.

[0440] Step 7:

[0441] The server cancels the unpleasant abnormal sounds by mixing the original sound signal with an out-of-phase sound signal. The input is the original sound signal obtained in step 1 and the out-of-phase sound signal generated in step 6, and the output is the cancelled sound signal.

[0442] Step 8:

[0443] The server sends the processed audio signal to the robot, which allows the robot to receive an audio signal that does not contain the canceled abnormal audio. The input is the canceled audio signal obtained in step 7, and the output is the digital audio signal received by the robot.

[0444] Step 9:

[0445] The server re-extracts features from the cancelled speech signal and inputs them into an emotion engine for identifying emotional states. The input is the cancelled speech signal obtained in step 7, and the output is a label indicating the emotional state.

[0446] Step 10:

[0447] The server provides appropriate feedback based on the emotional state identified by the emotion engine. For example, if many abnormal sounds are detected and the emotional state is identified as "stressed," the robot displays a message encouraging the operator to take a break. The input is the label indicating the emotional state obtained in step 9, and the output is the feedback message.

[0448] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0449] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0450] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0451] [Second embodiment]

[0452] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0453] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0454] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0455] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0456] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0457] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0458] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0459] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0460] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0461] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0462] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0463] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0464] This invention relates to a system that detects physiological phenomena such as sneezing, coughing, and hiccups that occur during online meetings and automatically mutes them. This system acquires audio signals in real time, uses generation AI to distinguish between normal and abnormal sounds, and if an abnormal sound is detected, generates and mixes an opposite-phase audio signal to cancel out the unpleasant sound.

[0465] System configuration

[0466] 1. Audio acquisition method

[0467] Terminal: This system is equipped with a microphone to capture the user's voice in real time. The microphone continuously captures the voice signal.

[0468] 2. Voice Analysis Methods

[0469] Terminal: The captured audio signal is analyzed using digital signal processing techniques, which apply noise reduction algorithms, divide it into frames, and perform a Fourier transform to extract the audio frequency components.

[0470] 3. Feature Extraction Method

[0471] Terminal: Calculate features such as Mel-Frequency Cepstrum Coefficients (MFCC) from the preprocessed speech data and generate a feature vector, which represents the characteristics of the speech as numerical data.

[0472] 4. Abnormal sound detection method

[0473] Terminal: The feature vector is input into the generation AI to distinguish between normal and abnormal voices. In particular, when abnormal voices such as sneezing, coughing, and hiccups are detected, the start and end times are recorded.

[0474] 5. Antiphase speech generation means

[0475] Terminal: When an abnormal audio is detected, it generates an audio signal with the opposite phase to the abnormal audio portion. The opposite phase audio signal has the phase information necessary to cancel out the original audio signal.

[0476] 6. Audio Mixing Methods

[0477] Terminal: Mix the original audio signal with an inverted audio signal to cancel out any unwanted sounds. This operation ensures that the final audio is output smoothly.

[0478] 7. Audio output means

[0479] Terminal: The processed audio signal is sent to the online meeting application, which prevents other participants from hearing the disturbing audio.

[0480] Specific examples

[0481] Situation:

[0482] User A suddenly sneezes during an online meeting.

[0483] 1. Audio capture

[0484] Device: The microphone captures user A's voice in real time and stores it in a buffer.

[0485] 2. Voice Analysis

[0486] Terminal: Noise reduction is applied to the acquired audio data, it is divided into frames, and a Fourier transform is performed.

[0487] 3. Feature Extraction

[0488] Terminal: Calculate Mel-frequency cepstrum coefficients from the audio data and generate a feature vector.

[0489] 4. Abnormal Audio Detection

[0490] Terminal: The feature vector is input into the generative AI, which detects that the sneeze occurred between 1.5 and 1.8 seconds.

[0491] 5. Antiphase speech generation

[0492] Terminal: Generates an audio signal in anti-phase with a sneeze sound between 1.5 and 1.8 seconds.

[0493] 6. Audio Mixing

[0494] Terminal: Mixes the original audio signal with an out-of-phase audio signal to cancel the sneeze sound.

[0495] 7. Audio Output

[0496] Device: Sends the processed audio signal to the online meeting application, where other participants hear the audio without hearing any sneeze sounds.

[0497] Using this system, users can participate in online meetings with peace of mind and prevent unintended audio from being heard by others.

[0498] The processing flow will be explained below.

[0499] Step 1: Getting voice input

[0500] On the device: The microphone is activated and the user's voice is captured in real time. The captured voice data is stored in a buffer.

[0501] Step 2: Preprocessing the audio data

[0502] Terminal: A noise reduction algorithm is applied to the acquired audio data. The audio data is then divided into short time frames, and frequency components are extracted by performing a Fourier transform on each frame.

[0503] Step 3: Extract audio features

[0504] Terminal: To extract features from the preprocessed speech data, Mel-Frequency Cepstral Coefficients (MFCCs) are calculated, which generates a feature vector that quantifies the speech features.

[0505] Step 4: Classify normal and abnormal voices

[0506] Terminal: The generated feature vector is input into a generative artificial intelligence (generative AI) to classify normal and abnormal speech. This AI identifies patterns in the speech signal based on past training data and detects abnormal speech, such as sneezing or coughing.

[0507] Step 5: Detecting Audio Anomalies

[0508] Terminal: When the artificial intelligence detects an abnormal sound, it records its start and end time. For example, if a certain part of the audio signal is recognized as a sneeze, that time frame will be marked.

[0509] Step 6: Generate an anti-phase audio signal

[0510] Terminal: Generates an audio signal with an inverse phase to the detected abnormal audio portion. The inverse phase audio signal is an audio signal with a phase difference of 180 degrees to cancel the original abnormal audio signal.

[0511] Step 7: Mixing the audio signals

[0512] Terminal: Mix the original audio signal with the generated out-of-phase audio signal. This mixing cancels out the unpleasant abnormal sounds.

[0513] Step 8: Output the audio signal

[0514] Terminal: Sends the processed audio data to the audio stream of the online meeting, so that other participants hear the audio with the abnormal audio canceled.

[0515] Step 9: Real-time monitoring and adjustment

[0516] Terminal: All processes from audio signal acquisition to output are repeated in real time, continuously monitoring for abnormal audio and canceling it as appropriate.

[0517] In this way, the system of the present invention can automatically cancel unpleasant sounds such as sneezing and coughing that occur during an online meeting in real time, allowing users to maintain smooth conversations.

[0518] Example 1

[0519] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0520] During online meetings, physiological phenomena such as sneezing, coughing, and hiccups can interrupt conversations and cause discomfort to other participants. Because these sounds are not automatically muted, users may unintentionally let other participants hear the unpleasant sounds.

[0521] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0522] In this invention, the server includes means for acquiring voice data, means for preprocessing the acquired voice data using digital signal processing technology, means for extracting voice features from the preprocessed voice data, means for inputting the extracted voice features into a generative artificial intelligence model to distinguish between normal voice and abnormal voice, means for generating out-of-phase voice data when abnormal voice is detected, means for mixing the original voice data with the out-of-phase voice data to cancel unpleasant voice, and means for outputting the processed voice data. This prevents unpleasant sounds caused by physiological phenomena such as sneezing, coughing, and hiccups from reaching other participants during online meetings, enabling comfortable communication.

[0523] A "means for acquiring voice data" is a device or method for digitally capturing a user's speech or voice in real time using hardware such as a microphone.

[0524] "Digital signal processing" is the process of converting analog audio signals into digital form and using techniques such as noise reduction and Fourier transform to improve the quality and characteristics of the audio.

[0525] "Speech features" are numerical representations of the characteristics of a speech signal, and are typically extracted using techniques such as Mel-Frequency Cepstrum Coefficients (MFCC).

[0526] A "generative artificial intelligence model" is an algorithm or network that is trained to use machine learning or deep learning techniques to analyze speech data and distinguish between normal and abnormal speech.

[0527] "Out-of-phase audio data" is audio data that has the opposite phase to the original audio signal and is generated to cancel or reduce audio by mixing it with the original audio signal.

[0528] The "means for outputting audio data" refers to a device or method for transferring the processed audio data to a user's terminal or an online meeting application.

[0529] "Abnormal voices" are sounds caused by noises or physiological phenomena that are not part of normal conversation, such as sneezing, coughing, or hiccuping, and are detected by the generative artificial intelligence model.

[0530] This invention relates to a system that detects physiological phenomena such as sneezing, coughing, and hiccups that occur during online meetings and automatically mutes them. This system acquires audio signals in real time, uses a generative AI model to distinguish between normal and abnormal sounds, and when abnormal sounds are detected, generates and mixes an opposite-phase audio signal to cancel out the unpleasant sounds.

[0531] The system includes the following components:

[0532] 1. How to obtain audio data

[0533] Terminal: This system is equipped with a microphone to capture the user's voice in real time. Voice data is continuously acquired from the microphone.

[0534] 2. Preprocessing using digital signal processing

[0535] Terminal: The captured audio data is pre-processed using digital signal processing techniques, which apply noise reduction algorithms, divide the audio data into frames, perform Fourier transforms, and extract frequency components.

[0536] 3. Methods for extracting speech features

[0537] Terminal: Calculate Mel-Frequency Cepstral Coefficients (MFCC) from the preprocessed speech data to generate a feature vector, which represents the characteristics of the speech as numerical data.

[0538] 4. Method for detecting abnormal voice using generative AI model

[0539] On the device: The feature vector is input into a generative AI model to distinguish between normal and abnormal speech. In particular, when abnormal speech such as sneezing, coughing, or hiccuping is detected, the start and end times are recorded.

[0540] 5. Means for generating out-of-phase audio data

[0541] Terminal: When an abnormal audio is detected, it generates audio data with an inverse phase to the abnormal audio portion. The inverse phase audio data has the phase information necessary to cancel out the original audio data.

[0542] 6. Means of mixing audio data

[0543] Terminal: Mix the original audio data with the opposite phase audio data to cancel out the unpleasant sounds. This operation makes the final audio output smooth.

[0544] 7. Means of outputting audio data

[0545] Terminal: The processed audio data is sent to the online meeting application, which prevents other participants from hearing the disturbing audio.

[0546] Specific examples are shown below.

[0547] If user A suddenly sneezes during an online meeting, the following occurs:

[0548] 1. How to obtain audio data

[0549] Device: The microphone captures user A's voice in real time and stores it in a buffer.

[0550] 2. Preprocessing using digital signal processing

[0551] Terminal: Apply noise reduction to the acquired audio data, divide it into frames, and perform a Fourier transform.

[0552] 3. Methods for extracting speech features

[0553] Terminal: Calculate Mel-Frequency Cepstrum Coefficients (MFCC) from the audio data and generate a feature vector.

[0554] 4. Method for detecting abnormal voice using generative AI model

[0555] On the device: The feature vector is fed into a generative AI model to detect that the sneeze occurred between 1.5 and 1.8 seconds.

[0556] 5. Means for generating out-of-phase audio data

[0557] Device: Generates audio data in antiphase to a sneeze sound between 1.5 and 1.8 seconds.

[0558] 6. Means of mixing audio data

[0559] Terminal: Mix the original audio data with out-of-phase audio data to cancel the sneeze sound.

[0560] 7. Means of outputting audio data

[0561] Device: Sends the processed audio data to the online meeting application, where other participants hear the audio without hearing the sneeze.

[0562] This system allows users to participate in online meetings with peace of mind and prevents unintended audio from being heard by others.

[0563] Prompt Sentence Examples

[0564] "Please explain the system that automatically mutes you when you sneeze during an online meeting."

[0565] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0566] Step 1:

[0567] Acquiring audio data

[0568] Device: Uses a microphone to capture the user's voice data in real time.

[0569] Input: User speech and surrounding audio.

[0570] Output: Digital audio data from the microphone.

[0571] What it does: The microphone is always on, capturing audio data in real time and storing it in a fixed buffer.

[0572] Step 2:

[0573] Audio signal preprocessing

[0574] Terminal: Performs digital signal processing on the captured audio data, applying noise reduction algorithms, dividing the audio data into frames, and performing a Fourier transform to extract frequency components.

[0575] Input: Raw audio data.

[0576] Output: Denoised per-frame frequency spectrum.

[0577] What it does: It filters the audio signal to remove noise, then divides it into small time segments (frames), then performs a Fourier transform on each frame to calculate its frequency spectrum.

[0578] Step 3:

[0579] Feature extraction

[0580] Terminal: Calculate the Mel-Frequency Cepstral Coefficients (MFCC) from each frame and generate a feature vector.

[0581] Input: Frequency spectrum obtained by Fourier transform.

[0582] Output: MFCC-based feature vectors representing audio features.

[0583] Specific operation: Calculate MFCC using frequency spectrum as input, and save the calculation result as a feature vector in list format.

[0584] Step 4:

[0585] Abnormal audio detection

[0586] On the device: The feature vector is fed into a generative AI model to distinguish between normal and abnormal speech. If abnormal speech is detected, the start and end times of the speech are recorded.

[0587] Input: MFCC-based feature vector.

[0588] Output: Whether or not there is an abnormal sound, and its start and end times.

[0589] Specific operation: The generative AI model detects abnormal sounds based on the input feature vector and saves the period during which the abnormal sounds occurred as a timestamp.

[0590] Step 5:

[0591] Antiphase audio data generation

[0592] Terminal: Generates audio data in the opposite phase to the time range of the detected abnormal audio.

[0593] Input: Start and end times of the abnormal audio, and the original audio data.

[0594] Output: Out-of-phase audio data.

[0595] Specific operation: Based on the time range of the detected abnormal sound, calculate and save the opposite phase audio waveform.

[0596] Step 6:

[0597] Mixing of audio data

[0598] Terminal: Mixes the original audio data with out-of-phase audio data to cancel abnormal sounds.

[0599] Input: Original audio data, inverted audio data.

[0600] Output: Audio data with abnormal sounds canceled.

[0601] Specific operation: The original audio data is added to the out-of-phase audio data, and a process is performed to cancel out abnormal sounds.

[0602] Step 7:

[0603] Sending audio data

[0604] Terminal: Sends the processed audio data to the online meeting application.

[0605] Input: Mixed audio data.

[0606] Output: Final audio data (audio heard by other participants in the online meeting).

[0607] What it does: Stream the processed audio data in real time to an online meeting application.

[0608] (Application example 1)

[0609] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0610] In conventional online meetings and autonomous vehicle environments, sudden physiological events such as sneezing or coughing can cause unpleasant sounds to those around you and temporarily worsen the air quality, resulting in a decrease in voice quality, hindering conversations and the transmission of instructions, as well as reducing the comfort of the vehicle's interior.

[0611] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0612] In this invention, the server includes means for acquiring an audio signal, artificial intelligence means for analyzing the acquired audio signal and distinguishing between normal and abnormal audio, means for generating an opposite-phase audio signal when an abnormal audio signal is detected, means for mixing the original audio signal with the opposite-phase audio signal to cancel the unpleasant audio, means for outputting the audio signal, and means for temporarily strengthening related systems when canceling a new physiological phenomenon. This not only enables immediate detection and control of abnormal audio, but also temporarily improves the air environment, making it possible to maintain a comfortable and healthy environment.

[0613] plaintext

[0614] The "means for acquiring an audio signal" refers to a microphone and associated hardware for converting sound into an electronic signal.

[0615] The "generative artificial intelligence means" is an algorithm using artificial intelligence to analyze the acquired voice signal and distinguish between normal voice and abnormal voice.

[0616] The "means for generating an inverted phase audio signal" refers to a device or algorithm for generating an audio signal with the phase of an abnormal audio signal inverted when the abnormal audio signal is detected.

[0617] "Means for canceling unpleasant sounds" is a technology that synthesizes an audio signal that is out of phase with the original audio signal to cancel out unpleasant sounds.

[0618] The "means for outputting an audio signal" refers to a device such as a speaker or earphone for reproducing the processed audio signal.

[0619] The "means for temporarily enhancing related systems" is a control device for temporarily enhancing the air filter system and other related systems in the vehicle when a physiological phenomenon is detected.

[0620] plaintext

[0621] This invention is a system primarily designed to improve the riding experience in self-driving vehicles. It detects sudden physiological phenomena such as sneezing and coughing, cancels the unpleasant sounds that occur during these events, and temporarily improves the in-car environment.

[0622] The elements of the server, terminal, and user are as follows:

[0623] server

[0624] The vehicle is equipped with a microphone that captures audio signals and a computer system that includes an artificial intelligence algorithm for analyzing the audio. Audio signals are captured in real time and analyzed using digital signal processing technology. If an abnormal sound is detected, an opposite-phase audio signal is generated, thereby canceling out the unpleasant sound. This detection also temporarily strengthens the cabin air filter system.

[0625] Terminal

[0626] The terminal is a component of the in-vehicle system that acquires and analyzes audio signals in real time. Specifically, it performs noise reduction on the audio signal, frequency analysis using Fourier transform, and calculation of Mel Frequency Cepstrum Coefficients (MFCC). It also uses generative artificial intelligence to distinguish between normal and abnormal audio, and generates and mixes an inverse audio signal for detected abnormal audio to cancel it. It also includes a device that outputs the processed audio signal.

[0627] User

[0628] As passengers in self-driving vehicles, users can benefit from the ability to avoid physiological events such as sneezing or coughing while in an online meeting or using the navigation system, causing discomfort to other passengers and the system.

[0629] Specific examples

[0630] For example, the following prompt sentences can be used to input to a generative AI model to recreate a situation in which the present invention is implemented:

[0631] "Detect the sound of a sneeze in the car, mute the audio system at that moment, and strengthen the air filtration system."

[0632] This allows the in-car environment to be controlled quickly and automatically, providing a comfortable riding experience.

[0633] Hardware and Software

[0634] The system utilizes the following hardware and software:

[0635] Hardware: microphones, on-board computers, speakers, air filtration systems.

[0636] Software: Noise reduction algorithm, Fourier transform algorithm, Mel-Frequency Cepstral Coefficient (MFCC) extraction algorithm, generative AI model (TensorFlow, PyTorch, etc.).

[0637] By using the above prompt sentences, the generative AI model can immediately detect abnormal sounds and improve the air quality inside the car while canceling out unpleasant sounds.

[0638] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0639] plaintext

[0640] Step 1:

[0641] The server captures audio signals in real time from microphones inside the vehicle. The captured analog audio signals are converted to digital signals and passed to the next processing step. This input includes various sounds, such as ambient sounds inside the vehicle and conversation sounds.

[0642] Step 2:

[0643] The terminal performs noise reduction filtering on the captured voice signal, specifically using digital signal processing techniques to remove unwanted background noise, resulting in a clean voice signal being output.

[0644] Step 3:

[0645] The terminal divides the noise-reduced speech signal into frames and performs a Fourier transform to extract frequency components. The input is the speech signal after noise reduction, and the output is the frequency spectrum.

[0646] Step 4:

[0647] The device extracts speech features by calculating Mel-Frequency Cepstrum Coefficients (MFCC) from the frequency spectrum. This method expresses speech characteristics as numerical data. The input is the frequency spectrum, and the output is an MFCC feature vector.

[0648] Step 5:

[0649] The device inputs the MFCC feature vectors into a generative AI model (e.g., a model using TensorFlow or PyTorch) to distinguish between normal and abnormal speech. If abnormal speech (such as sneezing or coughing) is detected in this step, the device outputs the start and end times of the speech.

[0650] Step 6:

[0651] The terminal generates an audio signal with an opposite phase to the detected abnormal audio. To generate the opposite phase audio signal, an operation is performed to invert the phase of the original audio signal. The input is the audio signal of the abnormal audio, and the output is the audio signal with the opposite phase.

[0652] Step 7:

[0653] The terminal mixes the original audio signal with the generated out-of-phase audio signal, thereby performing a calculation to cancel the abnormal audio. The input is the original audio signal and the out-of-phase audio signal, and the output is the cancelled audio signal.

[0654] Step 8:

[0655] The device outputs the processed audio signal through the vehicle's speakers, allowing other passengers to hear the audio with the unpleasant sounds canceled out.

[0656] Step 9:

[0657] When an abnormal sound is detected, the server triggers a temporary strengthening of the air filter system inside the vehicle. Specifically, it increases the airflow of the air filter to quickly remove allergens and viruses in the air. This operation improves the air environment inside the vehicle.

[0658] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0659] This invention relates to a system that detects physiological phenomena such as sneezing, coughing, and hiccups that occur during online meetings and automatically mutes them. This system acquires audio signals in real time and uses generative AI to distinguish between normal and abnormal sounds. If abnormal sounds are detected, it generates and mixes an opposite-phase audio signal to cancel out the unpleasant sounds. Furthermore, by combining it with an emotion engine that recognizes the user's emotions, it has the ability to identify the user's emotional state and provide appropriate feedback in addition to processing the audio signals.

[0660] System configuration

[0661] 1. Audio acquisition method

[0662] Device: Equipped with a microphone for capturing audio signals in real time. Audio signals are continuously captured from this microphone.

[0663] 2. Voice Analysis Methods

[0664] Terminal: The captured audio data is analyzed using digital signal processing techniques, which apply a noise reduction algorithm, divide it into frames, and perform a Fourier transform to extract the frequency components of the audio.

[0665] 3. Feature Extraction Method

[0666] Terminal: To extract features from the preprocessed speech data, Mel-Frequency Cepstral Coefficients (MFCCs) are calculated, which generates a feature vector that quantifies the speech features.

[0667] 4. Classification of normal and abnormal voices

[0668] Terminal: The generated feature vector is input into the generative artificial intelligence (generative AI) to classify normal and abnormal voices. The generative AI identifies patterns in the voice signal based on past training data and detects abnormal voices such as sneezing and coughing.

[0669] 5. Abnormal Audio Detection

[0670] Terminal: When an abnormal sound is detected, its start and end time are recorded. For example, if a certain part of the sound signal is recognized as a sneeze, that time frame is marked.

[0671] 6. Generation of anti-phase audio signals

[0672] Terminal: Generates an audio signal with an inverse phase to the detected abnormal audio portion. The inverse phase audio signal is an audio signal with a phase difference of 180 degrees to cancel the original abnormal audio signal.

[0673] 7. Mixing of audio signals

[0674] Terminal: Mix the original audio signal with the generated out-of-phase audio signal. This mixing cancels out the unpleasant abnormal sounds.

[0675] 8. Audio signal output

[0676] Terminal: Sends the processed audio data to the audio stream of the online meeting, so that other participants hear the audio with the abnormal audio canceled.

[0677] 9. Emotion Engine

[0678] Device: The emotion engine has algorithms to identify the emotional state from the user's voice signal. The emotion engine analyzes voice characteristics such as tone, pitch, and rate to identify the emotion the user is feeling.

[0679] Server: Provides different feedback depending on the emotional state identified by the emotion engine. For example, if the server determines that the user is feeling stressed, it displays a message on the screen suggesting that the user relax.

[0680] Specific examples

[0681] Situation:

[0682] User B suddenly sneezes during an online meeting and feels stressed.

[0683] 1. Audio capture

[0684] Device: The microphone captures user B's voice in real time and stores it in a buffer.

[0685] 2. Voice Analysis

[0686] Terminal: Noise reduction is applied to the acquired audio data, it is divided into frames, and a Fourier transform is performed.

[0687] 3. Feature Extraction

[0688] Terminal: Calculate Mel-frequency cepstrum coefficients from the audio data and generate a feature vector.

[0689] 4. Classification of normal and abnormal voices

[0690] On the device: The feature vector is fed into a generative AI to detect that a sneeze occurred within a specific time frame.

[0691] 5. Abnormal Audio Detection

[0692] Device: Detects that the sneeze occurred between 1.5 and 1.8 seconds and marks that time frame.

[0693] 6. Generation of anti-phase audio signals

[0694] Device: Generates an audio signal in antiphase to a sneeze sound lasting 1.5 to 1.8 seconds.

[0695] 7. Mixing of audio signals

[0696] Terminal: Mixes the original audio signal with an out-of-phase audio signal to cancel the sneeze sound.

[0697] 8. Audio signal output

[0698] Terminal: Sends the processed audio signal to the online meeting application. Other participants receive the audio without hearing the sneeze.

[0699] 9. Emotion recognition

[0700] Device: The emotion engine identifies User B's emotional state from audio signals including the sneeze sound.

[0701] 10. Emotional Feedback

[0702] Server: If the emotion engine identifies User B's emotional state as "stressed," a message such as "Relax" will be displayed on the screen of the online meeting application.

[0703] Using this system, users can not only automatically cancel unpleasant sounds such as sneezing and coughing during online meetings in real time, but also receive support based on their emotional state using an emotion engine, allowing users to participate in online meetings with greater peace of mind.

[0704] The processing flow will be explained below.

[0705] Step 1: Getting voice input

[0706] On the device: The microphone is activated and the user's voice is captured in real time. The captured voice data is stored in a buffer.

[0707] Step 2: Preprocessing the audio data

[0708] Terminal: A noise reduction algorithm is applied to the acquired audio data to obtain a clear audio signal. The audio data is then divided into short time frames, and a Fourier transform is performed on each frame to extract the frequency components of the audio.

[0709] Step 3: Extract audio features

[0710] Terminal: To extract features from the preprocessed speech data, Mel-Frequency Cepstral Coefficients (MFCCs) are calculated, which generates a feature vector that quantifies the speech features.

[0711] Step 4: Classify normal and abnormal voices

[0712] Terminal: The generated feature vector is input into the generative artificial intelligence (generative AI) to classify normal and abnormal voices. The generative AI identifies patterns in the voice signal based on past training data and detects abnormal voices such as sneezing and coughing.

[0713] Step 5: Detecting Audio Anomalies

[0714] Terminal: When abnormal sounds are detected by the AI ​​generator, their start and end times are recorded. For example, if the audio signal is recognized as a sneeze from 1.5 seconds to 1.8 seconds, that time frame is marked.

[0715] Step 6: Generate an anti-phase audio signal

[0716] Terminal: Generates an audio signal with an inverse phase to the detected abnormal audio portion. The inverse phase audio signal is an audio signal with a phase difference of 180 degrees to cancel the original abnormal audio signal.

[0717] Step 7: Mixing the audio signals

[0718] Terminal: Mix the original audio signal with the generated out-of-phase audio signal. This mixing cancels out the unpleasant abnormal sounds.

[0719] Step 8: Output the audio signal

[0720] Terminal: Sends the processed audio data to the audio stream of the online meeting, so that other participants hear the audio with the abnormal audio canceled.

[0721] Step 9: Emotion Recognition Processing

[0722] On the device: The voice signal is processed by an emotion engine that analyzes characteristics such as tone, pitch, and rate to identify the user's emotional state.

[0723] On the device: The emotion engine identifies the user's emotion based on the analyzed features. For example, if the voice is spoken with a strong tone, high pitch, and rapid speed, it will be identified as stressed or angry.

[0724] Step 10: Emotional State Feedback

[0725] Server: Provides appropriate feedback based on the emotional state identified by the emotion engine. For example, if the server determines that the user is feeling stressed, it displays a message such as "Relax" on the screen of an online meeting application.

[0726] Server: Based on the user's emotional state, it provides a feedback function that notifies other participants of the user's state, depending on the settings.

[0727] In this way, the entire system performs a series of processes in real time, from acquiring voice signals to canceling unpleasant sounds, and then identifying the user's emotions and providing appropriate feedback. This allows users to continue participating in online meetings with peace of mind, even when experiencing physiological or emotional changes.

[0728] Example 2

[0729] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0730] Physiological phenomena such as sneezing, coughing, and hiccups that occur during online meetings can be unpleasant for other participants and disrupt the flow of conversation. Furthermore, users often feel stressed if they become too aware of these physiological phenomena. Furthermore, it is difficult to grasp users' emotional state in real time, and there is a lack of means to provide appropriate feedback. This can reduce the quality of online meetings and impair participants' concentration and comfort.

[0731] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0732] In this invention, the terminal includes means for acquiring an audio signal, artificial intelligence means for analyzing the acquired audio signal and distinguishing between normal audio and abnormal audio, means for generating an opposite-phase audio signal when abnormal audio is detected, means for mixing the original audio signal with the opposite-phase audio signal and canceling unpleasant audio, means for outputting an audio signal, and emotion recognition means for recognizing the user's emotional state and providing feedback. This makes it possible to automatically cancel unpleasant physiological sounds during online meetings and provide support according to the user's emotional state.

[0733] The "audio signal" is a user's speech converted into an electrical signal, and is acquired through an input device such as a microphone.

[0734] "Means for acquiring" refers to a device or method for acquiring audio signals in real time using a microphone, sensor device, etc.

[0735] "Analyzing" is the process of analyzing an audio signal using digital signal processing techniques to extract various features.

[0736] "Normal voice" refers to voice that is not an abnormal voice, such as a normal conversation or speech.

[0737] "Abnormal sounds" refers to physiological sounds that interfere with normal conversation, such as sneezing, coughing, and hiccups.

[0738] "Artificial intelligence tools" are algorithms or models that use historical training data to identify speech patterns and categorize or differentiate between speech sounds.

[0739] "Detecting abnormal sounds" means identifying abnormal sounds from the acquired audio signals and recognizing the time of occurrence and their content.

[0740] An "opposite phase audio signal" is an audio signal generated by reversing the phase of the original abnormal audio signal by 180 degrees, and is used to cancel the abnormal audio.

[0741] "Mixing" is a process of synthesizing an audio signal that is out of phase with the original audio signal to cancel out abnormal audio.

[0742] "Unpleasant sounds" are sounds generated by physiological phenomena such as sneezing, coughing, and hiccuping, which are unpleasant to other participants and disrupt conversation.

[0743] "Cancellation means" refers to a method for filtering out abnormal sounds through a mixed audio signal, allowing other sounds to be heard normally.

[0744] "Means for outputting audio signals" refers to devices or methods for converting the processed audio data back into an audio stream for transmission to other participants.

[0745] "Emotion recognition means" refers to an algorithm or system that grasps the user's emotional state from their voice signal and analyzes it in real time.

[0746] A "means for providing feedback" is a system or method for displaying appropriate advice or suggestions to a user depending on the perceived emotional state.

[0747] This invention is a system that detects unpleasant physiological sounds (sneezing, coughing, hiccups, etc.) that occur during online meetings in real time and automatically cancels them. This system also aims to improve the quality of online meetings by recognizing the user's emotional state in real time and providing appropriate feedback.

[0748] System Configuration

[0749] Audio acquisition means

[0750] The device has a built-in microphone that captures the user's voice signal in real time, which is then stored in a buffer and used in subsequent processing steps.

[0751] Voice analysis methods

[0752] The captured audio signal is pre-processed using digital signal processing techniques, which involves applying a noise reduction algorithm to remove background noise, dividing the signal into frames, and then performing a Fourier transform to extract the audio frequency components.

[0753] Feature extraction method

[0754] Mel-frequency cepstral coefficients (MFCCs) are calculated from the preprocessed speech data, and a feature vector is generated that quantifies the speech features, enabling detailed analysis of the speech signal.

[0755] Classification of normal and abnormal voices

[0756] The device inputs the generated feature vector into a generative AI model to distinguish between normal and abnormal voices. The generative AI model uses past training data to identify patterns in the voice signal and detect abnormal sounds (e.g., sneezing, coughing).

[0757] Abnormal voice detection

[0758] If the generative AI model detects an abnormal sound, it records the time it occurred. For example, if a sneeze occurs between 1.5 and 1.8 seconds, that time frame will be recorded.

[0759] Generation of out-of-phase speech

[0760] An audio signal with an opposite phase is generated for the detected abnormal sound portion. The opposite phase audio signal is an audio signal with a phase difference of 180 degrees to cancel the original abnormal sound signal.

[0761] Audio signal mixing

[0762] By mixing the original audio signal with the generated out-of-phase audio signal, the unpleasant abnormal sounds are cancelled out.

[0763] Audio signal output

[0764] The processed audio data is sent to the audio stream of the online meeting application, so that other participants can hear the audio with the abnormal sounds canceled out.

[0765] emotion recognition means

[0766] The device is equipped with an emotion recognition engine that recognizes the user's emotional state from their voice signal. The engine analyzes the tone, pitch, and speed of the voice to determine the user's current emotional state.

[0767] Feedback methods

[0768] The server then provides appropriate feedback based on the emotional state identified by the emotion recognition engine. For example, if the server determines that the user is feeling stressed, a message such as "Relax" will be displayed on the screen of an online meeting application.

[0769] Specific situations and prompt sentence examples

[0770] Situation

[0771] User B sneezes during an online meeting and feels stressed.

[0772] 1. Audio capture

[0773] The device uses a microphone to capture user B's voice in real time and stores it in a buffer.

[0774] 2. Audio analysis

[0775] Noise reduction is applied to the voice data acquired by the device, and frequency components are extracted.

[0776] 3. Feature Extraction

[0777] The terminal calculates the Mel-frequency cepstral coefficients and generates a feature vector.

[0778] 4. Audio Classification

[0779] The device inputs the feature vector into a generative AI model to identify the time frame of the sneeze sound.

[0780] 5. Abnormal Audio Detection

[0781] The device records the time frame in which the sneeze occurred.

[0782] 6. Generation of out-of-phase speech

[0783] The terminal generates an audio signal that is in the opposite phase to the sneeze sound.

[0784] 7. Mixing of audio signals

[0785] The device mixes the original audio with an out-of-phase audio to cancel out the sneeze sound.

[0786] 8. Audio signal output

[0787] The terminal sends the processed audio signal to the online meeting.

[0788] 9. Emotion recognition

[0789] The device uses an emotion engine to recognize the emotional state of user B.

[0790] 10. Emotional Feedback

[0791] The server displays the message "Relax" based on User B's emotional state.

[0792] Prompt Sentence Examples

[0793] "When a user sneezes during an online meeting, create a program that analyzes the audio signal in real time, generates and mixes an opposite-phase audio signal, and cancels out the unpleasant sound. Also, use an emotion engine to identify the user's emotional state and provide feedback."

[0794] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0795] Step 1: Acquire audio

[0796] The device uses a microphone to capture the user's voice in real time. The input is the user's speech, and the output is the audio signal stored in the buffer. Specifically, the microphone is always on, and the user's voice is continuously sampled and stored in the buffer as digital data.

[0797] Step 2: Preprocessing the audio signal

[0798] The terminal performs digital signal processing (DSP) on the captured audio signal. The input is the raw audio signal stored in the buffer, and the output is the pre-processed audio signal. Specifically, it first applies a noise reduction algorithm to remove background noise. Then it divides the audio signal into frames and applies a Fourier transform to each frame to extract the frequency spectrum.

[0799] Step 3: Feature extraction

[0800] The device calculates Mel-Frequency Cepstral Coefficients (MFCCs) from the preprocessed audio signal. The input is the preprocessed audio signal, and the output is a feature vector. Specifically, the device calculates MFCCs for each frame and combines these MFCC values ​​to generate a single feature vector.

[0801] Step 4: Classify the audio

[0802] The device inputs the generated feature vector into a generative AI model to classify normal and abnormal speech. The input is the feature vector, and the output is the speech classification result. Specifically, the feature vector is input into the AI ​​model, and the AI ​​executes a process to determine whether the speech is normal or abnormal.

[0803] Step 5: Detecting Audio Anomalies

[0804] The device records the time when the abnormal sound occurred based on the classification result. The input is the sound classification result, and the output is the time information of the abnormal sound occurrence. Specifically, it identifies the time frame in which the AI ​​judged the sound to be abnormal and saves that information as a timestamp.

[0805] Step 6: Generate out-of-phase audio

[0806] The terminal generates an audio signal with an inverse phase to the detected abnormal audio portion. The input is the time frame of the abnormal audio and the original audio signal, and the output is an audio signal with an inverse phase. Specifically, it generates an inverse phase signal by inverting the frequency components of the abnormal audio and places that signal in the specified time frame.

[0807] Step 7: Mixing the audio signals

[0808] The terminal mixes the original audio signal with the generated out-of-phase audio signal. The input is the original audio signal and the out-of-phase audio signal, and the output is the mixed audio signal. Specifically, the original audio signal and the out-of-phase audio signal are added together to generate an audio signal that cancels the abnormal audio.

[0809] Step 8: Output the audio signal

[0810] The terminal transmits the processed audio data to the audio stream of the online meeting application. The input is the mixed audio signal, and the output is the processed audio that is transmitted to the participants of the online meeting. In specific operations, the terminal sends the processed audio signal to the transmission channel of the meeting application, allowing other participants to receive the audio with the abnormal sound canceled.

[0811] Step 9: Emotion Recognition

[0812] The device uses an emotion recognition engine to identify the user's emotional state from the voice signal. The input is the real-time voice signal, and the output is the user's emotional state. Specifically, the device analyzes the tone, pitch, and speed of the voice signal to determine emotions such as stress, joy, and sadness.

[0813] Step 10: Emotional Feedback

[0814] The server provides appropriate feedback according to the emotional state identified by the emotion recognition engine. The input is the user's emotional state, and the output is a feedback message. Specifically, the server generates a message that matches the user's emotional state and displays appropriate feedback, such as "Relax," on the screen of the online meeting application.

[0815] (Application example 2)

[0816] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0817] Robots in factories are required to accurately receive voice instructions from operators without misrecognition. However, if abnormal sounds such as sneezing or coughing are included in the voice, the risk of the robot malfunctioning increases. Furthermore, if the operator is stressed, it will affect work efficiency and safety. To solve these problems, a system is needed that combines the detection and removal of abnormal sounds and the recognition of the operator's emotional state.

[0818] The identification processing by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for acquiring a voice signal, artificial intelligence means for analyzing the acquired voice signal and distinguishing between normal voice and abnormal voice, means for generating an opposite-phase voice signal when an abnormal voice is detected, means for mixing the original voice signal with the opposite-phase voice signal to cancel out unpleasant voice, means for outputting a voice signal, emotion engine means for extracting features from the voice signal and identifying the emotional state, and means for providing feedback according to the emotional state. This allows the robot to accurately receive the operator's voice instructions, is not affected by abnormal sounds (such as sneezing or coughing), and can provide appropriate feedback according to the operator's emotional state.

[0819] The "means for acquiring an audio signal" refers to a means for obtaining audio data by a microphone or other audio input device.

[0820] "Generative artificial intelligence means for analyzing acquired voice signals and distinguishing between normal voices and abnormal voices" refers to a means for analyzing acquired voice data and using a generative AI to distinguish between normal voices and abnormal voices such as sneezing and coughing.

[0821] The "means for generating an audio signal of an opposite phase when an abnormal sound is detected" refers to a means for generating an audio signal of a phase shift of 180 degrees when an abnormal sound is detected in order to cancel out the sound.

[0822] The "means for mixing the original audio signal with an out-of-phase audio signal to cancel unpleasant sounds" is a means for canceling abnormal sounds, specifically unpleasant sounds, by combining the original audio with the generated out-of-phase audio.

[0823] "Means for outputting audio signals" refers to means for transmitting processed audio data to other systems or devices.

[0824] "Emotion engine means for extracting features from voice signals and identifying emotional states" refers to a means for extracting emotion-related features such as tone and pitch from voice data and analyzing and identifying the emotional state of the operator.

[0825] The "means for providing feedback according to emotional state" is a means for providing appropriate feedback or a message based on the emotional state of the operator identified by the emotion engine.

[0826] This invention is a system that enables robots in factories to accurately receive voice instructions from operators without misrecognition. In particular, it eliminates the effects of abnormal sounds such as sneezing and coughing, and provides feedback according to the operator's emotional state.

[0827] The system includes many means, but the main parts are as follows:

[0828] 1. Acquisition of audio signals:

[0829] The server uses a microphone or other audio input device to capture the operator's voice signals in real time, which consist of instructions and other environmental sounds in the factory.

[0830] 2. Means of analyzing the audio signal:

[0831] The server analyzes the captured audio signal using digital signal processing techniques. Specifically, it applies a noise reduction algorithm, divides the audio signal into frames, performs a Fourier transform, and extracts the audio frequency components. This process is performed using the feature.mfcc function from librosa.

[0832] 3. Generative AI means:

[0833] The server inputs the features extracted from the voice signal into a generative AI model, which distinguishes between normal voices and abnormal voices such as sneezing and coughing. This generative AI model uses a pre-trained convolutional neural network (CNN).

[0834] 4. Means for generating anti-phase audio signals:

[0835] When an abnormal sound is detected, the server generates an audio signal that is 180 degrees out of phase with the detected sound to cancel it out, effectively canceling the abnormal sound.

[0836] 5. Audio signal mixing means:

[0837] The server cancels out abnormal sounds by mixing the original audio signal with the generated out-of-phase audio, so that only normal instruction audio reaches the robot.

[0838] 6. Audio output means:

[0839] The server transmits the processed voice data to the robot so that the robot receives the voice instructions accurately.

[0840] 7. Emotion Engine Means:

[0841] The server analyzes the operator's emotional state using features extracted from the voice signal, using algorithms that use voice features such as tone and pitch.

[0842] 8. Feedback methods:

[0843] The server provides appropriate feedback based on the operator's emotional state as identified by the emotion engine. For example, if stress is detected, the robot will display a message encouraging the operator to take a break.

[0844] Specific examples

[0845] In a real-world scenario, an operator sneezes multiple times while working and then becomes extremely stressed. Detecting this state, the robot pauses the work and displays a message encouraging the operator to take a break. The system's processing steps are seamlessly handled by the server, and appropriate actions are automatically taken.

[0846] Prompt Sentence Examples

[0847] Build a system where if an operator sneezes multiple times and then becomes very stressed, the robot will display a message encouraging them to take a break.

[0848] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0849] Step 1:

[0850] The server uses a microphone to capture the operator's voice signal in real time. The voice signal obtained from this microphone includes environmental sounds in the factory and the operator's voice instructions. The input is an analog voice signal, and the output is digitized voice data. The server uses a buffer memory to store this data.

[0851] Step 2:

[0852] The server applies digital signal processing technology to the acquired audio signal, performs noise reduction, divides the audio signal into frames, performs a Fourier transform, and extracts the frequency components of the audio. The input is the digitized audio data obtained in step 1, and the output is frame data converted into frequency components.

[0853] Step 3:

[0854] The server calculates Mel-Frequency Cepstrum Coefficients (MFCC) from the frame data and extracts speech features. The input is the frequency component data obtained in step 2, and the output is a feature vector obtained using MFCC.

[0855] Step 4:

[0856] The server inputs the extracted feature vector into a generative AI model to classify normal and abnormal speech. The generative AI model identifies speech signal patterns based on past training data. The input is the feature vector obtained in step 3, and the output is a label: "normal speech" or "abnormal speech."

[0857] Step 5:

[0858] When an abnormal sound (sneezing or coughing) is detected, the server records its start and end times. The input is the classification label obtained in step 4, and the output is the time frame information of the abnormal sound.

[0859] Step 6:

[0860] The server generates an audio signal with an inverse phase to the detected abnormal audio. This inverse phase audio signal is an audio signal with a phase difference of 180 degrees to cancel out the original abnormal audio. The input is the time frame information of the abnormal audio obtained in step 5, and the output is the inverse phase audio signal.

[0861] Step 7:

[0862] The server cancels the unpleasant abnormal sounds by mixing the original sound signal with an out-of-phase sound signal. The input is the original sound signal obtained in step 1 and the out-of-phase sound signal generated in step 6, and the output is the cancelled sound signal.

[0863] Step 8:

[0864] The server sends the processed audio signal to the robot, which allows the robot to receive an audio signal that does not contain the canceled abnormal audio. The input is the canceled audio signal obtained in step 7, and the output is the digital audio signal received by the robot.

[0865] Step 9:

[0866] The server re-extracts features from the cancelled speech signal and inputs them into an emotion engine for identifying emotional states. The input is the cancelled speech signal obtained in step 7, and the output is a label indicating the emotional state.

[0867] Step 10:

[0868] The server provides appropriate feedback based on the emotional state identified by the emotion engine. For example, if many abnormal sounds are detected and the emotional state is identified as "stressed," the robot displays a message encouraging the operator to take a break. The input is the label indicating the emotional state obtained in step 9, and the output is the feedback message.

[0869] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0870] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0871] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0872] [Third embodiment]

[0873] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0874] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[0875] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0876] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0877] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0878] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0879] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0880] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0881] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0882] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0883] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0884] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0885] This invention relates to a system that detects physiological phenomena such as sneezing, coughing, and hiccups that occur during online meetings and automatically mutes them. This system acquires audio signals in real time, uses generation AI to distinguish between normal and abnormal sounds, and if an abnormal sound is detected, generates and mixes an opposite-phase audio signal to cancel out the unpleasant sound.

[0886] System configuration

[0887] 1. Audio acquisition method

[0888] Terminal: This system is equipped with a microphone to capture the user's voice in real time. The microphone continuously captures the voice signal.

[0889] 2. Voice Analysis Methods

[0890] Terminal: The captured audio signal is analyzed using digital signal processing techniques, which apply noise reduction algorithms, divide it into frames, and perform a Fourier transform to extract the audio frequency components.

[0891] 3. Feature Extraction Method

[0892] Terminal: Calculate features such as Mel-Frequency Cepstrum Coefficients (MFCC) from the preprocessed speech data and generate a feature vector, which represents the characteristics of the speech as numerical data.

[0893] 4. Abnormal sound detection method

[0894] Terminal: The feature vector is input into the generation AI to distinguish between normal and abnormal voices. In particular, when abnormal voices such as sneezing, coughing, and hiccups are detected, the start and end times are recorded.

[0895] 5. Antiphase speech generation means

[0896] Terminal: When an abnormal audio is detected, it generates an audio signal with the opposite phase to the abnormal audio portion. The opposite phase audio signal has the phase information necessary to cancel out the original audio signal.

[0897] 6. Audio Mixing Methods

[0898] Terminal: Mix the original audio signal with an inverted audio signal to cancel out any unwanted sounds. This operation ensures that the final audio is output smoothly.

[0899] 7. Audio output means

[0900] Terminal: The processed audio signal is sent to the online meeting application, which prevents other participants from hearing the disturbing audio.

[0901] Specific examples

[0902] Situation:

[0903] User A suddenly sneezes during an online meeting.

[0904] 1. Audio capture

[0905] Device: The microphone captures user A's voice in real time and stores it in a buffer.

[0906] 2. Voice Analysis

[0907] Terminal: Noise reduction is applied to the acquired audio data, it is divided into frames, and a Fourier transform is performed.

[0908] 3. Feature Extraction

[0909] Terminal: Calculate Mel-frequency cepstrum coefficients from the audio data and generate a feature vector.

[0910] 4. Abnormal Audio Detection

[0911] Terminal: The feature vector is input into the generative AI, which detects that the sneeze occurred between 1.5 and 1.8 seconds.

[0912] 5. Antiphase speech generation

[0913] Terminal: Generates an audio signal in anti-phase with a sneeze sound between 1.5 and 1.8 seconds.

[0914] 6. Audio Mixing

[0915] Terminal: Mixes the original audio signal with an out-of-phase audio signal to cancel the sneeze sound.

[0916] 7. Audio Output

[0917] Device: Sends the processed audio signal to the online meeting application, where other participants hear the audio without hearing any sneeze sounds.

[0918] Using this system, users can participate in online meetings with peace of mind and prevent unintended audio from being heard by others.

[0919] The processing flow will be explained below.

[0920] Step 1: Getting voice input

[0921] On the device: The microphone is activated and the user's voice is captured in real time. The captured voice data is stored in a buffer.

[0922] Step 2: Preprocessing the audio data

[0923] Terminal: A noise reduction algorithm is applied to the acquired audio data. The audio data is then divided into short time frames, and frequency components are extracted by performing a Fourier transform on each frame.

[0924] Step 3: Extract audio features

[0925] Terminal: To extract features from the preprocessed speech data, Mel-Frequency Cepstral Coefficients (MFCCs) are calculated, which generates a feature vector that quantifies the speech features.

[0926] Step 4: Classify normal and abnormal voices

[0927] Terminal: The generated feature vector is input into a generative artificial intelligence (generative AI) to classify normal and abnormal speech. This AI identifies patterns in the speech signal based on past training data and detects abnormal speech, such as sneezing or coughing.

[0928] Step 5: Detecting Audio Anomalies

[0929] Terminal: When the artificial intelligence detects an abnormal sound, it records its start and end time. For example, if a certain part of the audio signal is recognized as a sneeze, that time frame will be marked.

[0930] Step 6: Generate an anti-phase audio signal

[0931] Terminal: Generates an audio signal with an inverse phase to the detected abnormal audio portion. The inverse phase audio signal is an audio signal with a phase difference of 180 degrees to cancel the original abnormal audio signal.

[0932] Step 7: Mixing the audio signals

[0933] Terminal: Mix the original audio signal with the generated out-of-phase audio signal. This mixing cancels out the unpleasant abnormal sounds.

[0934] Step 8: Output the audio signal

[0935] Terminal: Sends the processed audio data to the audio stream of the online meeting, so that other participants hear the audio with the abnormal audio canceled.

[0936] Step 9: Real-time monitoring and adjustment

[0937] Terminal: All processes from audio signal acquisition to output are repeated in real time, continuously monitoring for abnormal audio and canceling it as appropriate.

[0938] In this way, the system of the present invention can automatically cancel unpleasant sounds such as sneezing and coughing that occur during an online meeting in real time, allowing users to maintain smooth conversations.

[0939] Example 1

[0940] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0941] During online meetings, physiological phenomena such as sneezing, coughing, and hiccups can interrupt conversations and cause discomfort to other participants. Because these sounds are not automatically muted, users may unintentionally let other participants hear the unpleasant sounds.

[0942] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0943] In this invention, the server includes means for acquiring voice data, means for preprocessing the acquired voice data using digital signal processing technology, means for extracting voice features from the preprocessed voice data, means for inputting the extracted voice features into a generative artificial intelligence model to distinguish between normal voice and abnormal voice, means for generating out-of-phase voice data when abnormal voice is detected, means for mixing the original voice data with the out-of-phase voice data to cancel unpleasant voice, and means for outputting the processed voice data. This prevents unpleasant sounds caused by physiological phenomena such as sneezing, coughing, and hiccups from reaching other participants during online meetings, enabling comfortable communication.

[0944] A "means for acquiring voice data" is a device or method for digitally capturing a user's speech or voice in real time using hardware such as a microphone.

[0945] "Digital signal processing" is the process of converting analog audio signals into digital form and using techniques such as noise reduction and Fourier transform to improve the quality and characteristics of the audio.

[0946] "Speech features" are numerical representations of the characteristics of a speech signal, and are typically extracted using techniques such as Mel-Frequency Cepstrum Coefficients (MFCC).

[0947] A "generative artificial intelligence model" is an algorithm or network that is trained to use machine learning or deep learning techniques to analyze speech data and distinguish between normal and abnormal speech.

[0948] "Out-of-phase audio data" is audio data that has the opposite phase to the original audio signal and is generated to cancel or reduce audio by mixing it with the original audio signal.

[0949] The "means for outputting audio data" refers to a device or method for transferring the processed audio data to a user's terminal or an online meeting application.

[0950] "Abnormal voices" are sounds caused by noises or physiological phenomena that are not part of normal conversation, such as sneezing, coughing, or hiccuping, and are detected by the generative artificial intelligence model.

[0951] This invention relates to a system that detects physiological phenomena such as sneezing, coughing, and hiccups that occur during online meetings and automatically mutes them. This system acquires audio signals in real time, uses a generative AI model to distinguish between normal and abnormal sounds, and when abnormal sounds are detected, generates and mixes an opposite-phase audio signal to cancel out the unpleasant sounds.

[0952] The system includes the following components:

[0953] 1. How to obtain audio data

[0954] Terminal: This system is equipped with a microphone to capture the user's voice in real time. Voice data is continuously acquired from the microphone.

[0955] 2. Preprocessing using digital signal processing

[0956] Terminal: The captured audio data is pre-processed using digital signal processing techniques, which apply noise reduction algorithms, divide the audio data into frames, perform Fourier transforms, and extract frequency components.

[0957] 3. Methods for extracting speech features

[0958] Terminal: Calculate Mel-Frequency Cepstral Coefficients (MFCC) from the preprocessed speech data to generate a feature vector, which represents the characteristics of the speech as numerical data.

[0959] 4. Method for detecting abnormal voice using generative AI model

[0960] On the device: The feature vector is input into a generative AI model to distinguish between normal and abnormal speech. In particular, when abnormal speech such as sneezing, coughing, or hiccuping is detected, the start and end times are recorded.

[0961] 5. Means for generating out-of-phase audio data

[0962] Terminal: When an abnormal audio is detected, it generates audio data with an inverse phase to the abnormal audio portion. The inverse phase audio data has the phase information necessary to cancel out the original audio data.

[0963] 6. Means of mixing audio data

[0964] Terminal: Mix the original audio data with the opposite phase audio data to cancel out the unpleasant sounds. This operation makes the final audio output smooth.

[0965] 7. Means of outputting audio data

[0966] Terminal: The processed audio data is sent to the online meeting application, which prevents other participants from hearing the disturbing audio.

[0967] Specific examples are shown below.

[0968] If user A suddenly sneezes during an online meeting, the following occurs:

[0969] 1. How to obtain audio data

[0970] Device: The microphone captures user A's voice in real time and stores it in a buffer.

[0971] 2. Preprocessing using digital signal processing

[0972] Terminal: Apply noise reduction to the acquired audio data, divide it into frames, and perform a Fourier transform.

[0973] 3. Methods for extracting speech features

[0974] Terminal: Calculate Mel-Frequency Cepstrum Coefficients (MFCC) from the audio data and generate a feature vector.

[0975] 4. Method for detecting abnormal voice using generative AI model

[0976] On the device: The feature vector is fed into a generative AI model to detect that the sneeze occurred between 1.5 and 1.8 seconds.

[0977] 5. Means for generating out-of-phase audio data

[0978] Device: Generates audio data in antiphase to a sneeze sound between 1.5 and 1.8 seconds.

[0979] 6. Means of mixing audio data

[0980] Terminal: Mix the original audio data with out-of-phase audio data to cancel the sneeze sound.

[0981] 7. Means of outputting audio data

[0982] Device: Sends the processed audio data to the online meeting application, where other participants hear the audio without hearing the sneeze.

[0983] This system allows users to participate in online meetings with peace of mind and prevents unintended audio from being heard by others.

[0984] Prompt Sentence Examples

[0985] "Please explain the system that automatically mutes you when you sneeze during an online meeting."

[0986] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0987] Step 1:

[0988] Acquiring audio data

[0989] Device: Uses a microphone to capture the user's voice data in real time.

[0990] Input: User speech and surrounding audio.

[0991] Output: Digital audio data from the microphone.

[0992] What it does: The microphone is always on, capturing audio data in real time and storing it in a fixed buffer.

[0993] Step 2:

[0994] Audio signal preprocessing

[0995] Terminal: Performs digital signal processing on the captured audio data, applying noise reduction algorithms, dividing the audio data into frames, and performing a Fourier transform to extract frequency components.

[0996] Input: Raw audio data.

[0997] Output: Denoised per-frame frequency spectrum.

[0998] What it does: It filters the audio signal to remove noise, then divides it into small time segments (frames), then performs a Fourier transform on each frame to calculate its frequency spectrum.

[0999] Step 3:

[1000] Feature extraction

[1001] Terminal: Calculate the Mel-Frequency Cepstral Coefficients (MFCC) from each frame and generate a feature vector.

[1002] Input: Frequency spectrum obtained by Fourier transform.

[1003] Output: MFCC-based feature vectors representing audio features.

[1004] Specific operation: Calculate MFCC using frequency spectrum as input, and save the calculation result as a feature vector in list format.

[1005] Step 4:

[1006] Abnormal audio detection

[1007] On the device: The feature vector is fed into a generative AI model to distinguish between normal and abnormal speech. If abnormal speech is detected, the start and end times of the speech are recorded.

[1008] Input: MFCC-based feature vector.

[1009] Output: Whether or not there is an abnormal sound, and its start and end times.

[1010] Specific operation: The generative AI model detects abnormal sounds based on the input feature vector and saves the period during which the abnormal sounds occurred as a timestamp.

[1011] Step 5:

[1012] Antiphase audio data generation

[1013] Terminal: Generates audio data in the opposite phase to the time range of the detected abnormal audio.

[1014] Input: Start and end times of the abnormal audio, and the original audio data.

[1015] Output: Out-of-phase audio data.

[1016] Specific operation: Based on the time range of the detected abnormal sound, calculate and save the opposite phase audio waveform.

[1017] Step 6:

[1018] Mixing of audio data

[1019] Terminal: Mixes the original audio data with out-of-phase audio data to cancel abnormal sounds.

[1020] Input: Original audio data, inverted audio data.

[1021] Output: Audio data with abnormal sounds canceled.

[1022] Specific operation: The original audio data is added to the out-of-phase audio data, and a process is performed to cancel out abnormal sounds.

[1023] Step 7:

[1024] Sending audio data

[1025] Terminal: Sends the processed audio data to the online meeting application.

[1026] Input: Mixed audio data.

[1027] Output: Final audio data (audio heard by other participants in the online meeting).

[1028] What it does: Stream the processed audio data in real time to an online meeting application.

[1029] (Application example 1)

[1030] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1031] In conventional online meetings and autonomous vehicle environments, sudden physiological events such as sneezing or coughing can cause unpleasant sounds to those around you and temporarily worsen the air quality, resulting in a decrease in voice quality, hindering conversations and the transmission of instructions, as well as reducing the comfort of the vehicle's interior.

[1032] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1033] In this invention, the server includes means for acquiring an audio signal, artificial intelligence means for analyzing the acquired audio signal and distinguishing between normal and abnormal audio, means for generating an opposite-phase audio signal when an abnormal audio signal is detected, means for mixing the original audio signal with the opposite-phase audio signal to cancel the unpleasant audio, means for outputting the audio signal, and means for temporarily strengthening related systems when canceling a new physiological phenomenon. This not only enables immediate detection and control of abnormal audio, but also temporarily improves the air environment, making it possible to maintain a comfortable and healthy environment.

[1034] plaintext

[1035] The "means for acquiring an audio signal" refers to a microphone and associated hardware for converting sound into an electronic signal.

[1036] The "generative artificial intelligence means" is an algorithm using artificial intelligence to analyze the acquired voice signal and distinguish between normal voice and abnormal voice.

[1037] The "means for generating an inverted phase audio signal" refers to a device or algorithm for generating an audio signal with the phase of an abnormal audio signal inverted when the abnormal audio signal is detected.

[1038] "Means for canceling unpleasant sounds" is a technology that synthesizes an audio signal that is out of phase with the original audio signal to cancel out unpleasant sounds.

[1039] The "means for outputting an audio signal" refers to a device such as a speaker or earphone for reproducing the processed audio signal.

[1040] The "means for temporarily enhancing related systems" is a control device for temporarily enhancing the air filter system and other related systems in the vehicle when a physiological phenomenon is detected.

[1041] plaintext

[1042] This invention is a system primarily designed to improve the riding experience in self-driving vehicles. It detects sudden physiological phenomena such as sneezing and coughing, cancels the unpleasant sounds that occur during these events, and temporarily improves the in-car environment.

[1043] The elements of the server, terminal, and user are as follows:

[1044] server

[1045] The vehicle is equipped with a microphone that captures audio signals and a computer system that includes an artificial intelligence algorithm for analyzing the audio. Audio signals are captured in real time and analyzed using digital signal processing technology. If an abnormal sound is detected, an opposite-phase audio signal is generated, thereby canceling out the unpleasant sound. This detection also temporarily strengthens the cabin air filter system.

[1046] Terminal

[1047] The terminal is a component of the in-vehicle system that acquires and analyzes audio signals in real time. Specifically, it performs noise reduction on the audio signal, frequency analysis using Fourier transform, and calculation of Mel Frequency Cepstrum Coefficients (MFCC). It also uses generative artificial intelligence to distinguish between normal and abnormal audio, and generates and mixes an inverse audio signal for detected abnormal audio to cancel it. It also includes a device that outputs the processed audio signal.

[1048] User

[1049] As passengers in self-driving vehicles, users can benefit from the ability to avoid physiological events such as sneezing or coughing while in an online meeting or using the navigation system, causing discomfort to other passengers and the system.

[1050] Specific examples

[1051] For example, the following prompt sentences can be used to input to a generative AI model to recreate a situation in which the present invention is implemented:

[1052] "Detect the sound of a sneeze in the car, mute the audio system at that moment, and strengthen the air filtration system."

[1053] This allows the in-car environment to be controlled quickly and automatically, providing a comfortable riding experience.

[1054] Hardware and Software

[1055] The system utilizes the following hardware and software:

[1056] Hardware: microphones, on-board computers, speakers, air filtration systems.

[1057] Software: Noise reduction algorithm, Fourier transform algorithm, Mel-Frequency Cepstral Coefficient (MFCC) extraction algorithm, generative AI model (TensorFlow, PyTorch, etc.).

[1058] By using the above prompt sentences, the generative AI model can immediately detect abnormal sounds and improve the air quality inside the car while canceling out unpleasant sounds.

[1059] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1060] plaintext

[1061] Step 1:

[1062] The server captures audio signals in real time from microphones inside the vehicle. The captured analog audio signals are converted to digital signals and passed to the next processing step. This input includes various sounds, such as ambient sounds inside the vehicle and conversation sounds.

[1063] Step 2:

[1064] The terminal performs noise reduction filtering on the captured voice signal, specifically using digital signal processing techniques to remove unwanted background noise, resulting in a clean voice signal being output.

[1065] Step 3:

[1066] The terminal divides the noise-reduced speech signal into frames and performs a Fourier transform to extract frequency components. The input is the speech signal after noise reduction, and the output is the frequency spectrum.

[1067] Step 4:

[1068] The device extracts speech features by calculating Mel-Frequency Cepstrum Coefficients (MFCC) from the frequency spectrum. This method expresses speech characteristics as numerical data. The input is the frequency spectrum, and the output is an MFCC feature vector.

[1069] Step 5:

[1070] The device inputs the MFCC feature vectors into a generative AI model (e.g., a model using TensorFlow or PyTorch) to distinguish between normal and abnormal speech. If abnormal speech (such as sneezing or coughing) is detected in this step, the device outputs the start and end times of the speech.

[1071] Step 6:

[1072] The terminal generates an audio signal with an opposite phase to the detected abnormal audio. To generate the opposite phase audio signal, an operation is performed to invert the phase of the original audio signal. The input is the audio signal of the abnormal audio, and the output is the audio signal with the opposite phase.

[1073] Step 7:

[1074] The terminal mixes the original audio signal with the generated out-of-phase audio signal, thereby performing a calculation to cancel the abnormal audio. The input is the original audio signal and the out-of-phase audio signal, and the output is the cancelled audio signal.

[1075] Step 8:

[1076] The device outputs the processed audio signal through the vehicle's speakers, allowing other passengers to hear the audio with the unpleasant sounds canceled out.

[1077] Step 9:

[1078] When an abnormal sound is detected, the server triggers a temporary strengthening of the air filter system inside the vehicle. Specifically, it increases the airflow of the air filter to quickly remove allergens and viruses in the air. This operation improves the air environment inside the vehicle.

[1079] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1080] This invention relates to a system that detects physiological phenomena such as sneezing, coughing, and hiccups that occur during online meetings and automatically mutes them. This system acquires audio signals in real time and uses generative AI to distinguish between normal and abnormal sounds. If abnormal sounds are detected, it generates and mixes an opposite-phase audio signal to cancel out the unpleasant sounds. Furthermore, by combining it with an emotion engine that recognizes the user's emotions, it has the ability to identify the user's emotional state and provide appropriate feedback in addition to processing the audio signals.

[1081] System configuration

[1082] 1. Audio acquisition method

[1083] Device: Equipped with a microphone for capturing audio signals in real time. Audio signals are continuously captured from this microphone.

[1084] 2. Voice Analysis Methods

[1085] Terminal: The captured audio data is analyzed using digital signal processing techniques, which apply a noise reduction algorithm, divide it into frames, and perform a Fourier transform to extract the frequency components of the audio.

[1086] 3. Feature Extraction Method

[1087] Terminal: To extract features from the preprocessed speech data, Mel-Frequency Cepstral Coefficients (MFCCs) are calculated, which generates a feature vector that quantifies the speech features.

[1088] 4. Classification of normal and abnormal voices

[1089] Terminal: The generated feature vector is input into the generative artificial intelligence (generative AI) to classify normal and abnormal voices. The generative AI identifies patterns in the voice signal based on past training data and detects abnormal voices such as sneezing and coughing.

[1090] 5. Abnormal Audio Detection

[1091] Terminal: When an abnormal sound is detected, its start and end time are recorded. For example, if a certain part of the sound signal is recognized as a sneeze, that time frame is marked.

[1092] 6. Generation of anti-phase audio signals

[1093] Terminal: Generates an audio signal with an inverse phase to the detected abnormal audio portion. The inverse phase audio signal is an audio signal with a phase difference of 180 degrees to cancel the original abnormal audio signal.

[1094] 7. Mixing of audio signals

[1095] Terminal: Mix the original audio signal with the generated out-of-phase audio signal. This mixing cancels out the unpleasant abnormal sounds.

[1096] 8. Audio signal output

[1097] Terminal: Sends the processed audio data to the audio stream of the online meeting, so that other participants hear the audio with the abnormal audio canceled.

[1098] 9. Emotion Engine

[1099] Device: The emotion engine has algorithms to identify the emotional state from the user's voice signal. The emotion engine analyzes voice characteristics such as tone, pitch, and rate to identify the emotion the user is feeling.

[1100] Server: Provides different feedback depending on the emotional state identified by the emotion engine. For example, if the server determines that the user is feeling stressed, it displays a message on the screen suggesting that the user relax.

[1101] Specific examples

[1102] Situation:

[1103] User B suddenly sneezes during an online meeting and feels stressed.

[1104] 1. Audio capture

[1105] Device: The microphone captures user B's voice in real time and stores it in a buffer.

[1106] 2. Voice Analysis

[1107] Terminal: Noise reduction is applied to the acquired audio data, it is divided into frames, and a Fourier transform is performed.

[1108] 3. Feature Extraction

[1109] Terminal: Calculate Mel-frequency cepstrum coefficients from the audio data and generate a feature vector.

[1110] 4. Classification of normal and abnormal voices

[1111] On the device: The feature vector is fed into a generative AI to detect that a sneeze occurred within a specific time frame.

[1112] 5. Abnormal Audio Detection

[1113] Device: Detects that the sneeze occurred between 1.5 and 1.8 seconds and marks that time frame.

[1114] 6. Generation of anti-phase audio signals

[1115] Device: Generates an audio signal in antiphase to a sneeze sound lasting 1.5 to 1.8 seconds.

[1116] 7. Mixing of audio signals

[1117] Terminal: Mixes the original audio signal with an out-of-phase audio signal to cancel the sneeze sound.

[1118] 8. Audio signal output

[1119] Terminal: Sends the processed audio signal to the online meeting application. Other participants receive the audio without hearing the sneeze.

[1120] 9. Emotion recognition

[1121] Device: The emotion engine identifies User B's emotional state from audio signals including the sneeze sound.

[1122] 10. Emotional Feedback

[1123] Server: If the emotion engine identifies User B's emotional state as "stressed," a message such as "Relax" will be displayed on the screen of the online meeting application.

[1124] Using this system, users can not only automatically cancel unpleasant sounds such as sneezing and coughing during online meetings in real time, but also receive support based on their emotional state using an emotion engine, allowing users to participate in online meetings with greater peace of mind.

[1125] The processing flow will be explained below.

[1126] Step 1: Getting voice input

[1127] On the device: The microphone is activated and the user's voice is captured in real time. The captured voice data is stored in a buffer.

[1128] Step 2: Preprocessing the audio data

[1129] Terminal: A noise reduction algorithm is applied to the acquired audio data to obtain a clear audio signal. The audio data is then divided into short time frames, and a Fourier transform is performed on each frame to extract the frequency components of the audio.

[1130] Step 3: Extract audio features

[1131] Terminal: To extract features from the preprocessed speech data, Mel-Frequency Cepstral Coefficients (MFCCs) are calculated, which generates a feature vector that quantifies the speech features.

[1132] Step 4: Classify normal and abnormal voices

[1133] Terminal: The generated feature vector is input into the generative artificial intelligence (generative AI) to classify normal and abnormal voices. The generative AI identifies patterns in the voice signal based on past training data and detects abnormal voices such as sneezing and coughing.

[1134] Step 5: Detecting Audio Anomalies

[1135] Terminal: When abnormal sounds are detected by the AI ​​generator, their start and end times are recorded. For example, if the audio signal is recognized as a sneeze from 1.5 seconds to 1.8 seconds, that time frame is marked.

[1136] Step 6: Generate an anti-phase audio signal

[1137] Terminal: Generates an audio signal with an inverse phase to the detected abnormal audio portion. The inverse phase audio signal is an audio signal with a phase difference of 180 degrees to cancel the original abnormal audio signal.

[1138] Step 7: Mixing the audio signals

[1139] Terminal: Mix the original audio signal with the generated out-of-phase audio signal. This mixing cancels out the unpleasant abnormal sounds.

[1140] Step 8: Output the audio signal

[1141] Terminal: Sends the processed audio data to the audio stream of the online meeting, so that other participants hear the audio with the abnormal audio canceled.

[1142] Step 9: Emotion Recognition Processing

[1143] On the device: The voice signal is processed by an emotion engine that analyzes characteristics such as tone, pitch, and rate to identify the user's emotional state.

[1144] On the device: The emotion engine identifies the user's emotion based on the analyzed features. For example, if the voice is spoken with a strong tone, high pitch, and rapid speed, it will be identified as stressed or angry.

[1145] Step 10: Emotional State Feedback

[1146] Server: Provides appropriate feedback based on the emotional state identified by the emotion engine. For example, if the server determines that the user is feeling stressed, it displays a message such as "Relax" on the screen of an online meeting application.

[1147] Server: Based on the user's emotional state, it provides a feedback function that notifies other participants of the user's state, depending on the settings.

[1148] In this way, the entire system performs a series of processes in real time, from acquiring voice signals to canceling unpleasant sounds, and then identifying the user's emotions and providing appropriate feedback. This allows users to continue participating in online meetings with peace of mind, even when experiencing physiological or emotional changes.

[1149] Example 2

[1150] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1151] Physiological phenomena such as sneezing, coughing, and hiccups that occur during online meetings can be unpleasant for other participants and disrupt the flow of conversation. Furthermore, users often feel stressed if they become too aware of these physiological phenomena. Furthermore, it is difficult to grasp users' emotional state in real time, and there is a lack of means to provide appropriate feedback. This can reduce the quality of online meetings and impair participants' concentration and comfort.

[1152] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1153] In this invention, the terminal includes means for acquiring an audio signal, artificial intelligence means for analyzing the acquired audio signal and distinguishing between normal audio and abnormal audio, means for generating an opposite-phase audio signal when abnormal audio is detected, means for mixing the original audio signal with the opposite-phase audio signal and canceling unpleasant audio, means for outputting an audio signal, and emotion recognition means for recognizing the user's emotional state and providing feedback. This makes it possible to automatically cancel unpleasant physiological sounds during online meetings and provide support according to the user's emotional state.

[1154] The "audio signal" is a user's speech converted into an electrical signal, and is acquired through an input device such as a microphone.

[1155] "Means for acquiring" refers to a device or method for acquiring audio signals in real time using a microphone, sensor device, etc.

[1156] "Analyzing" is the process of analyzing an audio signal using digital signal processing techniques to extract various features.

[1157] "Normal voice" refers to voice that is not an abnormal voice, such as a normal conversation or speech.

[1158] "Abnormal sounds" refers to physiological sounds that interfere with normal conversation, such as sneezing, coughing, and hiccups.

[1159] "Artificial intelligence tools" are algorithms or models that use historical training data to identify speech patterns and categorize or differentiate between speech sounds.

[1160] "Detecting abnormal sounds" means identifying abnormal sounds from the acquired audio signals and recognizing the time of occurrence and their content.

[1161] An "opposite phase audio signal" is an audio signal generated by reversing the phase of the original abnormal audio signal by 180 degrees, and is used to cancel the abnormal audio.

[1162] "Mixing" is a process of synthesizing an audio signal that is out of phase with the original audio signal to cancel out abnormal audio.

[1163] "Unpleasant sounds" are sounds generated by physiological phenomena such as sneezing, coughing, and hiccuping, which are unpleasant to other participants and disrupt conversation.

[1164] "Cancellation means" refers to a method for filtering out abnormal sounds through a mixed audio signal, allowing other sounds to be heard normally.

[1165] "Means for outputting audio signals" refers to devices or methods for converting the processed audio data back into an audio stream for transmission to other participants.

[1166] "Emotion recognition means" refers to an algorithm or system that grasps the user's emotional state from their voice signal and analyzes it in real time.

[1167] A "means for providing feedback" is a system or method for displaying appropriate advice or suggestions to a user depending on the perceived emotional state.

[1168] This invention is a system that detects unpleasant physiological sounds (sneezing, coughing, hiccups, etc.) that occur during online meetings in real time and automatically cancels them. This system also aims to improve the quality of online meetings by recognizing the user's emotional state in real time and providing appropriate feedback.

[1169] System Configuration

[1170] Audio acquisition means

[1171] The device has a built-in microphone that captures the user's voice signal in real time, which is then stored in a buffer and used in subsequent processing steps.

[1172] Voice analysis methods

[1173] The captured audio signal is pre-processed using digital signal processing techniques, which involves applying a noise reduction algorithm to remove background noise, dividing the signal into frames, and then performing a Fourier transform to extract the audio frequency components.

[1174] Feature extraction method

[1175] Mel-frequency cepstral coefficients (MFCCs) are calculated from the preprocessed speech data, and a feature vector is generated that quantifies the speech features, enabling detailed analysis of the speech signal.

[1176] Classification of normal and abnormal voices

[1177] The device inputs the generated feature vector into a generative AI model to distinguish between normal and abnormal voices. The generative AI model uses past training data to identify patterns in the voice signal and detect abnormal sounds (e.g., sneezing, coughing).

[1178] Abnormal voice detection

[1179] If the generative AI model detects an abnormal sound, it records the time it occurred. For example, if a sneeze occurs between 1.5 and 1.8 seconds, that time frame will be recorded.

[1180] Generation of out-of-phase speech

[1181] An audio signal with an opposite phase is generated for the detected abnormal sound portion. The opposite phase audio signal is an audio signal with a phase difference of 180 degrees to cancel the original abnormal sound signal.

[1182] Audio signal mixing

[1183] By mixing the original audio signal with the generated out-of-phase audio signal, the unpleasant abnormal sounds are cancelled out.

[1184] Audio signal output

[1185] The processed audio data is sent to the audio stream of the online meeting application, so that other participants can hear the audio with the abnormal sounds canceled out.

[1186] emotion recognition means

[1187] The device is equipped with an emotion recognition engine that recognizes the user's emotional state from their voice signal. The engine analyzes the tone, pitch, and speed of the voice to determine the user's current emotional state.

[1188] Feedback methods

[1189] The server then provides appropriate feedback based on the emotional state identified by the emotion recognition engine. For example, if the server determines that the user is feeling stressed, a message such as "Relax" will be displayed on the screen of an online meeting application.

[1190] Specific situations and prompt sentence examples

[1191] Situation

[1192] User B sneezes during an online meeting and feels stressed.

[1193] 1. Audio capture

[1194] The device uses a microphone to capture user B's voice in real time and stores it in a buffer.

[1195] 2. Audio analysis

[1196] Noise reduction is applied to the voice data acquired by the device, and frequency components are extracted.

[1197] 3. Feature Extraction

[1198] The terminal calculates the Mel-frequency cepstral coefficients and generates a feature vector.

[1199] 4. Audio Classification

[1200] The device inputs the feature vector into a generative AI model to identify the time frame of the sneeze sound.

[1201] 5. Abnormal Audio Detection

[1202] The device records the time frame in which the sneeze occurred.

[1203] 6. Generation of out-of-phase speech

[1204] The terminal generates an audio signal that is in the opposite phase to the sneeze sound.

[1205] 7. Mixing of audio signals

[1206] The device mixes the original audio with an out-of-phase audio to cancel out the sneeze sound.

[1207] 8. Audio signal output

[1208] The terminal sends the processed audio signal to the online meeting.

[1209] 9. Emotion recognition

[1210] The device uses an emotion engine to recognize the emotional state of user B.

[1211] 10. Emotional Feedback

[1212] The server displays the message "Relax" based on User B's emotional state.

[1213] Prompt Sentence Examples

[1214] "When a user sneezes during an online meeting, create a program that analyzes the audio signal in real time, generates and mixes an opposite-phase audio signal, and cancels out the unpleasant sound. Also, use an emotion engine to identify the user's emotional state and provide feedback."

[1215] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1216] Step 1: Acquire audio

[1217] The device uses a microphone to capture the user's voice in real time. The input is the user's speech, and the output is the audio signal stored in the buffer. Specifically, the microphone is always on, and the user's voice is continuously sampled and stored in the buffer as digital data.

[1218] Step 2: Preprocessing the audio signal

[1219] The terminal performs digital signal processing (DSP) on the captured audio signal. The input is the raw audio signal stored in the buffer, and the output is the pre-processed audio signal. Specifically, it first applies a noise reduction algorithm to remove background noise. Then it divides the audio signal into frames and applies a Fourier transform to each frame to extract the frequency spectrum.

[1220] Step 3: Feature extraction

[1221] The device calculates Mel-Frequency Cepstral Coefficients (MFCCs) from the preprocessed audio signal. The input is the preprocessed audio signal, and the output is a feature vector. Specifically, the device calculates MFCCs for each frame and combines these MFCC values ​​to generate a single feature vector.

[1222] Step 4: Classify the audio

[1223] The device inputs the generated feature vector into a generative AI model to classify normal and abnormal speech. The input is the feature vector, and the output is the speech classification result. Specifically, the feature vector is input into the AI ​​model, and the AI ​​executes a process to determine whether the speech is normal or abnormal.

[1224] Step 5: Detecting Audio Anomalies

[1225] The device records the time when the abnormal sound occurred based on the classification result. The input is the sound classification result, and the output is the time information of the abnormal sound occurrence. Specifically, it identifies the time frame in which the AI ​​judged the sound to be abnormal and saves that information as a timestamp.

[1226] Step 6: Generate out-of-phase audio

[1227] The terminal generates an audio signal with an inverse phase to the detected abnormal audio portion. The input is the time frame of the abnormal audio and the original audio signal, and the output is an audio signal with an inverse phase. Specifically, it generates an inverse phase signal by inverting the frequency components of the abnormal audio and places that signal in the specified time frame.

[1228] Step 7: Mixing the audio signals

[1229] The terminal mixes the original audio signal with the generated out-of-phase audio signal. The input is the original audio signal and the out-of-phase audio signal, and the output is the mixed audio signal. Specifically, the original audio signal and the out-of-phase audio signal are added together to generate an audio signal that cancels the abnormal audio.

[1230] Step 8: Output the audio signal

[1231] The terminal transmits the processed audio data to the audio stream of the online meeting application. The input is the mixed audio signal, and the output is the processed audio that is transmitted to the participants of the online meeting. In specific operations, the terminal sends the processed audio signal to the transmission channel of the meeting application, allowing other participants to receive the audio with the abnormal sound canceled.

[1232] Step 9: Emotion Recognition

[1233] The device uses an emotion recognition engine to identify the user's emotional state from the voice signal. The input is the real-time voice signal, and the output is the user's emotional state. Specifically, the device analyzes the tone, pitch, and speed of the voice signal to determine emotions such as stress, joy, and sadness.

[1234] Step 10: Emotional Feedback

[1235] The server provides appropriate feedback according to the emotional state identified by the emotion recognition engine. The input is the user's emotional state, and the output is a feedback message. Specifically, the server generates a message that matches the user's emotional state and displays appropriate feedback, such as "Relax," on the screen of the online meeting application.

[1236] (Application example 2)

[1237] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1238] Robots in factories are required to accurately receive voice instructions from operators without misrecognition. However, if abnormal sounds such as sneezing or coughing are included in the voice, the risk of the robot malfunctioning increases. Furthermore, if the operator is stressed, it will affect work efficiency and safety. To solve these problems, a system is needed that combines the detection and removal of abnormal sounds and the recognition of the operator's emotional state.

[1239] The identification processing by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for acquiring a voice signal, artificial intelligence means for analyzing the acquired voice signal and distinguishing between normal voice and abnormal voice, means for generating an opposite-phase voice signal when an abnormal voice is detected, means for mixing the original voice signal with the opposite-phase voice signal to cancel out unpleasant voice, means for outputting a voice signal, emotion engine means for extracting features from the voice signal and identifying the emotional state, and means for providing feedback according to the emotional state. This allows the robot to accurately receive the operator's voice instructions, is not affected by abnormal sounds (such as sneezing or coughing), and can provide appropriate feedback according to the operator's emotional state.

[1240] The "means for acquiring an audio signal" refers to a means for obtaining audio data by a microphone or other audio input device.

[1241] "Generative artificial intelligence means for analyzing acquired voice signals and distinguishing between normal voices and abnormal voices" refers to a means for analyzing acquired voice data and using a generative AI to distinguish between normal voices and abnormal voices such as sneezing and coughing.

[1242] The "means for generating an audio signal of an opposite phase when an abnormal sound is detected" refers to a means for generating an audio signal of a phase shift of 180 degrees when an abnormal sound is detected in order to cancel out the sound.

[1243] The "means for mixing the original audio signal with an out-of-phase audio signal to cancel unpleasant sounds" is a means for canceling abnormal sounds, specifically unpleasant sounds, by combining the original audio with the generated out-of-phase audio.

[1244] "Means for outputting audio signals" refers to means for transmitting processed audio data to other systems or devices.

[1245] "Emotion engine means for extracting features from voice signals and identifying emotional states" refers to a means for extracting emotion-related features such as tone and pitch from voice data and analyzing and identifying the emotional state of the operator.

[1246] The "means for providing feedback according to emotional state" is a means for providing appropriate feedback or a message based on the emotional state of the operator identified by the emotion engine.

[1247] This invention is a system that enables robots in factories to accurately receive voice instructions from operators without misrecognition. In particular, it eliminates the effects of abnormal sounds such as sneezing and coughing, and provides feedback according to the operator's emotional state.

[1248] The system includes many means, but the main parts are as follows:

[1249] 1. Acquisition of audio signals:

[1250] The server uses a microphone or other audio input device to capture the operator's voice signals in real time, which consist of instructions and other environmental sounds in the factory.

[1251] 2. Means of analyzing the audio signal:

[1252] The server analyzes the captured audio signal using digital signal processing techniques. Specifically, it applies a noise reduction algorithm, divides the audio signal into frames, performs a Fourier transform, and extracts the audio frequency components. This process is performed using the feature.mfcc function from librosa.

[1253] 3. Generative AI means:

[1254] The server inputs the features extracted from the voice signal into a generative AI model, which distinguishes between normal voices and abnormal voices such as sneezing and coughing. This generative AI model uses a pre-trained convolutional neural network (CNN).

[1255] 4. Means for generating anti-phase audio signals:

[1256] When an abnormal sound is detected, the server generates an audio signal that is 180 degrees out of phase with the detected sound to cancel it out, effectively canceling the abnormal sound.

[1257] 5. Audio signal mixing means:

[1258] The server cancels out abnormal sounds by mixing the original audio signal with the generated out-of-phase audio, so that only normal instruction audio reaches the robot.

[1259] 6. Audio output means:

[1260] The server transmits the processed voice data to the robot so that the robot receives the voice instructions accurately.

[1261] 7. Emotion Engine Means:

[1262] The server analyzes the operator's emotional state using features extracted from the voice signal, using algorithms that use voice features such as tone and pitch.

[1263] 8. Feedback methods:

[1264] The server provides appropriate feedback based on the operator's emotional state as identified by the emotion engine. For example, if stress is detected, the robot will display a message encouraging the operator to take a break.

[1265] Specific examples

[1266] In a real-world scenario, an operator sneezes multiple times while working and then becomes extremely stressed. Detecting this state, the robot pauses the work and displays a message encouraging the operator to take a break. The system's processing steps are seamlessly handled by the server, and appropriate actions are automatically taken.

[1267] Prompt Sentence Examples

[1268] Build a system where if an operator sneezes multiple times and then becomes very stressed, the robot will display a message encouraging them to take a break.

[1269] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1270] Step 1:

[1271] The server uses a microphone to capture the operator's voice signal in real time. The voice signal obtained from this microphone includes environmental sounds in the factory and the operator's voice instructions. The input is an analog voice signal, and the output is digitized voice data. The server uses a buffer memory to store this data.

[1272] Step 2:

[1273] The server applies digital signal processing technology to the acquired audio signal, performs noise reduction, divides the audio signal into frames, performs a Fourier transform, and extracts the frequency components of the audio. The input is the digitized audio data obtained in step 1, and the output is frame data converted into frequency components.

[1274] Step 3:

[1275] The server calculates Mel-Frequency Cepstrum Coefficients (MFCC) from the frame data and extracts speech features. The input is the frequency component data obtained in step 2, and the output is a feature vector obtained using MFCC.

[1276] Step 4:

[1277] The server inputs the extracted feature vector into a generative AI model to classify normal and abnormal speech. The generative AI model identifies speech signal patterns based on past training data. The input is the feature vector obtained in step 3, and the output is a label: "normal speech" or "abnormal speech."

[1278] Step 5:

[1279] When an abnormal sound (sneezing or coughing) is detected, the server records its start and end times. The input is the classification label obtained in step 4, and the output is the time frame information of the abnormal sound.

[1280] Step 6:

[1281] The server generates an audio signal with an inverse phase to the detected abnormal audio. This inverse phase audio signal is an audio signal with a phase difference of 180 degrees to cancel out the original abnormal audio. The input is the time frame information of the abnormal audio obtained in step 5, and the output is the inverse phase audio signal.

[1282] Step 7:

[1283] The server cancels the unpleasant abnormal sounds by mixing the original sound signal with an out-of-phase sound signal. The input is the original sound signal obtained in step 1 and the out-of-phase sound signal generated in step 6, and the output is the cancelled sound signal.

[1284] Step 8:

[1285] The server sends the processed audio signal to the robot, which allows the robot to receive an audio signal that does not contain the canceled abnormal audio. The input is the canceled audio signal obtained in step 7, and the output is the digital audio signal received by the robot.

[1286] Step 9:

[1287] The server re-extracts features from the cancelled speech signal and inputs them into an emotion engine for identifying emotional states. The input is the cancelled speech signal obtained in step 7, and the output is a label indicating the emotional state.

[1288] Step 10:

[1289] The server provides appropriate feedback based on the emotional state identified by the emotion engine. For example, if many abnormal sounds are detected and the emotional state is identified as "stressed," the robot displays a message encouraging the operator to take a break. The input is the label indicating the emotional state obtained in step 9, and the output is the feedback message.

[1290] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1291] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1292] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1293] [Fourth embodiment]

[1294] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1295] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1296] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1297] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1298] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1299] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1300] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1301] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1302] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1303] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1304] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1305] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1306] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1307] This invention relates to a system that detects physiological phenomena such as sneezing, coughing, and hiccups that occur during online meetings and automatically mutes them. This system acquires audio signals in real time, uses generation AI to distinguish between normal and abnormal sounds, and if an abnormal sound is detected, generates and mixes an opposite-phase audio signal to cancel out the unpleasant sound.

[1308] System configuration

[1309] 1. Audio acquisition method

[1310] Terminal: This system is equipped with a microphone to capture the user's voice in real time. The microphone continuously captures the voice signal.

[1311] 2. Voice Analysis Methods

[1312] Terminal: The captured audio signal is analyzed using digital signal processing techniques, which apply noise reduction algorithms, divide it into frames, and perform a Fourier transform to extract the audio frequency components.

[1313] 3. Feature Extraction Method

[1314] Terminal: Calculate features such as Mel-Frequency Cepstrum Coefficients (MFCC) from the preprocessed speech data and generate a feature vector, which represents the characteristics of the speech as numerical data.

[1315] 4. Abnormal sound detection method

[1316] Terminal: The feature vector is input into the generation AI to distinguish between normal and abnormal voices. In particular, when abnormal voices such as sneezing, coughing, and hiccups are detected, the start and end times are recorded.

[1317] 5. Antiphase speech generation means

[1318] Terminal: When an abnormal audio is detected, it generates an audio signal with the opposite phase to the abnormal audio portion. The opposite phase audio signal has the phase information necessary to cancel out the original audio signal.

[1319] 6. Audio Mixing Methods

[1320] Terminal: Mix the original audio signal with an inverted audio signal to cancel out any unwanted sounds. This operation ensures that the final audio is output smoothly.

[1321] 7. Audio output means

[1322] Terminal: The processed audio signal is sent to the online meeting application, which prevents other participants from hearing the disturbing audio.

[1323] Specific examples

[1324] Situation:

[1325] User A suddenly sneezes during an online meeting.

[1326] 1. Audio capture

[1327] Device: The microphone captures user A's voice in real time and stores it in a buffer.

[1328] 2. Voice Analysis

[1329] Terminal: Noise reduction is applied to the acquired audio data, it is divided into frames, and a Fourier transform is performed.

[1330] 3. Feature Extraction

[1331] Terminal: Calculate Mel-frequency cepstrum coefficients from the audio data and generate a feature vector.

[1332] 4. Abnormal Audio Detection

[1333] Terminal: The feature vector is input into the generative AI, which detects that the sneeze occurred between 1.5 and 1.8 seconds.

[1334] 5. Antiphase speech generation

[1335] Terminal: Generates an audio signal in anti-phase with a sneeze sound between 1.5 and 1.8 seconds.

[1336] 6. Audio Mixing

[1337] Terminal: Mixes the original audio signal with an out-of-phase audio signal to cancel the sneeze sound.

[1338] 7. Audio Output

[1339] Device: Sends the processed audio signal to the online meeting application, where other participants hear the audio without hearing any sneeze sounds.

[1340] Using this system, users can participate in online meetings with peace of mind and prevent unintended audio from being heard by others.

[1341] The processing flow will be explained below.

[1342] Step 1: Getting voice input

[1343] On the device: The microphone is activated and the user's voice is captured in real time. The captured voice data is stored in a buffer.

[1344] Step 2: Preprocessing the audio data

[1345] Terminal: A noise reduction algorithm is applied to the acquired audio data. The audio data is then divided into short time frames, and frequency components are extracted by performing a Fourier transform on each frame.

[1346] Step 3: Extract audio features

[1347] Terminal: To extract features from the preprocessed speech data, Mel-Frequency Cepstral Coefficients (MFCCs) are calculated, which generates a feature vector that quantifies the speech features.

[1348] Step 4: Classify normal and abnormal voices

[1349] Terminal: The generated feature vector is input into a generative artificial intelligence (generative AI) to classify normal and abnormal speech. This AI identifies patterns in the speech signal based on past training data and detects abnormal speech, such as sneezing or coughing.

[1350] Step 5: Detecting Audio Anomalies

[1351] Terminal: When the artificial intelligence detects an abnormal sound, it records its start and end time. For example, if a certain part of the audio signal is recognized as a sneeze, that time frame will be marked.

[1352] Step 6: Generate an anti-phase audio signal

[1353] Terminal: Generates an audio signal with an inverse phase to the detected abnormal audio portion. The inverse phase audio signal is an audio signal with a phase difference of 180 degrees to cancel the original abnormal audio signal.

[1354] Step 7: Mixing the audio signals

[1355] Terminal: Mix the original audio signal with the generated out-of-phase audio signal. This mixing cancels out the unpleasant abnormal sounds.

[1356] Step 8: Output the audio signal

[1357] Terminal: Sends the processed audio data to the audio stream of the online meeting, so that other participants hear the audio with the abnormal audio canceled.

[1358] Step 9: Real-time monitoring and adjustment

[1359] Terminal: All processes from audio signal acquisition to output are repeated in real time, continuously monitoring for abnormal audio and canceling it as appropriate.

[1360] In this way, the system of the present invention can automatically cancel unpleasant sounds such as sneezing and coughing that occur during an online meeting in real time, allowing users to maintain smooth conversations.

[1361] Example 1

[1362] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1363] During online meetings, physiological phenomena such as sneezing, coughing, and hiccups can interrupt conversations and cause discomfort to other participants. Because these sounds are not automatically muted, users may unintentionally let other participants hear the unpleasant sounds.

[1364] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1365] In this invention, the server includes means for acquiring voice data, means for preprocessing the acquired voice data using digital signal processing technology, means for extracting voice features from the preprocessed voice data, means for inputting the extracted voice features into a generative artificial intelligence model to distinguish between normal voice and abnormal voice, means for generating out-of-phase voice data when abnormal voice is detected, means for mixing the original voice data with the out-of-phase voice data to cancel unpleasant voice, and means for outputting the processed voice data. This prevents unpleasant sounds caused by physiological phenomena such as sneezing, coughing, and hiccups from reaching other participants during online meetings, enabling comfortable communication.

[1366] A "means for acquiring voice data" is a device or method for digitally capturing a user's speech or voice in real time using hardware such as a microphone.

[1367] "Digital signal processing" is the process of converting analog audio signals into digital form and using techniques such as noise reduction and Fourier transform to improve the quality and characteristics of the audio.

[1368] "Speech features" are numerical representations of the characteristics of a speech signal, and are typically extracted using techniques such as Mel-Frequency Cepstrum Coefficients (MFCC).

[1369] A "generative artificial intelligence model" is an algorithm or network that is trained to use machine learning or deep learning techniques to analyze speech data and distinguish between normal and abnormal speech.

[1370] "Out-of-phase audio data" is audio data that has the opposite phase to the original audio signal and is generated to cancel or reduce audio by mixing it with the original audio signal.

[1371] The "means for outputting audio data" refers to a device or method for transferring the processed audio data to a user's terminal or an online meeting application.

[1372] "Abnormal voices" are sounds caused by noises or physiological phenomena that are not part of normal conversation, such as sneezing, coughing, or hiccuping, and are detected by the generative artificial intelligence model.

[1373] This invention relates to a system that detects physiological phenomena such as sneezing, coughing, and hiccups that occur during online meetings and automatically mutes them. This system acquires audio signals in real time, uses a generative AI model to distinguish between normal and abnormal sounds, and when abnormal sounds are detected, generates and mixes an opposite-phase audio signal to cancel out the unpleasant sounds.

[1374] The system includes the following components:

[1375] 1. How to obtain audio data

[1376] Terminal: This system is equipped with a microphone to capture the user's voice in real time. Voice data is continuously acquired from the microphone.

[1377] 2. Preprocessing using digital signal processing

[1378] Terminal: The captured audio data is pre-processed using digital signal processing techniques, which apply noise reduction algorithms, divide the audio data into frames, perform Fourier transforms, and extract frequency components.

[1379] 3. Methods for extracting speech features

[1380] Terminal: Calculate Mel-Frequency Cepstral Coefficients (MFCC) from the preprocessed speech data to generate a feature vector, which represents the characteristics of the speech as numerical data.

[1381] 4. Method for detecting abnormal voice using generative AI model

[1382] On the device: The feature vector is input into a generative AI model to distinguish between normal and abnormal speech. In particular, when abnormal speech such as sneezing, coughing, or hiccuping is detected, the start and end times are recorded.

[1383] 5. Means for generating out-of-phase audio data

[1384] Terminal: When an abnormal audio is detected, it generates audio data with an inverse phase to the abnormal audio portion. The inverse phase audio data has the phase information necessary to cancel out the original audio data.

[1385] 6. Means of mixing audio data

[1386] Terminal: Mix the original audio data with the opposite phase audio data to cancel out the unpleasant sounds. This operation makes the final audio output smooth.

[1387] 7. Means of outputting audio data

[1388] Terminal: The processed audio data is sent to the online meeting application, which prevents other participants from hearing the disturbing audio.

[1389] Specific examples are shown below.

[1390] If user A suddenly sneezes during an online meeting, the following occurs:

[1391] 1. How to obtain audio data

[1392] Device: The microphone captures user A's voice in real time and stores it in a buffer.

[1393] 2. Preprocessing using digital signal processing

[1394] Terminal: Apply noise reduction to the acquired audio data, divide it into frames, and perform a Fourier transform.

[1395] 3. Methods for extracting speech features

[1396] Terminal: Calculate Mel-Frequency Cepstrum Coefficients (MFCC) from the audio data and generate a feature vector.

[1397] 4. Method for detecting abnormal voice using generative AI model

[1398] On the device: The feature vector is fed into a generative AI model to detect that the sneeze occurred between 1.5 and 1.8 seconds.

[1399] 5. Means for generating out-of-phase audio data

[1400] Device: Generates audio data in antiphase to a sneeze sound between 1.5 and 1.8 seconds.

[1401] 6. Means of mixing audio data

[1402] Terminal: Mix the original audio data with out-of-phase audio data to cancel the sneeze sound.

[1403] 7. Means of outputting audio data

[1404] Device: Sends the processed audio data to the online meeting application, where other participants hear the audio without hearing the sneeze.

[1405] This system allows users to participate in online meetings with peace of mind and prevents unintended audio from being heard by others.

[1406] Prompt Sentence Examples

[1407] "Please explain the system that automatically mutes you when you sneeze during an online meeting."

[1408] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1409] Step 1:

[1410] Acquiring audio data

[1411] Device: Uses a microphone to capture the user's voice data in real time.

[1412] Input: User speech and surrounding audio.

[1413] Output: Digital audio data from the microphone.

[1414] What it does: The microphone is always on, capturing audio data in real time and storing it in a fixed buffer.

[1415] Step 2:

[1416] Audio signal preprocessing

[1417] Terminal: Performs digital signal processing on the captured audio data, applying noise reduction algorithms, dividing the audio data into frames, and performing a Fourier transform to extract frequency components.

[1418] Input: Raw audio data.

[1419] Output: Denoised per-frame frequency spectrum.

[1420] What it does: It filters the audio signal to remove noise, then divides it into small time segments (frames), then performs a Fourier transform on each frame to calculate its frequency spectrum.

[1421] Step 3:

[1422] Feature extraction

[1423] Terminal: Calculate the Mel-Frequency Cepstral Coefficients (MFCC) from each frame and generate a feature vector.

[1424] Input: Frequency spectrum obtained by Fourier transform.

[1425] Output: MFCC-based feature vectors representing audio features.

[1426] Specific operation: Calculate MFCC using frequency spectrum as input, and save the calculation result as a feature vector in list format.

[1427] Step 4:

[1428] Abnormal audio detection

[1429] On the device: The feature vector is fed into a generative AI model to distinguish between normal and abnormal speech. If abnormal speech is detected, the start and end times of the speech are recorded.

[1430] Input: MFCC-based feature vector.

[1431] Output: Whether or not there is an abnormal sound, and its start and end times.

[1432] Specific operation: The generative AI model detects abnormal sounds based on the input feature vector and saves the period during which the abnormal sounds occurred as a timestamp.

[1433] Step 5:

[1434] Antiphase audio data generation

[1435] Terminal: Generates audio data in the opposite phase to the time range of the detected abnormal audio.

[1436] Input: Start and end times of the abnormal audio, and the original audio data.

[1437] Output: Out-of-phase audio data.

[1438] Specific operation: Based on the time range of the detected abnormal sound, calculate and save the opposite phase audio waveform.

[1439] Step 6:

[1440] Mixing of audio data

[1441] Terminal: Mixes the original audio data with out-of-phase audio data to cancel abnormal sounds.

[1442] Input: Original audio data, inverted audio data.

[1443] Output: Audio data with abnormal sounds canceled.

[1444] Specific operation: The original audio data is added to the out-of-phase audio data, and a process is performed to cancel out abnormal sounds.

[1445] Step 7:

[1446] Sending audio data

[1447] Terminal: Sends the processed audio data to the online meeting application.

[1448] Input: Mixed audio data.

[1449] Output: Final audio data (audio heard by other participants in the online meeting).

[1450] What it does: Stream the processed audio data in real time to an online meeting application.

[1451] (Application example 1)

[1452] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1453] In conventional online meetings and autonomous vehicle environments, sudden physiological events such as sneezing or coughing can cause unpleasant sounds to those around you and temporarily worsen the air quality, resulting in a decrease in voice quality, hindering conversations and the transmission of instructions, as well as reducing the comfort of the vehicle's interior.

[1454] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1455] In this invention, the server includes means for acquiring an audio signal, artificial intelligence means for analyzing the acquired audio signal and distinguishing between normal and abnormal audio, means for generating an opposite-phase audio signal when an abnormal audio signal is detected, means for mixing the original audio signal with the opposite-phase audio signal to cancel the unpleasant audio, means for outputting the audio signal, and means for temporarily strengthening related systems when canceling a new physiological phenomenon. This not only enables immediate detection and control of abnormal audio, but also temporarily improves the air environment, making it possible to maintain a comfortable and healthy environment.

[1456] plaintext

[1457] The "means for acquiring an audio signal" refers to a microphone and associated hardware for converting sound into an electronic signal.

[1458] The "generative artificial intelligence means" is an algorithm using artificial intelligence to analyze the acquired voice signal and distinguish between normal voice and abnormal voice.

[1459] The "means for generating an inverted phase audio signal" refers to a device or algorithm for generating an audio signal with the phase of an abnormal audio signal inverted when the abnormal audio signal is detected.

[1460] "Means for canceling unpleasant sounds" is a technology that synthesizes an audio signal that is out of phase with the original audio signal to cancel out unpleasant sounds.

[1461] The "means for outputting an audio signal" refers to a device such as a speaker or earphone for reproducing the processed audio signal.

[1462] The "means for temporarily enhancing related systems" is a control device for temporarily enhancing the air filter system and other related systems in the vehicle when a physiological phenomenon is detected.

[1463] plaintext

[1464] This invention is a system primarily designed to improve the riding experience in self-driving vehicles. It detects sudden physiological phenomena such as sneezing and coughing, cancels the unpleasant sounds that occur during these events, and temporarily improves the in-car environment.

[1465] The elements of the server, terminal, and user are as follows:

[1466] server

[1467] The vehicle is equipped with a microphone that captures audio signals and a computer system that includes an artificial intelligence algorithm for analyzing the audio. Audio signals are captured in real time and analyzed using digital signal processing technology. If an abnormal sound is detected, an opposite-phase audio signal is generated, thereby canceling out the unpleasant sound. This detection also temporarily strengthens the cabin air filter system.

[1468] Terminal

[1469] The terminal is a component of the in-vehicle system that acquires and analyzes audio signals in real time. Specifically, it performs noise reduction on the audio signal, frequency analysis using Fourier transform, and calculation of Mel Frequency Cepstrum Coefficients (MFCC). It also uses generative artificial intelligence to distinguish between normal and abnormal audio, and generates and mixes an inverse audio signal for detected abnormal audio to cancel it. It also includes a device that outputs the processed audio signal.

[1470] User

[1471] As passengers in self-driving vehicles, users can benefit from the ability to avoid physiological events such as sneezing or coughing while in an online meeting or using the navigation system, causing discomfort to other passengers and the system.

[1472] Specific examples

[1473] For example, the following prompt sentences can be used to input to a generative AI model to recreate a situation in which the present invention is implemented:

[1474] "Detect the sound of a sneeze in the car, mute the audio system at that moment, and strengthen the air filtration system."

[1475] This allows the in-car environment to be controlled quickly and automatically, providing a comfortable riding experience.

[1476] Hardware and Software

[1477] The system utilizes the following hardware and software:

[1478] Hardware: microphones, on-board computers, speakers, air filtration systems.

[1479] Software: Noise reduction algorithm, Fourier transform algorithm, Mel-Frequency Cepstral Coefficient (MFCC) extraction algorithm, generative AI model (TensorFlow, PyTorch, etc.).

[1480] By using the above prompt sentences, the generative AI model can immediately detect abnormal sounds and improve the air quality inside the car while canceling out unpleasant sounds.

[1481] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1482] plaintext

[1483] Step 1:

[1484] The server captures audio signals in real time from microphones inside the vehicle. The captured analog audio signals are converted to digital signals and passed to the next processing step. This input includes various sounds, such as ambient sounds inside the vehicle and conversation sounds.

[1485] Step 2:

[1486] The terminal performs noise reduction filtering on the captured voice signal, specifically using digital signal processing techniques to remove unwanted background noise, resulting in a clean voice signal being output.

[1487] Step 3:

[1488] The terminal divides the noise-reduced speech signal into frames and performs a Fourier transform to extract frequency components. The input is the speech signal after noise reduction, and the output is the frequency spectrum.

[1489] Step 4:

[1490] The device extracts speech features by calculating Mel-Frequency Cepstrum Coefficients (MFCC) from the frequency spectrum. This method expresses speech characteristics as numerical data. The input is the frequency spectrum, and the output is an MFCC feature vector.

[1491] Step 5:

[1492] The device inputs the MFCC feature vectors into a generative AI model (e.g., a model using TensorFlow or PyTorch) to distinguish between normal and abnormal speech. If abnormal speech (such as sneezing or coughing) is detected in this step, the device outputs the start and end times of the speech.

[1493] Step 6:

[1494] The terminal generates an audio signal with an opposite phase to the detected abnormal audio. To generate the opposite phase audio signal, an operation is performed to invert the phase of the original audio signal. The input is the audio signal of the abnormal audio, and the output is the audio signal with the opposite phase.

[1495] Step 7:

[1496] The terminal mixes the original audio signal with the generated out-of-phase audio signal, thereby performing a calculation to cancel the abnormal audio. The input is the original audio signal and the out-of-phase audio signal, and the output is the cancelled audio signal.

[1497] Step 8:

[1498] The device outputs the processed audio signal through the vehicle's speakers, allowing other passengers to hear the audio with the unpleasant sounds canceled out.

[1499] Step 9:

[1500] When an abnormal sound is detected, the server triggers a temporary strengthening of the air filter system inside the vehicle. Specifically, it increases the airflow of the air filter to quickly remove allergens and viruses in the air. This operation improves the air environment inside the vehicle.

[1501] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1502] This invention relates to a system that detects physiological phenomena such as sneezing, coughing, and hiccups that occur during online meetings and automatically mutes them. This system acquires audio signals in real time and uses generative AI to distinguish between normal and abnormal sounds. If abnormal sounds are detected, it generates and mixes an opposite-phase audio signal to cancel out the unpleasant sounds. Furthermore, by combining it with an emotion engine that recognizes the user's emotions, it has the ability to identify the user's emotional state and provide appropriate feedback in addition to processing the audio signals.

[1503] System configuration

[1504] 1. Audio acquisition method

[1505] Device: Equipped with a microphone for capturing audio signals in real time. Audio signals are continuously captured from this microphone.

[1506] 2. Voice Analysis Methods

[1507] Terminal: The captured audio data is analyzed using digital signal processing techniques, which apply a noise reduction algorithm, divide it into frames, and perform a Fourier transform to extract the frequency components of the audio.

[1508] 3. Feature Extraction Method

[1509] Terminal: To extract features from the preprocessed speech data, Mel-Frequency Cepstral Coefficients (MFCCs) are calculated, which generates a feature vector that quantifies the speech features.

[1510] 4. Classification of normal and abnormal voices

[1511] Terminal: The generated feature vector is input into the generative artificial intelligence (generative AI) to classify normal and abnormal voices. The generative AI identifies patterns in the voice signal based on past training data and detects abnormal voices such as sneezing and coughing.

[1512] 5. Abnormal Audio Detection

[1513] Terminal: When an abnormal sound is detected, its start and end time are recorded. For example, if a certain part of the sound signal is recognized as a sneeze, that time frame is marked.

[1514] 6. Generation of anti-phase audio signals

[1515] Terminal: Generates an audio signal with an inverse phase to the detected abnormal audio portion. The inverse phase audio signal is an audio signal with a phase difference of 180 degrees to cancel the original abnormal audio signal.

[1516] 7. Mixing of audio signals

[1517] Terminal: Mix the original audio signal with the generated out-of-phase audio signal. This mixing cancels out the unpleasant abnormal sounds.

[1518] 8. Audio signal output

[1519] Terminal: Sends the processed audio data to the audio stream of the online meeting, so that other participants hear the audio with the abnormal audio canceled.

[1520] 9. Emotion Engine

[1521] Device: The emotion engine has algorithms to identify the emotional state from the user's voice signal. The emotion engine analyzes voice characteristics such as tone, pitch, and rate to identify the emotion the user is feeling.

[1522] Server: Provides different feedback depending on the emotional state identified by the emotion engine. For example, if the server determines that the user is feeling stressed, it displays a message on the screen suggesting that the user relax.

[1523] Specific examples

[1524] Situation:

[1525] User B suddenly sneezes during an online meeting and feels stressed.

[1526] 1. Audio capture

[1527] Device: The microphone captures user B's voice in real time and stores it in a buffer.

[1528] 2. Voice Analysis

[1529] Terminal: Noise reduction is applied to the acquired audio data, it is divided into frames, and a Fourier transform is performed.

[1530] 3. Feature Extraction

[1531] Terminal: Calculate Mel-frequency cepstrum coefficients from the audio data and generate a feature vector.

[1532] 4. Classification of normal and abnormal voices

[1533] On the device: The feature vector is fed into a generative AI to detect that a sneeze occurred within a specific time frame.

[1534] 5. Abnormal Audio Detection

[1535] Device: Detects that the sneeze occurred between 1.5 and 1.8 seconds and marks that time frame.

[1536] 6. Generation of anti-phase audio signals

[1537] Device: Generates an audio signal in antiphase to a sneeze sound lasting 1.5 to 1.8 seconds.

[1538] 7. Mixing of audio signals

[1539] Terminal: Mixes the original audio signal with an out-of-phase audio signal to cancel the sneeze sound.

[1540] 8. Audio signal output

[1541] Terminal: Sends the processed audio signal to the online meeting application. Other participants receive the audio without hearing the sneeze.

[1542] 9. Emotion recognition

[1543] Device: The emotion engine identifies User B's emotional state from audio signals including the sneeze sound.

[1544] 10. Emotional Feedback

[1545] Server: If the emotion engine identifies User B's emotional state as "stressed," a message such as "Relax" will be displayed on the screen of the online meeting application.

[1546] Using this system, users can not only automatically cancel unpleasant sounds such as sneezing and coughing during online meetings in real time, but also receive support based on their emotional state using an emotion engine, allowing users to participate in online meetings with greater peace of mind.

[1547] The processing flow will be explained below.

[1548] Step 1: Getting voice input

[1549] On the device: The microphone is activated and the user's voice is captured in real time. The captured voice data is stored in a buffer.

[1550] Step 2: Preprocessing the audio data

[1551] Terminal: A noise reduction algorithm is applied to the acquired audio data to obtain a clear audio signal. The audio data is then divided into short time frames, and a Fourier transform is performed on each frame to extract the frequency components of the audio.

[1552] Step 3: Extract audio features

[1553] Terminal: To extract features from the preprocessed speech data, Mel-Frequency Cepstral Coefficients (MFCCs) are calculated, which generates a feature vector that quantifies the speech features.

[1554] Step 4: Classify normal and abnormal voices

[1555] Terminal: The generated feature vector is input into the generative artificial intelligence (generative AI) to classify normal and abnormal voices. The generative AI identifies patterns in the voice signal based on past training data and detects abnormal voices such as sneezing and coughing.

[1556] Step 5: Detecting Audio Anomalies

[1557] Terminal: When abnormal sounds are detected by the AI ​​generator, their start and end times are recorded. For example, if the audio signal is recognized as a sneeze from 1.5 seconds to 1.8 seconds, that time frame is marked.

[1558] Step 6: Generate an anti-phase audio signal

[1559] Terminal: Generates an audio signal with an inverse phase to the detected abnormal audio portion. The inverse phase audio signal is an audio signal with a phase difference of 180 degrees to cancel the original abnormal audio signal.

[1560] Step 7: Mixing the audio signals

[1561] Terminal: Mix the original audio signal with the generated out-of-phase audio signal. This mixing cancels out the unpleasant abnormal sounds.

[1562] Step 8: Output the audio signal

[1563] Terminal: Sends the processed audio data to the audio stream of the online meeting, so that other participants hear the audio with the abnormal audio canceled.

[1564] Step 9: Emotion Recognition Processing

[1565] On the device: The voice signal is processed by an emotion engine that analyzes characteristics such as tone, pitch, and rate to identify the user's emotional state.

[1566] On the device: The emotion engine identifies the user's emotion based on the analyzed features. For example, if the voice is spoken with a strong tone, high pitch, and rapid speed, it will be identified as stressed or angry.

[1567] Step 10: Emotional State Feedback

[1568] Server: Provides appropriate feedback based on the emotional state identified by the emotion engine. For example, if the server determines that the user is feeling stressed, it displays a message such as "Relax" on the screen of an online meeting application.

[1569] Server: Based on the user's emotional state, it provides a feedback function that notifies other participants of the user's state, depending on the settings.

[1570] In this way, the entire system performs a series of processes in real time, from acquiring voice signals to canceling unpleasant sounds, and then identifying the user's emotions and providing appropriate feedback. This allows users to continue participating in online meetings with peace of mind, even when experiencing physiological or emotional changes.

[1571] Example 2

[1572] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1573] Physiological phenomena such as sneezing, coughing, and hiccups that occur during online meetings can be unpleasant for other participants and disrupt the flow of conversation. Furthermore, users often feel stressed if they become too aware of these physiological phenomena. Furthermore, it is difficult to grasp users' emotional state in real time, and there is a lack of means to provide appropriate feedback. This can reduce the quality of online meetings and impair participants' concentration and comfort.

[1574] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1575] In this invention, the terminal includes means for acquiring an audio signal, artificial intelligence means for analyzing the acquired audio signal and distinguishing between normal audio and abnormal audio, means for generating an opposite-phase audio signal when abnormal audio is detected, means for mixing the original audio signal with the opposite-phase audio signal and canceling unpleasant audio, means for outputting an audio signal, and emotion recognition means for recognizing the user's emotional state and providing feedback. This makes it possible to automatically cancel unpleasant physiological sounds during online meetings and provide support according to the user's emotional state.

[1576] The "audio signal" is a user's speech converted into an electrical signal, and is acquired through an input device such as a microphone.

[1577] "Means for acquiring" refers to a device or method for acquiring audio signals in real time using a microphone, sensor device, etc.

[1578] "Analyzing" is the process of analyzing an audio signal using digital signal processing techniques to extract various features.

[1579] "Normal voice" refers to voice that is not an abnormal voice, such as a normal conversation or speech.

[1580] "Abnormal sounds" refers to physiological sounds that interfere with normal conversation, such as sneezing, coughing, and hiccups.

[1581] "Artificial intelligence tools" are algorithms or models that use historical training data to identify speech patterns and categorize or differentiate between speech sounds.

[1582] "Detecting abnormal sounds" means identifying abnormal sounds from the acquired audio signals and recognizing the time of occurrence and their content.

[1583] An "opposite phase audio signal" is an audio signal generated by reversing the phase of the original abnormal audio signal by 180 degrees, and is used to cancel the abnormal audio.

[1584] "Mixing" is a process of synthesizing an audio signal that is out of phase with the original audio signal to cancel out abnormal audio.

[1585] "Unpleasant sounds" are sounds generated by physiological phenomena such as sneezing, coughing, and hiccuping, which are unpleasant to other participants and disrupt conversation.

[1586] "Cancellation means" refers to a method for filtering out abnormal sounds through a mixed audio signal, allowing other sounds to be heard normally.

[1587] "Means for outputting audio signals" refers to devices or methods for converting the processed audio data back into an audio stream for transmission to other participants.

[1588] "Emotion recognition means" refers to an algorithm or system that grasps the user's emotional state from their voice signal and analyzes it in real time.

[1589] A "means for providing feedback" is a system or method for displaying appropriate advice or suggestions to a user depending on the perceived emotional state.

[1590] This invention is a system that detects unpleasant physiological sounds (sneezing, coughing, hiccups, etc.) that occur during online meetings in real time and automatically cancels them. This system also aims to improve the quality of online meetings by recognizing the user's emotional state in real time and providing appropriate feedback.

[1591] System Configuration

[1592] Audio acquisition means

[1593] The device has a built-in microphone that captures the user's voice signal in real time, which is then stored in a buffer and used in subsequent processing steps.

[1594] Voice analysis methods

[1595] The captured audio signal is pre-processed using digital signal processing techniques, which involves applying a noise reduction algorithm to remove background noise, dividing the signal into frames, and then performing a Fourier transform to extract the audio frequency components.

[1596] Feature extraction method

[1597] Mel-frequency cepstral coefficients (MFCCs) are calculated from the preprocessed speech data, and a feature vector is generated that quantifies the speech features, enabling detailed analysis of the speech signal.

[1598] Classification of normal and abnormal voices

[1599] The device inputs the generated feature vector into a generative AI model to distinguish between normal and abnormal voices. The generative AI model uses past training data to identify patterns in the voice signal and detect abnormal sounds (e.g., sneezing, coughing).

[1600] Abnormal voice detection

[1601] If the generative AI model detects an abnormal sound, it records the time it occurred. For example, if a sneeze occurs between 1.5 and 1.8 seconds, that time frame will be recorded.

[1602] Generation of out-of-phase speech

[1603] An audio signal with an opposite phase is generated for the detected abnormal sound portion. The opposite phase audio signal is an audio signal with a phase difference of 180 degrees to cancel the original abnormal sound signal.

[1604] Audio signal mixing

[1605] By mixing the original audio signal with the generated out-of-phase audio signal, the unpleasant abnormal sounds are cancelled out.

[1606] Audio signal output

[1607] The processed audio data is sent to the audio stream of the online meeting application, so that other participants can hear the audio with the abnormal sounds canceled out.

[1608] emotion recognition means

[1609] The device is equipped with an emotion recognition engine that recognizes the user's emotional state from their voice signal. The engine analyzes the tone, pitch, and speed of the voice to determine the user's current emotional state.

[1610] Feedback methods

[1611] The server then provides appropriate feedback based on the emotional state identified by the emotion recognition engine. For example, if the server determines that the user is feeling stressed, a message such as "Relax" will be displayed on the screen of an online meeting application.

[1612] Specific situations and prompt sentence examples

[1613] Situation

[1614] User B sneezes during an online meeting and feels stressed.

[1615] 1. Audio capture

[1616] The device uses a microphone to capture user B's voice in real time and stores it in a buffer.

[1617] 2. Audio analysis

[1618] Noise reduction is applied to the voice data acquired by the device, and frequency components are extracted.

[1619] 3. Feature Extraction

[1620] The terminal calculates the Mel-frequency cepstral coefficients and generates a feature vector.

[1621] 4. Audio Classification

[1622] The device inputs the feature vector into a generative AI model to identify the time frame of the sneeze sound.

[1623] 5. Abnormal Audio Detection

[1624] The device records the time frame in which the sneeze occurred.

[1625] 6. Generation of out-of-phase speech

[1626] The terminal generates an audio signal that is in the opposite phase to the sneeze sound.

[1627] 7. Mixing of audio signals

[1628] The device mixes the original audio with an out-of-phase audio to cancel out the sneeze sound.

[1629] 8. Audio signal output

[1630] The terminal sends the processed audio signal to the online meeting.

[1631] 9. Emotion recognition

[1632] The device uses an emotion engine to recognize the emotional state of user B.

[1633] 10. Emotional Feedback

[1634] The server displays the message "Relax" based on User B's emotional state.

[1635] Prompt Sentence Examples

[1636] "When a user sneezes during an online meeting, create a program that analyzes the audio signal in real time, generates and mixes an opposite-phase audio signal, and cancels out the unpleasant sound. Also, use an emotion engine to identify the user's emotional state and provide feedback."

[1637] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1638] Step 1: Acquire audio

[1639] The device uses a microphone to capture the user's voice in real time. The input is the user's speech, and the output is the audio signal stored in the buffer. Specifically, the microphone is always on, and the user's voice is continuously sampled and stored in the buffer as digital data.

[1640] Step 2: Preprocessing the audio signal

[1641] The terminal performs digital signal processing (DSP) on the captured audio signal. The input is the raw audio signal stored in the buffer, and the output is the pre-processed audio signal. Specifically, it first applies a noise reduction algorithm to remove background noise. Then it divides the audio signal into frames and applies a Fourier transform to each frame to extract the frequency spectrum.

[1642] Step 3: Feature extraction

[1643] The device calculates Mel-Frequency Cepstral Coefficients (MFCCs) from the preprocessed audio signal. The input is the preprocessed audio signal, and the output is a feature vector. Specifically, the device calculates MFCCs for each frame and combines these MFCC values ​​to generate a single feature vector.

[1644] Step 4: Classify the audio

[1645] The device inputs the generated feature vector into a generative AI model to classify normal and abnormal speech. The input is the feature vector, and the output is the speech classification result. Specifically, the feature vector is input into the AI ​​model, and the AI ​​executes a process to determine whether the speech is normal or abnormal.

[1646] Step 5: Detecting Audio Anomalies

[1647] The device records the time when the abnormal sound occurred based on the classification result. The input is the sound classification result, and the output is the time information of the abnormal sound occurrence. Specifically, it identifies the time frame in which the AI ​​judged the sound to be abnormal and saves that information as a timestamp.

[1648] Step 6: Generate out-of-phase audio

[1649] The terminal generates an audio signal with an inverse phase to the detected abnormal audio portion. The input is the time frame of the abnormal audio and the original audio signal, and the output is an audio signal with an inverse phase. Specifically, it generates an inverse phase signal by inverting the frequency components of the abnormal audio and places that signal in the specified time frame.

[1650] Step 7: Mixing the audio signals

[1651] The terminal mixes the original audio signal with the generated out-of-phase audio signal. The input is the original audio signal and the out-of-phase audio signal, and the output is the mixed audio signal. Specifically, the original audio signal and the out-of-phase audio signal are added together to generate an audio signal that cancels the abnormal audio.

[1652] Step 8: Output the audio signal

[1653] The terminal transmits the processed audio data to the audio stream of the online meeting application. The input is the mixed audio signal, and the output is the processed audio that is transmitted to the participants of the online meeting. In specific operations, the terminal sends the processed audio signal to the transmission channel of the meeting application, allowing other participants to receive the audio with the abnormal sound canceled.

[1654] Step 9: Emotion Recognition

[1655] The device uses an emotion recognition engine to identify the user's emotional state from the voice signal. The input is the real-time voice signal, and the output is the user's emotional state. Specifically, the device analyzes the tone, pitch, and speed of the voice signal to determine emotions such as stress, joy, and sadness.

[1656] Step 10: Emotional Feedback

[1657] The server provides appropriate feedback according to the emotional state identified by the emotion recognition engine. The input is the user's emotional state, and the output is a feedback message. Specifically, the server generates a message that matches the user's emotional state and displays appropriate feedback, such as "Relax," on the screen of the online meeting application.

[1658] (Application example 2)

[1659] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1660] Robots in factories are required to accurately receive voice instructions from operators without misrecognition. However, if abnormal sounds such as sneezing or coughing are included in the voice, the risk of the robot malfunctioning increases. Furthermore, if the operator is stressed, it will affect work efficiency and safety. To solve these problems, a system is needed that combines the detection and removal of abnormal sounds and the recognition of the operator's emotional state.

[1661] The identification processing by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for acquiring a voice signal, artificial intelligence means for analyzing the acquired voice signal and distinguishing between normal voice and abnormal voice, means for generating an opposite-phase voice signal when an abnormal voice is detected, means for mixing the original voice signal with the opposite-phase voice signal to cancel out unpleasant voice, means for outputting a voice signal, emotion engine means for extracting features from the voice signal and identifying the emotional state, and means for providing feedback according to the emotional state. This allows the robot to accurately receive the operator's voice instructions, is not affected by abnormal sounds (such as sneezing or coughing), and can provide appropriate feedback according to the operator's emotional state.

[1662] The "means for acquiring an audio signal" refers to a means for obtaining audio data by a microphone or other audio input device.

[1663] "Generative artificial intelligence means for analyzing acquired voice signals and distinguishing between normal voices and abnormal voices" refers to a means for analyzing acquired voice data and using a generative AI to distinguish between normal voices and abnormal voices such as sneezing and coughing.

[1664] The "means for generating an audio signal of an opposite phase when an abnormal sound is detected" refers to a means for generating an audio signal of a phase shift of 180 degrees when an abnormal sound is detected in order to cancel out the sound.

[1665] The "means for mixing the original audio signal with an out-of-phase audio signal to cancel unpleasant sounds" is a means for canceling abnormal sounds, specifically unpleasant sounds, by combining the original audio with the generated out-of-phase audio.

[1666] "Means for outputting audio signals" refers to means for transmitting processed audio data to other systems or devices.

[1667] "Emotion engine means for extracting features from voice signals and identifying emotional states" refers to a means for extracting emotion-related features such as tone and pitch from voice data and analyzing and identifying the emotional state of the operator.

[1668] The "means for providing feedback according to emotional state" is a means for providing appropriate feedback or a message based on the emotional state of the operator identified by the emotion engine.

[1669] This invention is a system that enables robots in factories to accurately receive voice instructions from operators without misrecognition. In particular, it eliminates the effects of abnormal sounds such as sneezing and coughing, and provides feedback according to the operator's emotional state.

[1670] The system includes many means, but the main parts are as follows:

[1671] 1. Acquisition of audio signals:

[1672] The server uses a microphone or other audio input device to capture the operator's voice signals in real time, which consist of instructions and other environmental sounds in the factory.

[1673] 2. Means of analyzing the audio signal:

[1674] The server analyzes the captured audio signal using digital signal processing techniques. Specifically, it applies a noise reduction algorithm, divides the audio signal into frames, performs a Fourier transform, and extracts the audio frequency components. This process is performed using the feature.mfcc function from librosa.

[1675] 3. Generative AI means:

[1676] The server inputs the features extracted from the voice signal into a generative AI model, which distinguishes between normal voices and abnormal voices such as sneezing and coughing. This generative AI model uses a pre-trained convolutional neural network (CNN).

[1677] 4. Means for generating anti-phase audio signals:

[1678] When an abnormal sound is detected, the server generates an audio signal that is 180 degrees out of phase with the detected sound to cancel it out, effectively canceling the abnormal sound.

[1679] 5. Audio signal mixing means:

[1680] The server cancels out abnormal sounds by mixing the original audio signal with the generated out-of-phase audio, so that only normal instruction audio reaches the robot.

[1681] 6. Audio output means:

[1682] The server transmits the processed voice data to the robot so that the robot receives the voice instructions accurately.

[1683] 7. Emotion Engine Means:

[1684] The server analyzes the operator's emotional state using features extracted from the voice signal, using algorithms that use voice features such as tone and pitch.

[1685] 8. Feedback methods:

[1686] The server provides appropriate feedback based on the operator's emotional state as identified by the emotion engine. For example, if stress is detected, the robot will display a message encouraging the operator to take a break.

[1687] Specific examples

[1688] In a real-world scenario, an operator sneezes multiple times while working and then becomes extremely stressed. Detecting this state, the robot pauses the work and displays a message encouraging the operator to take a break. The system's processing steps are seamlessly handled by the server, and appropriate actions are automatically taken.

[1689] Prompt Sentence Examples

[1690] Build a system where if an operator sneezes multiple times and then becomes very stressed, the robot will display a message encouraging them to take a break.

[1691] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1692] Step 1:

[1693] The server uses a microphone to capture the operator's voice signal in real time. The voice signal obtained from this microphone includes environmental sounds in the factory and the operator's voice instructions. The input is an analog voice signal, and the output is digitized voice data. The server uses a buffer memory to store this data.

[1694] Step 2:

[1695] The server applies digital signal processing technology to the acquired audio signal, performs noise reduction, divides the audio signal into frames, performs a Fourier transform, and extracts the frequency components of the audio. The input is the digitized audio data obtained in step 1, and the output is frame data converted into frequency components.

[1696] Step 3:

[1697] The server calculates Mel-Frequency Cepstrum Coefficients (MFCC) from the frame data and extracts speech features. The input is the frequency component data obtained in step 2, and the output is a feature vector obtained using MFCC.

[1698] Step 4:

[1699] The server inputs the extracted feature vector into a generative AI model to classify normal and abnormal speech. The generative AI model identifies speech signal patterns based on past training data. The input is the feature vector obtained in step 3, and the output is a label: "normal speech" or "abnormal speech."

[1700] Step 5:

[1701] When an abnormal sound (sneezing or coughing) is detected, the server records its start and end times. The input is the classification label obtained in step 4, and the output is the time frame information of the abnormal sound.

[1702] Step 6:

[1703] The server generates an audio signal with an inverse phase to the detected abnormal audio. This inverse phase audio signal is an audio signal with a phase difference of 180 degrees to cancel out the original abnormal audio. The input is the time frame information of the abnormal audio obtained in step 5, and the output is the inverse phase audio signal.

[1704] Step 7:

[1705] The server cancels the unpleasant abnormal sounds by mixing the original sound signal with an out-of-phase sound signal. The input is the original sound signal obtained in step 1 and the out-of-phase sound signal generated in step 6, and the output is the cancelled sound signal.

[1706] Step 8:

[1707] The server sends the processed audio signal to the robot, which allows the robot to receive an audio signal that does not contain the canceled abnormal audio. The input is the canceled audio signal obtained in step 7, and the output is the digital audio signal received by the robot.

[1708] Step 9:

[1709] The server re-extracts features from the cancelled speech signal and inputs them into an emotion engine for identifying emotional states. The input is the cancelled speech signal obtained in step 7, and the output is a label indicating the emotional state.

[1710] Step 10:

[1711] The server provides appropriate feedback based on the emotional state identified by the emotion engine. For example, if many abnormal sounds are detected and the emotional state is identified as "stressed," the robot displays a message encouraging the operator to take a break. The input is the label indicating the emotional state obtained in step 9, and the output is the feedback message.

[1712] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1713] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1714] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1715] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1716] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1717] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1718] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1719] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1720] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1721] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1722] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1723] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1724] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1725] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1726] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1727] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1728] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1729] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1730] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1731] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1732] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1733] The following is further disclosed regarding the above embodiment.

[1734] (Claim 1)

[1735] means for acquiring an audio signal;

[1736] a generating artificial intelligence means for analyzing the acquired voice signal and distinguishing between normal voice and abnormal voice;

[1737] means for generating an audio signal of an opposite phase when an abnormal audio signal is detected;

[1738] a means for mixing an original audio signal with an audio signal of opposite phase to cancel out unpleasant audio;

[1739] means for outputting an audio signal;

[1740] A system including:

[1741] (Claim 2)

[1742] 10. The system of claim 1, further comprising means for processing said audio signal in real time.

[1743] (Claim 3)

[1744] 2. The system according to claim 1, wherein the artificial intelligence generating means includes means for extracting speech features using Mel-frequency cepstral coefficients.

[1745] "Example 1"

[1746] (Claim 1)

[1747] means for acquiring audio data;

[1748] means for preprocessing the acquired audio data using digital signal processing technology;

[1749] means for extracting speech features from the preprocessed speech data;

[1750] A means for inputting the extracted voice features into a generative artificial intelligence model to distinguish between normal voice and abnormal voice;

[1751] means for generating audio data of an opposite phase when an abnormal audio is detected;

[1752] A means for canceling unpleasant sounds by mixing original sound data with opposite-phase sound data;

[1753] means for outputting the processed audio data;

[1754] A system including:

[1755] (Claim 2)

[1756] 10. The system of claim 1, wherein the system processes the audio data in real time.

[1757] (Claim 3)

[1758] 2. The system of claim 1, wherein the generative artificial intelligence model extracts speech features using Mel-frequency cepstral coefficients.

[1759] "Application Example 1"

[1760] plaintext

[1761] (Claim 1)

[1762] means for acquiring an audio signal;

[1763] a generating artificial intelligence means for analyzing the acquired voice signal and distinguishing between normal voice and abnormal voice;

[1764] means for generating an audio signal of an opposite phase when an abnormal audio signal is detected;

[1765] a means for mixing an original audio signal with an audio signal of opposite phase to cancel out unpleasant audio;

[1766] means for outputting an audio signal;

[1767] A means for temporarily strengthening the relevant system when canceling a novel physiological phenomenon;

[1768] A system including:

[1769] (Claim 2)

[1770] 10. The system of claim 1, further comprising means for processing said audio signal in real time.

[1771] (Claim 3)

[1772] 10. The system of claim 1, further comprising means for extracting speech features using Mel-frequency cepstral coefficients.

[1773] "Example 2: Combining Emotion Engines"

[1774] (Claim 1)

[1775] means for acquiring an audio signal;

[1776] an artificial intelligence means for analyzing the acquired voice signal and distinguishing between normal voice and abnormal voice;

[1777] means for generating an audio signal of an opposite phase when an abnormal audio signal is detected;

[1778] a means for mixing an original audio signal with an audio signal of opposite phase to cancel out unpleasant audio;

[1779] means for outputting an audio signal;

[1780] emotion recognition means for recognizing the emotional state of a user and providing feedback;

[1781] A system including:

[1782] (Claim 2)

[1783] 10. The system of claim 1, further comprising means for processing said audio signal in real time.

[1784] (Claim 3)

[1785] 2. The system of claim 1, wherein the artificial intelligence means includes means for extracting speech features using Mel-frequency cepstral coefficients.

[1786] "Application example 2 when combining emotion engines"

[1787] (Claim 1)

[1788] means for acquiring an audio signal;

[1789] a generating artificial intelligence means for analyzing the acquired voice signal and distinguishing between normal voice and abnormal voice;

[1790] means for generating an audio signal of an opposite phase when an abnormal audio signal is detected;

[1791] a means for mixing an original audio signal with an audio signal of opposite phase to cancel out unpleasant audio;

[1792] means for outputting an audio signal;

[1793] an emotion engine means for extracting features from a voice signal and identifying an emotional state;

[1794] a means for providing feedback according to emotional state;

[1795] A system including:

[1796] (Claim 2)

[1797] 10. The system of claim 1, further comprising means for processing said audio signal in real time.

[1798] (Claim 3)

[1799] 2. The system according to claim 1, wherein the artificial intelligence generating means includes means for extracting speech features using Mel-frequency cepstral coefficients. [Explanation of symbols]

[1800] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. means for acquiring an audio signal; a generating artificial intelligence means for analyzing the acquired voice signal and distinguishing between normal voice and abnormal voice; means for generating an audio signal of an opposite phase when an abnormal audio signal is detected; a means for mixing an original audio signal with an audio signal of opposite phase to cancel out unpleasant audio; means for outputting an audio signal; A system including:

2. 2. The system of claim 1, further comprising means for processing said audio signal in real time.

3. 2. The system according to claim 1, wherein said generating artificial intelligence means includes means for extracting speech features using Mel-frequency cepstral coefficients.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A