System

A system that converts negative workplace comments into positive audio in real time improves productivity by maintaining information flow and enhancing psychological stability.

JP2026019145APending Publication Date: 2026-02-05SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024120554
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-25
Publication Date
2026-02-05

AI Technical Summary

Technical Problem

Negative comments and irritable speech in the workplace reduce team morale and productivity, while completely blocking audio reduces the amount of information available, disrupting work.

Method used

A system that captures ambient sound, converts it into text, analyzes and detects negative expressions, converts them into positive expressions using a generative AI model, and plays back the positive audio in real time.

Benefits of technology

Maintains the same amount of information available while improving psychological stability and productivity by replacing negative comments with positive expressions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026019145000001_ABST
    Figure 2026019145000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system comprising: means for acquiring ambient audio; means for converting the acquired audio into text data; means for analyzing and detecting negative expressions from the text data; means for converting the detected negative expressions into positive expressions; means for converting the converted positive expressions into audio data; and means for playing back the audio data.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] In today's workplace, the spread of negative comments and irritable speech is known to have a negative impact on the productivity of the entire team and the organization as a whole. In particular, angry comments from superiors often lower the morale of the entire team and cause a decline in work efficiency. Furthermore, completely blocking audio reduces the amount of information available, which can disrupt work. Therefore, there is a need for a method to effectively suppress negative comments from those around you and replace them with positive ones. [Means for solving the problem]

[0005] The present invention solves the above-mentioned problems by providing a system including a means for acquiring ambient sound, a means for converting the acquired sound into text data, a means for analyzing and detecting negative expressions from the text data, a means for converting the detected negative expressions into positive expressions, a means for converting the converted positive expressions into sound data, and a means for playing back the sound data.

[0006] The device picks up surrounding sounds using a microphone, and the server converts the sounds into text. The server uses a sentiment analysis library to detect negative expressions and converts them into positive ones using a generative artificial intelligence model. The converted text is then converted back into audio data and played back to the user from the device in real time. This allows the user to receive only positive audio information without being influenced by the negative emotions around them, while maintaining the same amount of information.

[0007] "Means for acquiring surrounding audio" refers to the function of collecting surrounding audio in real time using a microphone or other audio collection device and outputting it as digital data.

[0008] "Means for converting acquired voice data into text data" refers to a function that uses voice recognition technology to convert collected voice data into text information. For example, a voice recognition API would fall under this category.

[0009] "Means for analyzing and detecting negative expressions from text data" refers to a function that uses natural language processing and sentiment analysis to identify and distinguish negative emotions and unhappy expressions in text data.

[0010] "Means for converting detected negative expressions into positive expressions" refers to a function that uses a generative artificial intelligence model to convert detected negative text into positive text that preserves meaning.

[0011] "Means for converting the converted positive expressions into audio data" refers to a function for converting the positive text data back into audio data using a text-to-audio conversion tool.

[0012] "Means for playing back audio data" refers to a function for playing back the generated audio data to the user in real time using an audio output device such as a speaker or earphones. [Brief explanation of the drawings]

[0013] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12]FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0014] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0015] First, the terms used in the following description will be explained.

[0016] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0017] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0018] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0019] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0020] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0021] [First embodiment]

[0022] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0023] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0024] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0025] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0026] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0027] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0028] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0029] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0030] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0031] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0032] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0033] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0034] This invention is a system that converts negative comments from people around you into positive expressions and provides positive voice information to the user. Each element and its operation will be described in detail below.

[0035] 1. System Configuration

[0036] (a) Terminal

[0037] The device is equipped with a microphone to capture surrounding sounds, and hardware and software for processing and playing back the audio data. The device transmits the collected audio data to a server in real time and plays back the audio data from the server in real time.

[0038] (b) Server

[0039] The server has the following functions:

[0040] Speech recognition function: Converts voice data into text.

[0041] Sentiment analysis function: Analyzes text data and detects negative expressions.

[0042] Conversion function: Converts detected negative expressions into positive expressions.

[0043] Speech synthesis function: Converts positive text data back into speech data.

[0044] 2. Program Processing

[0045] (a) Audio capture and transmission

[0046] The device picks up ambient sounds through a microphone, and the acquired audio data is sent to the server in streaming format.

[0047] (b) Converting voice data into text

[0048] The server receives the voice data sent from the device, converts it into text using a speech recognition API, and then prepares the text data for analysis.

[0049] (c) Detection of negative expressions

[0050] The server analyzes the text data using a sentiment analysis library, and based on the analysis results, negative expressions and unpleasant speech patterns are identified.

[0051] (d) Transforming negative expressions into positive ones

[0052] The server converts the detected negative expressions into positive expressions using a generative artificial intelligence model.

[0053] (e) Vocalization and delivery of positive expressions

[0054] The server converts the positively converted text data into voice data using a voice synthesis tool, and the voice data is transmitted to the terminal.

[0055] (f) Audio playback

[0056] The terminal receives the positive voice data sent from the server and plays it back to the user in real time, so that the user can obtain positive voice information.

[0057] 3. Specific Examples

[0058] Example 1: Converting an angry boss's words

[0059] Boss' statement:

[0060] "Why can't you do something so simple? I can't believe it!"

[0061] The device picks up this speech with a microphone and sends it to the server.

[0062] The server uses a speech recognition API to generate the text "Why can't you do something so simple? I can't believe it!"

[0063] The server uses a sentiment analysis library to determine that this text is negative.

[0064] The server uses a generative artificial intelligence model to convert this into "You made a lot of mistakes today. Let's think together about how to solve the problem."

[0065] The server converts the converted text into voice data using a voice synthesis tool and sends it to the terminal.

[0066] The terminal receives the converted audio data and plays it back to the user.

[0067] Example 2: Converting a colleague's grumpy tone

[0068] Colleagues say:

[0069] "I'm tired. I'm starting to hate this job."

[0070] The device picks up this speech with a microphone and sends it to the server.

[0071] The server uses a speech recognition API to generate the text "I'm tired. I hate this job."

[0072] The server uses a sentiment analysis library to determine that this text is negative.

[0073] The server uses a generative artificial intelligence model to translate this into "I'm a little tired today, but I believe I can get through it."

[0074] The server converts the converted text into voice data using a voice synthesis tool and sends it to the terminal.

[0075] The terminal receives the converted audio data and plays it back to the user.

[0076] In this way, this system converts negative expressions into positive ones, thereby improving the user's productivity.

[0077] The processing flow will be explained below.

[0078] Step 1:

[0079] The device picks up ambient sound through a microphone, and the sound data is recorded in real time and sent to a server via streaming.

[0080] Step 2:

[0081] The server receives the voice data sent from the device. This voice data is converted into text using the Google speech recognition API or other voice recognition technology. The converted text data is temporarily stored in an internal buffer.

[0082] Step 3:

[0083] The server inputs the converted text data into a sentiment analysis library, which analyzes the text data for negative expressions and detects negative or irritable speech patterns. The results are saved as negative flags.

[0084] Step 4:

[0085] The server checks the analysis results from the sentiment analysis library and extracts negative phrases, which are then converted into positive ones using a generative AI model.

[0086] Step 5:

[0087] The server inputs the generated positive expressions into a text-to-speech (TTS) tool to generate voice data, which is then temporarily stored.

[0088] Step 6:

[0089] The server then transmits the generated audio data to the terminal in streaming format, delivering the audio to the user in real time.

[0090] Step 7:

[0091] The device decodes the voice data received from the server and plays it through a speaker or earphones, allowing the user to receive voice information converted into positive expressions in real time.

[0092] Example 1

[0093] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0094] In conventional systems, negative comments from those around you could be directly conveyed to the user, causing psychological stress and reducing productivity. Furthermore, the process of converting negative comments into positive ones was complicated, making it difficult to respond in real time. This made it difficult for users to work in a comfortable environment.

[0095] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0096] In this invention, the server includes means for converting received speech into text data, means for analyzing the text data to detect negative expressions, and means for converting the detected negative expressions into positive expressions using a generative artificial intelligence model. This makes it possible to acquire surrounding speech, convert the negative expressions into positive expressions in real time, and provide them to the user.

[0097] "Audio capture device" refers to a device for capturing ambient sound. Examples include a microphone.

[0098] A "speech recognition service API" refers to an application programming interface for converting voice data into text, often provided as a cloud-based service.

[0099] "Server" refers to a computer system with data processing capabilities that receives, converts, and analyzes voice data.

[0100] A "generative artificial intelligence model" refers to an algorithm that generates new text based on a given prompt, based on deep learning techniques, such as GPT-3.

[0101] "Real-time" refers to data processing and communication occurring instantaneously with minimal delay.

[0102] "Text data" refers to the result of converting audio data into text information, which serves as the basis for analysis and other processing.

[0103] "Negative expressions" refer to negative content or expressions identified through sentiment analysis, which may cause psychological stress to users.

[0104] "Positive expressions" refer to positive content or expressions that replace negative expressions and are used to maintain the user's psychological stability.

[0105] "Audio data" refers to the digital representation of sound that is used to be heard by a user through a playback device.

[0106] This invention is a system that converts negative comments from people around you into positive expressions and provides positive voice information to the user. Each element and its operation will be described in detail below.

[0107] 1. System Configuration

[0108] (a) Terminal

[0109] The device is equipped with a sound capture device (microphone) for capturing surrounding sounds, and hardware and software for processing and playing back the sound data. The device transmits the collected sound data to a server in real time and plays back the sound data from the server in real time.

[0110] (b) Server

[0111] The server has the following functions:

[0112] Speech recognition function: Converts voice data into text. Specifically, it uses a speech recognition service API (e.g., Google Cloud Speech-to-Text).

[0113] Sentiment analysis function: Analyzes text data and detects negative expressions. Specifically, it uses a sentiment analysis library (e.g., NLTK or Google Cloud Natural Language API).

[0114] Conversion function: Converts detected negative expressions into positive expressions using a generative artificial intelligence model (e.g., GPT-3).

[0115] Speech synthesis function: Positive text data is converted back into voice data, specifically using a speech synthesis tool (e.g., Google Cloud Text-to-Speech).

[0116] 2. Operation overview

[0117] The device captures ambient sounds and sends the audio data in streaming format to the server. The server then converts the received audio data into text data using a speech recognition API. The server then analyzes the text data using a sentiment analysis library to identify negative expressions. The identified negative expressions are converted into positive expressions using a generative artificial intelligence model. The converted positive text data is then reconverted into audio data using a speech synthesis tool and sent to the device. Finally, the device plays the transmitted positive audio data to the user.

[0118] 3. Specific Examples

[0119] Example 1: Converting an angry boss's words

[0120] Boss' statement:

[0121] "Why can't you do something so simple? I can't believe it!"

[0122] 1. The device picks up this speech with a microphone and sends it to the server.

[0123] 2. The server uses a speech recognition API to generate the text "Why can't you do something so simple? I can't believe it!"

[0124] 3. The server uses a sentiment analysis library to determine that the text is negative.

[0125] 4. The server uses a generative artificial intelligence model to convert this into "You made a lot of mistakes today. Let's think together about how to solve the problem."

[0126] 5. The server converts the converted text into voice data using a speech synthesis tool and sends it to the device.

[0127] 6. The device receives the converted audio data and plays it back to the user.

[0128] Example 2: Converting a colleague's grumpy tone

[0129] Colleagues say:

[0130] "I'm tired. I'm starting to hate this job."

[0131] 1. The device picks up this speech with a microphone and sends it to the server.

[0132] 2. The server uses a speech recognition API to generate the text "I'm tired. I hate this job."

[0133] 3. The server uses a sentiment analysis library to determine that the text is negative.

[0134] 4. The server uses a generative artificial intelligence model to convert this to "I'm a little tired today, but I believe I can get it done."

[0135] 5. The server converts the converted text into voice data using a speech synthesis tool and sends it to the device.

[0136] 6. The device receives the converted audio data and plays it back to the user.

[0137] Prompt Sentence Examples

[0138] Change the following negative text into a positive one: "I'm exhausted. I hate this job."

[0139] In this way, this system converts negative expressions from the surrounding environment into positive ones, improving the user's psychological stability and productivity.

[0140] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0141] Step 1:

[0142] The device captures the surrounding sounds using an audio capture device, specifically, a microphone collects audio data and transmits it to a server in real time in a streaming format.

[0143] Input: Ambient audio

[0144] Output: Streaming audio data

[0145] Step 2:

[0146] The server receives the voice data sent from the device. It sends the received voice data to a voice recognition service API and converts the voice data into text data. Specifically, it uses a voice recognition API (e.g., Google Cloud Speech-to-Text) to convert the voice data into text format.

[0147] Input: Streaming audio data

[0148] Output: Text data

[0149] Step 3:

[0150] The server analyzes the text data using a sentiment analysis library to detect negative expressions. Specifically, it uses a sentiment analysis library (e.g., NLTK or Google Cloud Natural Language API) to calculate the sentiment score of the text data. If the negative score exceeds a certain threshold, the server identifies the text as negative.

[0151] Input: Text data

[0152] Output: Text data containing negative expressions

[0153] Step 4:

[0154] The server converts the identified negative expressions into positive expressions using a generative artificial intelligence model. Specifically, it uses a generative artificial intelligence model (e.g., GPT-3) and provides a specific prompt. An example of the prompt is "Please convert the following negative text into a positive expression:"

[0155] Input: Text data containing negative expressions

[0156] Output: Text data containing positive expressions

[0157] Step 5:

[0158] The server converts the text data containing positive expressions into voice data using a voice synthesis tool (e.g., Google Cloud Text-to-Speech).

[0159] Input: Text data containing positive expressions

[0160] Output: Positive voice data

[0161] Step 6:

[0162] The server transmits the generated positive voice data to the terminal, and the terminal plays the received voice data to the user through a speaker, allowing the user to receive positive voice information in real time.

[0163] Input: Positive voice data

[0164] Output: Providing positive audio information to the user

[0165] (Application example 1)

[0166] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0167] In the work environment within a factory, workers are often exposed to negative comments and inappropriate language, which can reduce work efficiency and even lead to deterioration of interpersonal relationships. In such an environment, workers experience increased psychological stress and the risk of work errors and accidents increases. The objective of the present invention is to solve these problems and provide a means to provide an environment in which workers can work with a positive attitude.

[0168] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0169] In this invention, the server includes means for converting surrounding voices into text data using a voice recognition function, means for analyzing and detecting negative expressions using a sentiment analysis function, and means for converting the detected negative expressions into positive expressions using a generative artificial intelligence model. This makes it possible to provide negative comments as positive voice information in real time to workers wearing head-mounted displays in factories.

[0170] "Ambient audio" refers to any audio signals emanating within the environment in which the user is present.

[0171] "Capturing means" refers to a device or software that collects audio data and provides it in a form that can be used for further processing.

[0172] "Means for converting into text data" refers to a process or system that analyzes collected voice data and converts it into corresponding text information.

[0173] "Means for analyzing and detecting negative expressions" refers to an algorithm or system for detecting and identifying expressions with negative or downside meanings contained in text data.

[0174] "Means for converting into positive expressions" refers to a generative artificial intelligence model that has the ability to automatically change detected negative expressions into positive or forward-looking expressions.

[0175] "Means for converting into audio data" refers to a speech synthesis system or tool that can convert text data into audio format.

[0176] "Means for playing" refers to a device or system that outputs the converted audio data in an audible form to a user.

[0177] "Head-mounted display" refers to a visual and audio output device that is worn on a user's head and can provide audio information in real time.

[0178] "Worker" refers to a person who performs work in a specific environment such as a factory or workshop.

[0179] "Real-time presentation means" refers to a system or process for immediately conveying captured and processed audio data to a user.

[0180] This invention is a system used in factories in which workers wearing head-mounted displays convert negative comments made by those around them into positive expressions and provide them as voice in real time.

[0181] System configuration and specific operation

[0182] 1. System Configuration

[0183] The system mainly consists of the following components:

[0184] Terminal: A device for workers wearing a head-mounted display (HMD). It has a built-in microphone and speaker, picks up surrounding sounds, and plays back the audio data.

[0185] Server: Converts voice data to text, performs sentiment analysis and expression conversion, and converts the conversion results back into voice data.

[0186] 2. Hardware and Software Used

[0187] Head-mounted display (HMD): A device worn by the worker that inputs and outputs voice.

[0188] Microphone: A sound capture device built into the HMD.

[0189] Server: A powerful computer system.

[0190] Speech Recognition API: Uses the speech_recognition library.

[0191] Sentiment Analysis Library: Uses the sentiment analysis functionality from the transformers library.

[0192] Generative AI models: Generative artificial intelligence models such as GPT-3.

[0193] Speech synthesis tool: Uses the gTTS (Google Text-to-Speech) library.

[0194] 3. Operational Details

[0195] The system operates by having the server receive the voice data sent from the terminal and perform the following processing. First, it converts the voice data into text data using a voice recognition API and analyzes the text data using a sentiment analysis library. Next, it uses a generative artificial intelligence model to convert any negative expressions detected into positive ones. After that, it converts the converted positive text data into voice data using a voice synthesis tool and sends the result to the terminal. The terminal then plays this positive voice data in real time and provides it to the worker.

[0196] 4. Usage example

[0197] Example 1:

[0198] If a worker in a factory hears the following statement:

[0199] "Too many mistakes today!"

[0200] The microphone picks up this speech and sends it to the server, which converts it into text using a speech recognition API and uses a sentiment analysis library to identify negative expressions. The generative AI model is then given the following prompt:

[0201] text

[0202] Turn the following statement into a positive: I made too many mistakes today!

[0203] The generative AI model transforms this negative expression into:

[0204] text

[0205] I made a few mistakes today, but I'm sure I can improve next time.

[0206] This positive text data is converted into voice data using a speech synthesis tool and sent to the terminal. The head-mounted display plays the converted voice data and provides it to the worker.

[0207] This allows workers to obtain positive information in real time, reducing psychological stress and improving work efficiency.

[0208] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0209] Step 1:

[0210] The device captures the surrounding audio through a microphone. This captured audio data is sent to the server in streaming format. The input data is the surrounding audio, and the output data is the audio stream sent to the server.

[0211] Step 2:

[0212] The server converts the received voice data into text data using a voice recognition API. Here, the voice data is converted into text information. The input data is an audio stream, and the output data is text data.

[0213] Step 3:

[0214] The server analyzes the converted text data using a sentiment analysis library to detect negative expressions. The input data is text data, and the output data is the analysis result of negative expressions. Specifically, it determines whether the words and expressions contained in the text are negative.

[0215] Step 4:

[0216] The server uses a generative AI model to convert the detected negative expressions into positive expressions. At this time, it passes a prompt sentence to the generative AI model to instruct the conversion. The input data is negative text data, and the output data is the text data converted into positive. An example of a specific prompt sentence is as follows:

[0217] text

[0218] Turn the following statement into a positive: I made too many mistakes today!

[0219] Step 5:

[0220] The server converts the text data converted into positive data into voice data using a voice synthesis tool. The input data is the positive text data, and the output data is voice data.

[0221] Step 6:

[0222] The server sends the generated audio data to the terminal, where the input data is the audio data and the output data is the audio stream sent to the terminal.

[0223] Step 7:

[0224] The device plays back the received positive voice data and provides it to the user. The input data is the voice data, and the output data is the voice that the user hears. Specifically, the voice is played back through the speaker of the head-mounted display.

[0225] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0226] This invention achieves more advanced emotion management by combining a system that captures the sounds around the user, converts them into text, converts negative expressions into positive ones, and provides them to the user again as audio, with an emotion engine that recognizes the user's emotions. Each element and its operation will be explained in detail below.

[0227] 1. System Configuration

[0228] (a) Terminal

[0229] The device is equipped with a microphone, a speaker or earphones, and an emotion engine (camera, microphone, biometric sensors, etc.) to recognize the user's emotions. The device captures the surrounding sounds and sends them to the server. It also plays back the audio data sent from the server.

[0230] (b) Server

[0231] The server has the following functions:

[0232] Speech recognition function: Converts voice data into text.

[0233] Sentiment analysis function: Analyzes text data and detects negative expressions.

[0234] Transformation function: Transform negative expressions into positive expressions.

[0235] Speech synthesis function: Converts positive text data back into speech data.

[0236] Emotion recognition function: Recognizes the user's emotions and reflects the results in other analysis and conversion processes.

[0237] 2. Program Processing

[0238] (a) Audio capture and transmission

[0239] The device picks up ambient sounds through a microphone and transmits the audio data to the server in streaming format. At the same time, the device's emotion engine recognizes the user's emotions (voice tone, facial expressions, biometric signals, etc.) and transmits the data to the server.

[0240] (b) Converting voice data into text

[0241] The server analyzes the received voice data and converts it into text data using a speech recognition API, which is used as input data for sentiment analysis.

[0242] (c) Detection of negative expressions

[0243] The server analyzes the text data using a sentiment analysis library to detect negative expressions. It also takes into account the user's emotional data from the emotion engine to improve the accuracy of the analysis. For example, if the user is feeling stressed, the analysis results will be set to be more sensitive.

[0244] (d) Transforming negative expressions into positive ones

[0245] The server uses a generative artificial intelligence model to convert negative expressions into positive ones, and uses the user's emotional data to make more appropriate conversions. For example, if the user is relaxed, the tone will be adjusted to be calmer.

[0246] (e) Vocalization and delivery of positive expressions

[0247] The server converts the converted positive text data into voice data using a speech synthesis tool, which is then sent to the device and played back in real time.

[0248] (f) Audio playback

[0249] The device receives the audio data sent from the server and plays it back to the user through a speaker or earphone, protecting the user from the negative audio environment around them and allowing them to receive only positive information.

[0250] 3. Specific Examples

[0251] Example 1: Converting an angry boss's words

[0252] Boss' statement:

[0253] "Why can't you do something so simple? I can't believe it!"

[0254] The device picks up this speech with a microphone and sends the audio data to the server. At the same time, the device captures the user's facial expression with a camera and sends the emotional data to the server.

[0255] The server uses a speech recognition API to generate the text "Why can't you do something so simple? I can't believe it!"

[0256] The server analyzes this text using a sentiment analysis library and determines that it is a negative comment. Data from the emotion engine determines that the user is stressed.

[0257] The server uses a generative artificial intelligence model to translate this into "You've made a lot of mistakes today. Let's work together to figure out how to solve the problem." The tone is adjusted to be gentler depending on the user's stress level.

[0258] The server converts the converted text into voice data using a voice synthesis tool and sends it to the terminal.

[0259] The terminal receives the converted audio data and plays it back to the user.

[0260] Example 2: Converting a colleague's grumpy tone

[0261] Colleagues say:

[0262] "I'm tired. I'm starting to hate this job."

[0263] The device picks up this speech with a microphone and sends the voice data to the server, where the emotion engine analyzes the user's voice tone and sends the data to the server.

[0264] The server uses a speech recognition API to generate the text "I'm tired. I hate this job."

[0265] The server analyzes this text using a sentiment analysis library and determines that it is a negative comment. Data from the sentiment engine determines that the user is relaxed.

[0266] The server uses a generative artificial intelligence model to translate this into "I'm a little tired today, but I believe I can get through it," adjusting the tone to match the user's state of relaxation.

[0267] The server converts the converted text into voice data using a voice synthesis tool and sends it to the terminal.

[0268] The terminal receives the converted audio data and plays it back to the user.

[0269] In this way, this system converts negative expressions around the user into positive ones, and by providing optimal voice information taking into account the user's emotional state, it is possible to improve the user's mental health and productivity.

[0270] The processing flow will be explained below.

[0271] Step 1:

[0272] The device captures surrounding sounds through a microphone. At the same time, it uses a camera and biometric sensors to collect the user's facial expressions and biometric signals (e.g., heart rate) to recognize the user's emotions. The collected voice data and emotion data are sent to a server in real time.

[0273] Step 2:

[0274] The server receives the voice data sent from the device, converts it into text data using the Google speech recognition API or other voice recognition technology, and simultaneously receives and analyzes the user's emotion data sent from the device.

[0275] Step 3:

[0276] The server inputs the converted text data into a sentiment analysis library for analysis. The sentiment analysis library detects negative expressions in the text data and identifies negative or unhappy speech. At the same time, it determines the user's current emotional state based on the user's emotional data.

[0277] Step 4:

[0278] The server uses a generative AI model to convert negative expressions in the text into positive ones, taking into account the user's emotional data. For example, if the user is feeling stressed, the text will be converted into a softer, more positive one.

[0279] Step 5:

[0280] The server inputs the generated positive expressions into a text-to-speech (TTS) tool to generate voice data, which is adjusted in tone and speed depending on the user's emotional state.

[0281] Step 6:

[0282] The server then transmits the generated audio data to the terminal in streaming format, delivering the audio to the user in real time.

[0283] Step 7:

[0284] The device decodes the voice data received from the server and plays it through a speaker or earphones, allowing the user to receive voice information converted into positive expressions in real time.

[0285] Specific examples

[0286] Example 1: Converting an angry boss's words

[0287] Step 1:

[0288] The device picks up the boss's statement, "Why can't you do something so simple? I can't believe it!", through a microphone. At the same time, the device captures the user's facial expression with a camera and also obtains emotional data (facial expression, heart rate). This data is then sent to the server.

[0289] Step 2:

[0290] The server receives the voice data and uses a speech recognition API to generate the text "Why can't you do something so simple? I can't believe it!" At the same time, it determines from the emotion data that the user is feeling stressed.

[0291] Step 3:

[0292] The server uses a sentiment analysis library to determine that the text is a negative statement.

[0293] Step 4:

[0294] The server uses a generative artificial intelligence model to translate the text into something like, "You've made a lot of mistakes today. Let's work together to figure out how to solve the problem." The tone is adjusted to calm the user based on their stress level.

[0295] Step 5:

[0296] The server inputs positively-phrased text into a speech synthesis tool to generate speech data, the tone and speed of which are adjusted to help the user relax.

[0297] Step 6:

[0298] The server transmits this audio data to the terminal.

[0299] Step 7:

[0300] The device then plays back the received voice data and delivers it to the user, who can hear a positive, calm voice saying in real time, "You've made a lot of mistakes today. Let's think together about how to solve the problem."

[0301] Example 2: Converting a colleague's grumpy tone

[0302] Step 1:

[0303] The device uses a microphone to capture a colleague's statement, "I'm tired. I'm starting to hate this job." The emotion engine analyzes the user's tone of voice and sends the data to the server.

[0304] Step 2:

[0305] The server receives the voice data and uses a speech recognition API to generate the text "I'm tired. I hate this job." At the same time, it checks the emotion engine to see if the user is relaxed.

[0306] Step 3:

[0307] The server uses a sentiment analysis library to determine that this text is negative.

[0308] Step 4:

[0309] The server uses a generative artificial intelligence model to translate the text into "I'm a little tired today, but I believe I can get through it," adjusting the tone and speed to match the user's level of relaxation.

[0310] Step 5:

[0311] The server inputs the text of the positive expression into a speech synthesis tool to generate speech data.

[0312] Step 6:

[0313] The server transmits this audio data to the terminal.

[0314] Step 7:

[0315] The device plays back the received voice data and delivers it to the user, who can hear a positive, calm tone in real time: "I'm a little tired today, but I believe I can get through it."

[0316] In this way, this system converts negative expressions from the surrounding environment into positive ones, and by providing optimal voice information taking into account the user's emotional state, it is possible to improve the user's mental health and productivity.

[0317] Example 2

[0318] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0319] In recent years, many people have become aware of the psychological impact of negative expressions and stressful environments at work and at home. However, there are no systems that can convert these negative expressions into positive expressions in real time and take the user's emotional state into account. Therefore, there is a need to develop a system that allows users to receive positive information without being exposed to negative environments.

[0320] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for acquiring ambient sound, means for converting the acquired sound into text data, means for analyzing and detecting negative expressions from the text data, means for acquiring user emotion data, means for converting the detected negative expressions into positive expressions, means for converting the converted positive expressions into voice data, and means for playing the voice data. This allows the user to receive voice information in a state where negative expressions have been converted into positive expressions.

[0321] "Means for acquiring ambient sound" refers to a device or method for collecting environmental sounds and conversations using a microphone or other sensors.

[0322] "Means for converting acquired voice into text data" refers to a device or method that converts voice signals into text information using technology such as a voice recognition service API.

[0323] A "means for analyzing and detecting negative expressions from text data" is an apparatus or method that uses natural language processing techniques to identify negative sentiments and expressions within text data.

[0324] "Means for acquiring user emotional data" refers to devices or methods that use cameras, microphones, biometric sensors, etc. to collect emotional information such as the user's facial expressions, voice tone, and heart rate.

[0325] A "means for converting detected negative expressions into positive expressions" is a device or method that uses a generative artificial intelligence model to modify negative parts of text data into semantically positive ones.

[0326] The "means for converting the converted positive expression into voice data" refers to a device or method for converting positive text data into a voice signal using a voice synthesis tool that generates voice from text.

[0327] "Means for reproducing audio data" refers to a device or method for transmitting audio data to a user in real time through a speaker, earphones, etc.

[0328] This invention is a system that captures the voices around the user, converts them into text, converts negative expressions into positive ones, and provides them to the user as voice again. Furthermore, by combining it with an emotion engine that recognizes the user's emotions, it provides more advanced emotion management. Each element and its operation procedure are explained in detail below.

[0329] 1. System Configuration

[0330] (a) Terminal

[0331] The device is equipped with a microphone, a speaker or earphones, and an emotion engine (camera, microphone, biometric sensors, etc.) to recognize the user's emotions. The device captures the surrounding sounds and sends them to the server. It also receives the audio data sent from the server and plays it back to the user.

[0332] (b) Server

[0333] The server has the following functions:

[0334] Speech recognition function: Converts voice data into text. Specifically, it uses a speech recognition service API (e.g., a general cloud-based speech recognition API).

[0335] Sentiment analysis function: Analyzes text data and detects negative expressions. Uses a sentiment analysis library.

[0336] Conversion function: Uses a generative artificial intelligence model to convert negative expressions into positive ones.

[0337] Speech synthesis function: Use a speech synthesis tool to convert positive text data back into speech data.

[0338] Emotion recognition function: Recognizes the user's emotions and reflects that emotional data in other analysis and conversion processes.

[0339] 2. Explanation of specific operations

[0340] Audio capture and transmission

[0341] The device picks up ambient sounds through a microphone. This audio data is sent to the server in real time in streaming format. At the same time, the device's emotion engine recognizes the user's emotions and sends data such as facial expressions, voice tone, and biometric signals to the server.

[0342] Converting audio data to text

[0343] The server receives the voice data sent from the device and converts it into text using a voice recognition API.

[0344] Detecting negative expressions

[0345] The server analyzes text data using a sentiment analysis library to detect negative expressions, and also considers user sentiment data from an emotion engine to improve analysis accuracy.

[0346] Transforming negative expressions into positive ones

[0347] The server uses a generative artificial intelligence model to convert negative expressions into positive ones, taking into account the user's emotional data to make the conversion more appropriate.

[0348] Vocalization and delivery of positive expressions

[0349] The server reconverts the converted positive text data into voice data using a voice synthesis tool, and the voice data is sent to the terminal.

[0350] Playing audio

[0351] The terminal receives the audio data sent from the server and plays it back to the user through a speaker or earphones.

[0352] 3. Specific Examples

[0353] Example 1: Converting an angry boss's words

[0354] Boss says: "Why can't you do something so simple? I can't believe it!"

[0355] The device picks up this speech with a microphone and transmits the audio data to the server in real time. At the same time, the device captures the user's facial expression with a camera and transmits the data to the server.

[0356] The server uses a speech recognition API to generate text data of the utterance.

[0357] The server uses a sentiment analysis library to detect that the generated text is negative, and based on data from the emotion engine, recognizes that the user is feeling stressed.

[0358] The server uses a generative artificial intelligence model to translate the utterance into a positive expression, such as "You made a lot of mistakes today. Let's think together about how to solve the problem."

[0359] The server uses a speech synthesis tool to convert the text into voice data and send it to the terminal.

[0360] The device receives the converted audio and plays it back to the user through speakers or earphones.

[0361] Example 2: Converting a colleague's grumpy tone

[0362] A colleague says: "I'm tired. I'm going to hate this job."

[0363] The device picks up this speech with a microphone and sends the voice data to the server in real time. At the same time, the emotion engine analyzes the user's voice tone and sends the data to the server.

[0364] The server uses a speech recognition API to generate text data of the utterance.

[0365] The server uses a sentiment analysis library to detect that the generated text is negative, and based on data from the emotion engine, recognizes that the user is relaxed.

[0366] The server uses a generative artificial intelligence model to translate the utterance into a positive expression, generating the text, "I'm a little tired today, but I believe I can get through it."

[0367] The server uses a speech synthesis tool to convert the text into voice data and send it to the terminal.

[0368] The device receives the converted audio and plays it back to the user through speakers or earphones.

[0369] As described above, this system converts negative comments made by users around them into positive ones, and furthermore, by taking into account the user's emotional state and providing optimal voice information, it is possible to improve the user's mental health and productivity.

[0370] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0371] Step 1: Acquire audio

[0372] The device uses a microphone to capture surrounding sounds in real time. For example, it collects the user's voice while speaking and environmental sounds. The input is audio signals from the external environment, and the output is digital audio data. At the same time, the device's emotion engine (camera, microphone, biometric sensors, etc.) captures the user's facial expressions, voice tone, and biometric signals. The input is the user's emotional state, and the output is the user's emotional data.

[0373] Step 2: Sending audio data

[0374] The terminal transmits the acquired voice data and emotion data to the server in streaming format. The input here is the acquired voice data and the user's emotion data, and the output is data transmitted to the server via a network. Specifically, data transmission is performed using wireless communication.

[0375] Step 3: Converting audio data to text

[0376] The server receives the voice data sent from the device and converts it into text data using a speech recognition API (e.g., a common cloud-based speech recognition service). The input is digital voice data, and the output is text data. For example, the generated text might say, "Why can't you do something so simple?" This process requires network and computer resources.

[0377] Step 4: Detecting negative expressions

[0378] The server uses a sentiment analysis library to analyze the converted text data and detect negative expressions. It also uses the user's emotional data from the emotion engine for analysis. The input is the text data and the user's emotional data, and the output is the identification result of text containing negative expressions. Specifically, the text "Why can't you do something so simple?" is determined to be negative.

[0379] Step 5: Transform negative expressions into positive ones

[0380] The server uses a generative artificial intelligence model (e.g., a generative AI model) to convert the detected negative expressions into positive expressions. The input is text data containing negative expressions and the user's emotional data, and the output is text data converted into positive expressions. For example, "Why can't you do something so simple?" is converted into "You made quite a few mistakes today, but let's find a solution together." This process requires natural language processing technology and a generative AI model.

[0381] Step 6: Vocalize and deliver positive feedback

[0382] The server converts the converted positive text data into voice data using a voice synthesis tool (e.g., a voice synthesis engine). The input is the positive text data, and the output is voice data. For example, a voice saying, "You made quite a few mistakes today, but let's find a solution together" is generated. This voice data is then sent to the device.

[0383] Step 7: Playing Audio

[0384] The terminal receives the voice data sent from the server and plays it back to the user through a speaker or earphone. The input is the voice data sent from the server, and the output is the voice played back to the user. This process allows the user to receive positively converted voice information in real time.

[0385] (Application example 2)

[0386] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0387] In communication between employees and customers in physical stores, responding appropriately to negative comments or complaints from customers can be mentally stressful for employees and can reduce work efficiency. Furthermore, if an employee's emotional state is unstable, their response to customers will be poor. To solve this problem, a system is needed that can convert surrounding negative vocal expressions into positive ones and provide optimal vocal information that takes into account the employee's emotional state.

[0388] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0389] In this invention, the server includes means for converting voice into text data, means for analyzing and detecting negative expressions from the text data, means for converting the detected negative expressions into positive expressions, and means for utilizing an emotion engine that recognizes the user's emotions. This enables employees to receive negative comments from customers as voice that has been converted into positive ones, thereby reducing their mental burden and improving the quality of customer service.

[0390] The "means for acquiring ambient sound" is a function for acquiring sound data in the environment using a microphone, a sound acquisition device, or the like.

[0391] "Means for converting acquired voice data into text data" refers to the process of converting acquired voice data into character information (text data) using voice recognition technology or API.

[0392] "Means for analyzing and detecting negative expressions from text data" refers to a function that uses a sentiment analysis library and natural language processing technology to identify expressions that indicate negative emotions or dissatisfaction from input text data.

[0393] The "means for converting detected negative expressions into positive expressions" is a means for converting detected negative text into positive content using a generative artificial intelligence model.

[0394] "Means of utilizing an emotion engine that recognizes the user's emotions" refers to a technology that uses a camera or biometric sensor to monitor the user's facial expressions and biometric signals, and analyzes them to recognize the user's emotional state.

[0395] The "means for reproducing audio data" is a function for allowing the user to hear the generated audio data using an audio reproduction device such as a speaker or earphones.

[0396] This invention is a system designed to support store employees when dealing with customers. Specifically, the system captures surrounding voices, converts them into text data, converts negative expressions into positive ones, and provides the voice data to the store employees. Furthermore, by combining it with an emotion engine that recognizes the emotions of the user (employee), the system aims to achieve advanced emotion management by providing optimal voice information according to the employee's emotional state.

[0397] System Configuration

[0398] (a) Terminal

[0399] The device is equipped with smart glasses, a microphone, a speaker or earphones, and an emotion engine (camera, biometric sensors, etc.) that recognizes the user's emotions. The device captures the surrounding audio and sends it to the server. It also plays back the audio data sent from the server.

[0400] (b) Server

[0401] The server has the following functions:

[0402] Speech recognition function: Converts voice data into text (e.g., Google Speech-to-Text API).

[0403] Sentiment analysis functions: Analyze text data and detect negative expressions (e.g., NLTK or Transformers libraries).

[0404] Transformation function: Transforming negative expressions into positive expressions (e.g., generative AI models such as GPT-3).

[0405] Speech synthesis function: Converts positive text data back into speech data (e.g. gTTS).

[0406] Emotion recognition: Recognizes the user's emotions and incorporates the results into other analysis and transformation processes (e.g., camera and biometric sensors).

[0407] Program processing

[0408] The device uses a microphone to capture the conversation between the customer and employee and sends the audio data in streaming format to the server. At the same time, the device's emotion engine recognizes the user's emotions (voice tone, facial expressions, biometric signals, etc.) and sends this data to the server. The server analyzes the received audio data and converts it into text data using a speech recognition API. This data is used as input data for emotion analysis.

[0409] The server uses a sentiment analysis library to analyze text data and detect negative expressions. It also takes into account the user's emotional data from the emotion engine to improve analysis accuracy. For example, if the user is feeling stressed, the analysis results will be set more sensitively. The server then uses a generative artificial intelligence model to convert negative expressions into positive ones. The server uses the user's emotional data to perform more appropriate conversions. For example, if the user is relaxed, the tone will be adjusted to be gentler.

[0410] The server then converts the converted positive text data into audio data using a speech synthesis tool. This data is then sent to the device and played in real time. The device then receives the audio data sent from the server and plays it back to the user through a speaker or earphones. This protects the user from the negative audio environment around them and allows them to receive only positive information.

[0411] Specific examples

[0412] Example 1: When a customer complains

[0413] Customer says: "I can't use this product at all, what's going on?"

[0414] What employees ask: "Are you having trouble with this product? Is there anything I can help you with?"

[0415] Prompt Sentence Examples

[0416] "This product is completely unusable. What's going on?" Please change this to a positive expression.

[0417] By using this system, employees can reduce their mental burden and improve the quality of customer service.

[0418] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0419] Step 1:

[0420] The device captures the surrounding audio through a microphone, and this audio data is sent to the server in real time. The input is the surrounding audio data, and the output is sending the audio data to the server.

[0421] Step 2:

[0422] The device's emotion engine (including the camera and biometric sensors) acquires the user's emotional data (facial expressions, biometric signals, etc.). This data is also sent to the server at the same time as the voice data. The input is facial expressions and biometric signals, and the output is sending emotional data to the server.

[0423] Step 3:

[0424] The server converts the received voice data into text data using a speech recognition API. The input is the voice data, and the output is the corresponding text data.

[0425] Step 4:

[0426] The server analyzes the converted text data using a sentiment analysis library to detect negative expressions. The input is text data, and the output is the result of whether or not it contains negative expressions.

[0427] Step 5:

[0428] The server uses a generative artificial intelligence model (e.g., GPT-3) to convert text containing negative expressions into positive ones, while also taking into account the user's emotional data to adjust the conversion process. The input is negative text data and emotional data, and the output is the text data converted into positive ones.

[0429] Step 6:

[0430] The server converts the positive text data into speech data using a speech synthesis tool (e.g., gTTS). The input is the positive text data, and the output is speech data.

[0431] Step 7:

[0432] The server sends the generated audio data to the terminal. The input is audio data, and the output is sending audio data to the terminal.

[0433] Step 8:

[0434] The device then plays the received audio data through a speaker or earphone. The input is audio data, and the output is the audio the user hears. This protects the user from the negative audio environment around them and allows them to receive only positive information.

[0435] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0436] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0437] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0438] [Second embodiment]

[0439] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0440] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0441] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0442] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0443] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0444] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0445] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0446] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0447] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0448] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0449] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0450] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0451] This invention is a system that converts negative comments from people around you into positive expressions and provides positive voice information to the user. Each element and its operation will be described in detail below.

[0452] 1. System Configuration

[0453] (a) Terminal

[0454] The device is equipped with a microphone to capture surrounding sounds, and hardware and software for processing and playing back the audio data. The device transmits the collected audio data to a server in real time and plays back the audio data from the server in real time.

[0455] (b) Server

[0456] The server has the following functions:

[0457] Speech recognition function: Converts voice data into text.

[0458] Sentiment analysis function: Analyzes text data and detects negative expressions.

[0459] Conversion function: Converts detected negative expressions into positive expressions.

[0460] Speech synthesis function: Converts positive text data back into speech data.

[0461] 2. Program Processing

[0462] (a) Audio capture and transmission

[0463] The device picks up ambient sound through a microphone, and the acquired sound data is sent to the server in streaming format.

[0464] (b) Converting voice data into text

[0465] The server receives the voice data sent from the device, converts it into text using a speech recognition API, and then prepares the text data for analysis.

[0466] (c) Detection of negative expressions

[0467] The server analyzes the text data using a sentiment analysis library, and based on the analysis results, negative expressions and unpleasant speech patterns are identified.

[0468] (d) Transforming negative expressions into positive ones

[0469] The server converts the detected negative expressions into positive expressions using a generative artificial intelligence model.

[0470] (e) Vocalization and delivery of positive expressions

[0471] The server converts the positively converted text data into voice data using a voice synthesis tool, and the voice data is sent to the terminal.

[0472] (f) Audio playback

[0473] The terminal receives the positive voice data sent from the server and plays it back to the user in real time, so that the user can obtain positive voice information.

[0474] 3. Specific Examples

[0475] Example 1: Converting an angry boss's words

[0476] Boss' statement:

[0477] "Why can't you do something so simple? I can't believe it!"

[0478] The device picks up this speech with a microphone and sends it to the server.

[0479] The server uses a speech recognition API to generate the text "Why can't you do something so simple? I can't believe it!"

[0480] The server uses a sentiment analysis library to determine that this text is negative.

[0481] The server uses a generative artificial intelligence model to convert this into "You made a lot of mistakes today. Let's think together about how to solve the problem."

[0482] The server converts the converted text into voice data using a voice synthesis tool and sends it to the terminal.

[0483] The terminal receives the converted audio data and plays it back to the user.

[0484] Example 2: Converting a colleague's grumpy tone

[0485] Colleagues say:

[0486] "I'm tired. I'm starting to hate this job."

[0487] The device picks up this speech with a microphone and sends it to the server.

[0488] The server uses a speech recognition API to generate the text "I'm tired. I hate this job."

[0489] The server uses a sentiment analysis library to determine that this text is negative.

[0490] The server uses a generative artificial intelligence model to translate this into "I'm a little tired today, but I believe I can get through it."

[0491] The server converts the converted text into voice data using a voice synthesis tool and sends it to the terminal.

[0492] The terminal receives the converted audio data and plays it back to the user.

[0493] In this way, this system converts negative expressions into positive ones, thereby improving the user's productivity.

[0494] The processing flow will be explained below.

[0495] Step 1:

[0496] The device picks up ambient sound through a microphone, and the sound data is recorded in real time and sent to a server via streaming.

[0497] Step 2:

[0498] The server receives the voice data sent from the device. This voice data is converted into text using Google Speech Recognition API or other voice recognition technology. The converted text data is temporarily stored in an internal buffer.

[0499] Step 3:

[0500] The server inputs the converted text data into a sentiment analysis library, which analyzes the text data for negative expressions and detects negative or irritable speech patterns. The results are saved as negative flags.

[0501] Step 4:

[0502] The server checks the analysis results from the sentiment analysis library and extracts negative phrases, which are then converted into positive ones using a generative AI model.

[0503] Step 5:

[0504] The server inputs the generated positive expressions into a text-to-speech (TTS) tool to generate voice data, which is then temporarily stored.

[0505] Step 6:

[0506] The server then transmits the generated audio data to the terminal in streaming format, delivering the audio to the user in real time.

[0507] Step 7:

[0508] The device decodes the voice data received from the server and plays it through a speaker or earphones, allowing the user to receive voice information converted into positive expressions in real time.

[0509] Example 1

[0510] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0511] In conventional systems, negative comments from those around you could be directly conveyed to the user, causing psychological stress and reducing productivity. Furthermore, the process of converting negative comments into positive ones was complicated, making it difficult to respond in real time. This made it difficult for users to work in a comfortable environment.

[0512] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0513] In this invention, the server includes means for converting received speech into text data, means for analyzing the text data to detect negative expressions, and means for converting the detected negative expressions into positive expressions using a generative artificial intelligence model. This makes it possible to acquire surrounding speech, convert the negative expressions into positive expressions in real time, and provide them to the user.

[0514] "Audio capture device" refers to a device for capturing ambient sound. Examples include a microphone.

[0515] A "speech recognition service API" refers to an application programming interface for converting voice data into text, often provided as a cloud-based service.

[0516] "Server" refers to a computer system with data processing capabilities that receives, converts, and analyzes voice data.

[0517] A "generative artificial intelligence model" refers to an algorithm that generates new text based on a given prompt, based on deep learning techniques, such as GPT-3.

[0518] "Real-time" refers to data processing and communication occurring instantaneously with minimal delay.

[0519] "Text data" refers to the result of converting audio data into text information, which serves as the basis for analysis and other processing.

[0520] "Negative expressions" refer to negative content or expressions identified through sentiment analysis, which may cause psychological stress to users.

[0521] "Positive expressions" refer to positive content or expressions that replace negative expressions and are used to maintain the user's psychological stability.

[0522] "Audio data" refers to the digital representation of sound that is used to be heard by a user through a playback device.

[0523] This invention is a system that converts negative comments from people around you into positive expressions and provides positive voice information to the user. Each element and its operation will be described in detail below.

[0524] 1. System Configuration

[0525] (a) Terminal

[0526] The device is equipped with a sound capture device (microphone) for capturing surrounding sounds, and hardware and software for processing and playing back the sound data. The device transmits the collected sound data to a server in real time and plays back the sound data from the server in real time.

[0527] (b) Server

[0528] The server has the following functions:

[0529] Speech recognition function: Converts voice data into text. Specifically, it uses a speech recognition service API (e.g., Google Cloud Speech-to-Text).

[0530] Sentiment analysis function: Analyzes text data and detects negative expressions. Specifically, it uses a sentiment analysis library (e.g., NLTK or Google Cloud Natural Language API).

[0531] Conversion function: Converts detected negative expressions into positive expressions using a generative artificial intelligence model (e.g., GPT-3).

[0532] Speech synthesis function: Positive text data is converted back into voice data, specifically using a speech synthesis tool (e.g., Google Cloud Text-to-Speech).

[0533] 2. Operation overview

[0534] The device captures ambient sounds and sends the audio data in streaming format to the server. The server then converts the received audio data into text data using a speech recognition API. The server then analyzes the text data using a sentiment analysis library to identify negative expressions. The identified negative expressions are converted into positive expressions using a generative artificial intelligence model. The converted positive text data is then reconverted into audio data using a speech synthesis tool and sent to the device. Finally, the device plays the transmitted positive audio data to the user.

[0535] 3. Specific Examples

[0536] Example 1: Converting an angry boss's words

[0537] Boss' statement:

[0538] "Why can't you do something so simple? I can't believe it!"

[0539] 1. The device picks up this speech with a microphone and sends it to the server.

[0540] 2. The server uses a speech recognition API to generate the text "Why can't you do something so simple? I can't believe it!"

[0541] 3. The server uses a sentiment analysis library to determine that the text is negative.

[0542] 4. The server uses a generative artificial intelligence model to convert this into "You made a lot of mistakes today. Let's think together about how to solve the problem."

[0543] 5. The server converts the converted text into voice data using a speech synthesis tool and sends it to the device.

[0544] 6. The device receives the converted audio data and plays it back to the user.

[0545] Example 2: Converting a colleague's grumpy tone

[0546] Colleagues say:

[0547] "I'm tired. I'm starting to hate this job."

[0548] 1. The device picks up this speech with a microphone and sends it to the server.

[0549] 2. The server uses a speech recognition API to generate the text "I'm tired. I hate this job."

[0550] 3. The server uses a sentiment analysis library to determine that the text is negative.

[0551] 4. The server uses a generative artificial intelligence model to convert this to "I'm a little tired today, but I believe I can get it done."

[0552] 5. The server converts the converted text into voice data using a speech synthesis tool and sends it to the device.

[0553] 6. The device receives the converted audio data and plays it back to the user.

[0554] Prompt Sentence Examples

[0555] Change the following negative text into a positive one: "I'm exhausted. I hate this job."

[0556] In this way, this system converts negative expressions from the surrounding environment into positive ones, improving the user's psychological stability and productivity.

[0557] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0558] Step 1:

[0559] The device captures the surrounding sounds using an audio capture device, specifically, a microphone collects audio data and transmits it to a server in real time in a streaming format.

[0560] Input: Ambient audio

[0561] Output: Streaming audio data

[0562] Step 2:

[0563] The server receives the voice data sent from the device. It sends the received voice data to a voice recognition service API and converts the voice data into text data. Specifically, it uses a voice recognition API (e.g., Google Cloud Speech-to-Text) to convert the voice data into text format.

[0564] Input: Streaming audio data

[0565] Output: Text data

[0566] Step 3:

[0567] The server analyzes the text data using a sentiment analysis library to detect negative expressions. Specifically, it uses a sentiment analysis library (e.g., NLTK or Google Cloud Natural Language API) to calculate the sentiment score of the text data. It then identifies any part of the text where the negative score exceeds a certain threshold as a negative expression.

[0568] Input: Text data

[0569] Output: Text data containing negative expressions

[0570] Step 4:

[0571] The server converts the identified negative expressions into positive expressions using a generative artificial intelligence model. Specifically, it uses a generative artificial intelligence model (e.g., GPT-3) and provides a specific prompt. An example of the prompt is "Please convert the following negative text into a positive expression:"

[0572] Input: Text data containing negative expressions

[0573] Output: Text data containing positive expressions

[0574] Step 5:

[0575] The server converts the text data containing positive expressions into voice data using a voice synthesis tool (e.g., Google Cloud Text-to-Speech).

[0576] Input: Text data containing positive expressions

[0577] Output: Positive voice data

[0578] Step 6:

[0579] The server transmits the generated positive voice data to the terminal, and the terminal plays the received voice data to the user through a speaker, allowing the user to receive positive voice information in real time.

[0580] Input: Positive voice data

[0581] Output: Providing positive audio information to the user

[0582] (Application example 1)

[0583] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0584] In the work environment within a factory, workers are often exposed to negative comments and inappropriate language, which can reduce work efficiency and even lead to deterioration of interpersonal relationships. In such an environment, workers experience increased psychological stress and the risk of work errors and accidents increases. The objective of the present invention is to solve these problems and provide a means to provide an environment in which workers can work with a positive attitude.

[0585] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0586] In this invention, the server includes means for converting surrounding voices into text data using a voice recognition function, means for analyzing and detecting negative expressions using a sentiment analysis function, and means for converting the detected negative expressions into positive expressions using a generative artificial intelligence model. This makes it possible to provide negative comments as positive voice information in real time to workers wearing head-mounted displays in factories.

[0587] "Ambient audio" refers to any audio signals emanating within the environment in which the user is present.

[0588] "Capturing means" refers to a device or software that collects audio data and provides it in a form that can be used for further processing.

[0589] "Means for converting into text data" refers to a process or system that analyzes collected voice data and converts it into corresponding text information.

[0590] "Means for analyzing and detecting negative expressions" refers to an algorithm or system for detecting and identifying expressions with negative or downside meanings contained in text data.

[0591] "Means for converting into positive expressions" refers to a generative artificial intelligence model that has the ability to automatically change detected negative expressions into positive or forward-looking expressions.

[0592] "Means for converting into audio data" refers to a speech synthesis system or tool that can convert text data into audio format.

[0593] "Means for playing" refers to a device or system that outputs the converted audio data in an audible form to a user.

[0594] "Head-mounted display" refers to a visual and audio output device that is worn on a user's head and can provide audio information in real time.

[0595] "Worker" refers to a person who performs work in a specific environment such as a factory or workshop.

[0596] "Real-time presentation means" refers to a system or process for immediately conveying captured and processed audio data to a user.

[0597] This invention is a system used in factories in which workers wearing head-mounted displays convert negative comments made by those around them into positive expressions and provide them as voice in real time.

[0598] System configuration and specific operation

[0599] 1. System Configuration

[0600] The system mainly consists of the following components:

[0601] Terminal: A device for workers wearing a head-mounted display (HMD). It has a built-in microphone and speaker, picks up surrounding sounds, and plays back the audio data.

[0602] Server: Converts voice data to text, performs sentiment analysis and expression conversion, and converts the conversion results back into voice data.

[0603] 2. Hardware and Software Used

[0604] Head-mounted display (HMD): A device worn by the worker that inputs and outputs voice.

[0605] Microphone: A sound capture device built into the HMD.

[0606] Server: A powerful computer system.

[0607] Speech Recognition API: Uses the speech_recognition library.

[0608] Sentiment Analysis Library: Uses the sentiment analysis functionality from the transformers library.

[0609] Generative AI models: Generative artificial intelligence models such as GPT-3.

[0610] Speech synthesis tool: Uses the gTTS (Google Text-to-Speech) library.

[0611] 3. Operational Details

[0612] The system operates by having the server receive the voice data sent from the terminal and perform the following processing. First, it converts the voice data into text data using a voice recognition API and analyzes the text data using a sentiment analysis library. Next, it uses a generative artificial intelligence model to convert any negative expressions detected into positive ones. After that, it converts the converted positive text data into voice data using a voice synthesis tool and sends the result to the terminal. The terminal then plays this positive voice data in real time and provides it to the worker.

[0613] 4. Usage example

[0614] Example 1:

[0615] If a worker in a factory hears the following statement:

[0616] "Too many mistakes today!"

[0617] The microphone picks up this speech and sends it to the server, which converts it into text using a speech recognition API and uses a sentiment analysis library to identify negative expressions. The generative AI model is then given the following prompt:

[0618] text

[0619] Turn the following statement into a positive: I made too many mistakes today!

[0620] The generative AI model transforms this negative expression into:

[0621] text

[0622] I made a few mistakes today, but I'm sure I can improve next time.

[0623] This positive text data is converted into voice data using a speech synthesis tool and sent to the terminal. The head-mounted display plays the converted voice data and provides it to the worker.

[0624] This allows workers to obtain positive information in real time, reducing psychological stress and improving work efficiency.

[0625] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0626] Step 1:

[0627] The device captures the surrounding audio through a microphone. This captured audio data is sent to the server in streaming format. The input data is the surrounding audio, and the output data is the audio stream sent to the server.

[0628] Step 2:

[0629] The server converts the received voice data into text data using a voice recognition API. Here, the voice data is converted into text information. The input data is an audio stream, and the output data is text data.

[0630] Step 3:

[0631] The server analyzes the converted text data using a sentiment analysis library to detect negative expressions. The input data is text data, and the output data is the analysis result of negative expressions. Specifically, it determines whether the words and expressions contained in the text are negative.

[0632] Step 4:

[0633] The server uses a generative AI model to convert the detected negative expressions into positive expressions. At this time, it passes a prompt sentence to the generative AI model to instruct the conversion. The input data is negative text data, and the output data is the text data converted into positive. An example of a specific prompt sentence is as follows:

[0634] text

[0635] Turn the following statement into a positive: I made too many mistakes today!

[0636] Step 5:

[0637] The server converts the text data converted into positive data into voice data using a voice synthesis tool. The input data is the positive text data, and the output data is voice data.

[0638] Step 6:

[0639] The server sends the generated audio data to the terminal, where the input data is the audio data and the output data is the audio stream sent to the terminal.

[0640] Step 7:

[0641] The device plays back the received positive voice data and provides it to the user. The input data is the voice data, and the output data is the voice that the user hears. Specifically, the voice is played back through the speaker of the head-mounted display.

[0642] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0643] This invention achieves more advanced emotion management by combining a system that captures the sounds around the user, converts them into text, converts negative expressions into positive ones, and provides them to the user again as audio, with an emotion engine that recognizes the user's emotions. Each element and its operation will be explained in detail below.

[0644] 1. System Configuration

[0645] (a) Terminal

[0646] The device is equipped with a microphone, a speaker or earphones, and an emotion engine (camera, microphone, biometric sensors, etc.) to recognize the user's emotions. The device captures the surrounding sounds and sends them to the server. It also plays back the audio data sent from the server.

[0647] (b) Server

[0648] The server has the following functions:

[0649] Speech recognition function: Converts voice data into text.

[0650] Sentiment analysis function: Analyzes text data and detects negative expressions.

[0651] Transformation function: Transform negative expressions into positive expressions.

[0652] Speech synthesis function: Converts positive text data back into speech data.

[0653] Emotion recognition function: Recognizes the user's emotions and reflects the results in other analysis and conversion processes.

[0654] 2. Program Processing

[0655] (a) Audio capture and transmission

[0656] The device picks up ambient sounds through a microphone and transmits the audio data to the server in streaming format. At the same time, the device's emotion engine recognizes the user's emotions (voice tone, facial expressions, biometric signals, etc.) and transmits the data to the server.

[0657] (b) Converting voice data into text

[0658] The server analyzes the received voice data and converts it into text data using a speech recognition API, which is used as input data for sentiment analysis.

[0659] (c) Detection of negative expressions

[0660] The server analyzes the text data using a sentiment analysis library to detect negative expressions. It also takes into account the user's emotional data from the emotion engine to improve the accuracy of the analysis. For example, if the user is feeling stressed, the analysis results will be set to be more sensitive.

[0661] (d) Transforming negative expressions into positive ones

[0662] The server uses a generative artificial intelligence model to convert negative expressions into positive ones, and uses the user's emotional data to make more appropriate conversions. For example, if the user is relaxed, the tone will be adjusted to be calmer.

[0663] (e) Vocalization and delivery of positive expressions

[0664] The server converts the converted positive text data into voice data using a speech synthesis tool, which is then sent to the device and played back in real time.

[0665] (f) Audio playback

[0666] The device receives the audio data sent from the server and plays it back to the user through a speaker or earphone, protecting the user from the negative audio environment around them and allowing them to receive only positive information.

[0667] 3. Specific Examples

[0668] Example 1: Converting an angry boss's words

[0669] Boss' statement:

[0670] "Why can't you do something so simple? I can't believe it!"

[0671] The device picks up this speech with a microphone and sends the audio data to the server. At the same time, the device captures the user's facial expression with a camera and sends the emotional data to the server.

[0672] The server uses a speech recognition API to generate the text "Why can't you do something so simple? I can't believe it!"

[0673] The server analyzes this text using a sentiment analysis library and determines that it is a negative comment. Data from the emotion engine determines that the user is stressed.

[0674] The server uses a generative artificial intelligence model to translate this into "You've made a lot of mistakes today. Let's work together to figure out how to solve the problem." The tone is adjusted to be gentler depending on the user's stress level.

[0675] The server converts the converted text into voice data using a voice synthesis tool and sends it to the terminal.

[0676] The terminal receives the converted audio data and plays it back to the user.

[0677] Example 2: Converting a colleague's grumpy tone

[0678] Colleagues say:

[0679] "I'm tired. I'm starting to hate this job."

[0680] The device picks up this speech with a microphone and sends the voice data to the server, where the emotion engine analyzes the user's voice tone and sends the data to the server.

[0681] The server uses a speech recognition API to generate the text "I'm tired. I hate this job."

[0682] The server analyzes this text using a sentiment analysis library and determines that it is a negative comment. Data from the sentiment engine determines that the user is relaxed.

[0683] The server uses a generative artificial intelligence model to translate this into "I'm a little tired today, but I believe I can get through it," adjusting the tone to match the user's state of relaxation.

[0684] The server converts the converted text into voice data using a voice synthesis tool and sends it to the terminal.

[0685] The terminal receives the converted audio data and plays it back to the user.

[0686] In this way, this system converts negative expressions around the user into positive ones, and by providing optimal voice information taking into account the user's emotional state, it is possible to improve the user's mental health and productivity.

[0687] The processing flow will be explained below.

[0688] Step 1:

[0689] The device captures surrounding sounds through a microphone. At the same time, it uses a camera and biometric sensors to collect the user's facial expressions and biometric signals (e.g., heart rate) to recognize the user's emotions. The collected voice data and emotion data are sent to a server in real time.

[0690] Step 2:

[0691] The server receives the voice data sent from the device, converts it into text data using the Google speech recognition API or other voice recognition technology, and simultaneously receives and analyzes the user's emotion data sent from the device.

[0692] Step 3:

[0693] The server inputs the converted text data into a sentiment analysis library for analysis. The sentiment analysis library detects negative expressions in the text data and identifies negative or unhappy speech. At the same time, it determines the user's current emotional state based on the user's emotional data.

[0694] Step 4:

[0695] The server uses a generative AI model to convert negative expressions in the text into positive ones, taking into account the user's emotional data. For example, if the user is feeling stressed, the text will be converted into a softer, more positive one.

[0696] Step 5:

[0697] The server inputs the generated positive expressions into a text-to-speech (TTS) tool to generate voice data, which is adjusted in tone and speed depending on the user's emotional state.

[0698] Step 6:

[0699] The server then transmits the generated audio data to the terminal in streaming format, delivering the audio to the user in real time.

[0700] Step 7:

[0701] The device decodes the voice data received from the server and plays it through a speaker or earphones, allowing the user to receive voice information converted into positive expressions in real time.

[0702] Specific examples

[0703] Example 1: Converting an angry boss's words

[0704] Step 1:

[0705] The device picks up the boss's statement, "Why can't you do something so simple? I can't believe it!", through a microphone. At the same time, the device captures the user's facial expression with a camera and also obtains emotional data (facial expression, heart rate). This data is then sent to the server.

[0706] Step 2:

[0707] The server receives the voice data and uses a speech recognition API to generate the text "Why can't you do something so simple? I can't believe it!" At the same time, it determines from the emotion data that the user is feeling stressed.

[0708] Step 3:

[0709] The server uses a sentiment analysis library to determine that the text is a negative statement.

[0710] Step 4:

[0711] The server uses a generative artificial intelligence model to translate the text into something like, "You've made a lot of mistakes today. Let's work together to figure out how to solve the problem." The tone is adjusted to calm the user based on their stress level.

[0712] Step 5:

[0713] The server inputs positively-phrased text into a speech synthesis tool to generate speech data, the tone and speed of which are adjusted to help the user relax.

[0714] Step 6:

[0715] The server transmits this audio data to the terminal.

[0716] Step 7:

[0717] The device then plays back the received voice data and delivers it to the user, who can hear a positive, calm voice saying in real time, "You've made a lot of mistakes today. Let's think together about how to solve the problem."

[0718] Example 2: Converting a colleague's grumpy tone

[0719] Step 1:

[0720] The device uses a microphone to capture a colleague's statement, "I'm tired. I'm starting to hate this job." The emotion engine analyzes the user's tone of voice and sends the data to the server.

[0721] Step 2:

[0722] The server receives the voice data and uses a speech recognition API to generate the text "I'm tired. I hate this job." At the same time, it checks the emotion engine to see if the user is relaxed.

[0723] Step 3:

[0724] The server uses a sentiment analysis library to determine that this text is negative.

[0725] Step 4:

[0726] The server uses a generative artificial intelligence model to translate the text into "I'm a little tired today, but I believe I can get through it," adjusting the tone and speed to match the user's level of relaxation.

[0727] Step 5:

[0728] The server inputs the text of the positive expression into a speech synthesis tool to generate speech data.

[0729] Step 6:

[0730] The server transmits this audio data to the terminal.

[0731] Step 7:

[0732] The device plays back the received voice data and delivers it to the user, who can hear a positive, calm tone in real time: "I'm a little tired today, but I believe I can get through it."

[0733] In this way, this system converts negative expressions from the surrounding environment into positive ones, and by providing optimal voice information taking into account the user's emotional state, it is possible to improve the user's mental health and productivity.

[0734] Example 2

[0735] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0736] In recent years, many people have become aware of the psychological impact of negative expressions and stressful environments at work and at home. However, there are no systems that can convert these negative expressions into positive expressions in real time and take the user's emotional state into account. Therefore, there is a need to develop a system that allows users to receive positive information without being exposed to negative environments.

[0737] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for acquiring ambient sound, means for converting the acquired sound into text data, means for analyzing and detecting negative expressions from the text data, means for acquiring user emotion data, means for converting the detected negative expressions into positive expressions, means for converting the converted positive expressions into voice data, and means for playing the voice data. This allows the user to receive voice information in a state where negative expressions have been converted into positive expressions.

[0738] "Means for acquiring ambient sound" refers to a device or method for collecting environmental sounds and conversations using a microphone or other sensors.

[0739] "Means for converting acquired voice into text data" refers to a device or method that converts voice signals into text information using technology such as a voice recognition service API.

[0740] A "means for analyzing and detecting negative expressions from text data" is an apparatus or method that uses natural language processing techniques to identify negative sentiments and expressions within text data.

[0741] "Means for acquiring user emotional data" refers to devices or methods that use cameras, microphones, biometric sensors, etc. to collect emotional information such as the user's facial expressions, voice tone, and heart rate.

[0742] A "means for converting detected negative expressions into positive expressions" is a device or method that uses a generative artificial intelligence model to modify negative parts of text data into semantically positive ones.

[0743] The "means for converting the converted positive expression into voice data" refers to a device or method for converting positive text data into a voice signal using a voice synthesis tool that generates voice from text.

[0744] "Means for reproducing audio data" refers to a device or method for transmitting audio data to a user in real time through a speaker, earphones, etc.

[0745] This invention is a system that captures the voices around the user, converts them into text, converts negative expressions into positive ones, and provides them to the user as voice again. Furthermore, by combining it with an emotion engine that recognizes the user's emotions, it provides more advanced emotion management. Each element and its operation procedure are explained in detail below.

[0746] 1. System Configuration

[0747] (a) Terminal

[0748] The device is equipped with a microphone, a speaker or earphones, and an emotion engine (camera, microphone, biometric sensors, etc.) to recognize the user's emotions. The device captures the surrounding sounds and sends them to the server. It also receives the audio data sent from the server and plays it back to the user.

[0749] (b) Server

[0750] The server has the following functions:

[0751] Speech recognition function: Converts voice data into text. Specifically, it uses a speech recognition service API (e.g., a general cloud-based speech recognition API).

[0752] Sentiment analysis function: Analyzes text data and detects negative expressions. Uses a sentiment analysis library.

[0753] Conversion function: Uses a generative artificial intelligence model to convert negative expressions into positive ones.

[0754] Speech synthesis function: Use a speech synthesis tool to convert positive text data back into speech data.

[0755] Emotion recognition function: Recognizes the user's emotions and reflects that emotional data in other analysis and conversion processes.

[0756] 2. Explanation of specific operations

[0757] Audio capture and transmission

[0758] The device picks up ambient sounds through a microphone. This audio data is sent to the server in real time in streaming format. At the same time, the device's emotion engine recognizes the user's emotions and sends data such as facial expressions, voice tone, and biometric signals to the server.

[0759] Converting audio data to text

[0760] The server receives the voice data sent from the device and converts it into text using a voice recognition API.

[0761] Detecting negative expressions

[0762] The server analyzes text data using a sentiment analysis library to detect negative expressions, and also considers user sentiment data from an emotion engine to improve analysis accuracy.

[0763] Transforming negative expressions into positive ones

[0764] The server uses a generative artificial intelligence model to convert negative expressions into positive ones, taking into account the user's emotional data to make the conversion more appropriate.

[0765] Vocalization and delivery of positive expressions

[0766] The server reconverts the converted positive text data into voice data using a voice synthesis tool, and the voice data is sent to the terminal.

[0767] Playing audio

[0768] The terminal receives the audio data sent from the server and plays it back to the user through a speaker or earphones.

[0769] 3. Specific Examples

[0770] Example 1: Converting an angry boss's words

[0771] Boss says: "Why can't you do something so simple? I can't believe it!"

[0772] The device picks up this speech with a microphone and transmits the audio data to the server in real time. At the same time, the device captures the user's facial expression with a camera and transmits the data to the server.

[0773] The server uses a speech recognition API to generate text data of the utterance.

[0774] The server uses a sentiment analysis library to detect that the generated text is negative, and based on data from the emotion engine, recognizes that the user is feeling stressed.

[0775] The server uses a generative artificial intelligence model to translate the utterance into a positive expression, such as "You made a lot of mistakes today. Let's think together about how to solve the problem."

[0776] The server uses a speech synthesis tool to convert the text into voice data and send it to the terminal.

[0777] The device receives the converted audio and plays it back to the user through speakers or earphones.

[0778] Example 2: Converting a colleague's grumpy tone

[0779] A colleague says: "I'm tired. I'm going to hate this job."

[0780] The device picks up this speech with a microphone and sends the voice data to the server in real time. At the same time, the emotion engine analyzes the user's voice tone and sends the data to the server.

[0781] The server uses a speech recognition API to generate text data of the utterance.

[0782] The server uses a sentiment analysis library to detect that the generated text is negative, and based on data from the emotion engine, recognizes that the user is relaxed.

[0783] The server uses a generative artificial intelligence model to translate the utterance into a positive expression, generating the text, "I'm a little tired today, but I believe I can get through it."

[0784] The server uses a speech synthesis tool to convert the text into voice data and send it to the terminal.

[0785] The device receives the converted audio and plays it back to the user through speakers or earphones.

[0786] As described above, this system converts negative comments made by users around them into positive ones, and furthermore, by taking into account the user's emotional state and providing optimal voice information, it is possible to improve the user's mental health and productivity.

[0787] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0788] Step 1: Acquire audio

[0789] The device uses a microphone to capture surrounding sounds in real time. For example, it collects the user's voice while speaking and environmental sounds. The input is audio signals from the external environment, and the output is digital audio data. At the same time, the device's emotion engine (camera, microphone, biometric sensors, etc.) captures the user's facial expressions, voice tone, and biometric signals. The input is the user's emotional state, and the output is the user's emotional data.

[0790] Step 2: Sending audio data

[0791] The terminal transmits the acquired voice data and emotion data to the server in streaming format. The input here is the acquired voice data and the user's emotion data, and the output is data transmitted to the server via a network. Specifically, data transmission is performed using wireless communication.

[0792] Step 3: Converting audio data to text

[0793] The server receives the voice data sent from the device and converts it into text data using a speech recognition API (e.g., a common cloud-based speech recognition service). The input is digital voice data, and the output is text data. For example, the generated text might say, "Why can't you do something so simple?" This process requires network and computer resources.

[0794] Step 4: Detecting negative expressions

[0795] The server uses a sentiment analysis library to analyze the converted text data and detect negative expressions. It also uses the user's emotional data from the emotion engine for analysis. The input is the text data and the user's emotional data, and the output is the identification result of text containing negative expressions. Specifically, the text "Why can't you do something so simple?" is determined to be negative.

[0796] Step 5: Transform negative expressions into positive ones

[0797] The server uses a generative artificial intelligence model (e.g., a generative AI model) to convert the detected negative expressions into positive expressions. The input is text data containing negative expressions and the user's emotional data, and the output is text data converted into positive expressions. For example, "Why can't you do something so simple?" is converted into "You made quite a few mistakes today, but let's find a solution together." This process requires natural language processing technology and a generative AI model.

[0798] Step 6: Vocalize and deliver positive feedback

[0799] The server converts the converted positive text data into voice data using a voice synthesis tool (e.g., a voice synthesis engine). The input is the positive text data, and the output is voice data. For example, a voice saying, "You made quite a few mistakes today, but let's find a solution together" is generated. This voice data is then sent to the device.

[0800] Step 7: Playing Audio

[0801] The terminal receives the voice data sent from the server and plays it back to the user through a speaker or earphone. The input is the voice data sent from the server, and the output is the voice played back to the user. This process allows the user to receive positively converted voice information in real time.

[0802] (Application example 2)

[0803] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0804] In communication between employees and customers in physical stores, responding appropriately to negative comments or complaints from customers can be mentally stressful for employees and can reduce work efficiency. Furthermore, if an employee's emotional state is unstable, their response to customers will be poor. To solve this problem, a system is needed that can convert surrounding negative vocal expressions into positive ones and provide optimal vocal information that takes into account the employee's emotional state.

[0805] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0806] In this invention, the server includes means for converting voice into text data, means for analyzing and detecting negative expressions from the text data, means for converting the detected negative expressions into positive expressions, and means for utilizing an emotion engine that recognizes the user's emotions. This enables employees to receive negative comments from customers as voice that has been converted into positive ones, thereby reducing their mental burden and improving the quality of customer service.

[0807] The "means for acquiring ambient sound" is a function for acquiring sound data in the environment using a microphone, a sound acquisition device, or the like.

[0808] "Means for converting acquired voice data into text data" refers to the process of converting acquired voice data into character information (text data) using voice recognition technology or API.

[0809] "Means for analyzing and detecting negative expressions from text data" refers to a function that uses a sentiment analysis library and natural language processing technology to identify expressions that indicate negative emotions or dissatisfaction from input text data.

[0810] The "means for converting detected negative expressions into positive expressions" is a means for converting detected negative text into positive content using a generative artificial intelligence model.

[0811] "Means of utilizing an emotion engine that recognizes the user's emotions" refers to a technology that uses a camera or biometric sensor to monitor the user's facial expressions and biometric signals, and analyzes them to recognize the user's emotional state.

[0812] The "means for reproducing audio data" is a function for allowing the user to hear the generated audio data using an audio reproduction device such as a speaker or earphones.

[0813] This invention is a system designed to support store employees when dealing with customers. Specifically, the system captures surrounding voices, converts them into text data, converts negative expressions into positive ones, and provides the voice data to the store employees. Furthermore, by combining it with an emotion engine that recognizes the emotions of the user (employee), the system aims to achieve advanced emotion management by providing optimal voice information according to the employee's emotional state.

[0814] System Configuration

[0815] (a) Terminal

[0816] The device is equipped with smart glasses, a microphone, a speaker or earphones, and an emotion engine (camera, biometric sensors, etc.) that recognizes the user's emotions. The device captures the surrounding audio and sends it to the server. It also plays back the audio data sent from the server.

[0817] (b) Server

[0818] The server has the following functions:

[0819] Speech recognition function: Converts voice data into text (e.g., Google Speech-to-Text API).

[0820] Sentiment analysis functions: Analyze text data and detect negative expressions (e.g., NLTK or Transformers libraries).

[0821] Transformation function: Transforming negative expressions into positive expressions (e.g., generative AI models such as GPT-3).

[0822] Speech synthesis function: Converts positive text data back into speech data (e.g. gTTS).

[0823] Emotion recognition: Recognizes the user's emotions and incorporates the results into other analysis and transformation processes (e.g., camera and biometric sensors).

[0824] Program processing

[0825] The device uses a microphone to capture the conversation between the customer and employee and sends the audio data in streaming format to the server. At the same time, the device's emotion engine recognizes the user's emotions (voice tone, facial expressions, biometric signals, etc.) and sends this data to the server. The server analyzes the received audio data and converts it into text data using a speech recognition API. This data is used as input data for emotion analysis.

[0826] The server uses a sentiment analysis library to analyze text data and detect negative expressions. It also takes into account the user's emotional data from the emotion engine to improve analysis accuracy. For example, if the user is feeling stressed, the analysis results will be set more sensitively. The server then uses a generative artificial intelligence model to convert negative expressions into positive ones. The server uses the user's emotional data to perform more appropriate conversions. For example, if the user is relaxed, the tone will be adjusted to be gentler.

[0827] The server then converts the converted positive text data into audio data using a speech synthesis tool. This data is then sent to the device and played in real time. The device then receives the audio data sent from the server and plays it back to the user through a speaker or earphones. This protects the user from the negative audio environment around them and allows them to receive only positive information.

[0828] Specific examples

[0829] Example 1: When a customer complains

[0830] Customer says: "I can't use this product at all, what's going on?"

[0831] What employees ask: "Are you having trouble with this product? Is there anything I can help you with?"

[0832] Prompt Sentence Examples

[0833] "This product is completely unusable. What's going on?" Please change this to a positive expression.

[0834] By using this system, employees can reduce their mental burden and improve the quality of customer service.

[0835] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0836] Step 1:

[0837] The device captures the surrounding audio through a microphone, and this audio data is sent to the server in real time. The input is the surrounding audio data, and the output is sending the audio data to the server.

[0838] Step 2:

[0839] The device's emotion engine (including the camera and biometric sensors) acquires the user's emotional data (facial expressions, biometric signals, etc.). This data is also sent to the server at the same time as the voice data. The input is facial expressions and biometric signals, and the output is sending emotional data to the server.

[0840] Step 3:

[0841] The server converts the received voice data into text data using a speech recognition API. The input is the voice data, and the output is the corresponding text data.

[0842] Step 4:

[0843] The server analyzes the converted text data using a sentiment analysis library to detect negative expressions. The input is text data, and the output is the result of whether or not it contains negative expressions.

[0844] Step 5:

[0845] The server uses a generative artificial intelligence model (e.g., GPT-3) to convert text containing negative expressions into positive ones, while also taking into account the user's emotional data to adjust the conversion process. The input is negative text data and emotional data, and the output is the text data converted into positive ones.

[0846] Step 6:

[0847] The server converts the positive text data into speech data using a speech synthesis tool (e.g., gTTS). The input is the positive text data, and the output is speech data.

[0848] Step 7:

[0849] The server sends the generated audio data to the terminal. The input is audio data, and the output is sending audio data to the terminal.

[0850] Step 8:

[0851] The device then plays the received audio data through a speaker or earphone. The input is audio data, and the output is the audio the user hears. This protects the user from the negative audio environment around them and allows them to receive only positive information.

[0852] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0853] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0854] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0855] [Third embodiment]

[0856] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0857] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[0858] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0859] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0860] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0861] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0862] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0863] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0864] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0865] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0866] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0867] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0868] This invention is a system that converts negative comments from people around you into positive expressions and provides positive voice information to the user. Each element and its operation will be described in detail below.

[0869] 1. System Configuration

[0870] (a) Terminal

[0871] The device is equipped with a microphone to capture surrounding sounds, and hardware and software for processing and playing back the audio data. The device transmits the collected audio data to a server in real time and plays back the audio data from the server in real time.

[0872] (b) Server

[0873] The server has the following functions:

[0874] Speech recognition function: Converts voice data into text.

[0875] Sentiment analysis function: Analyzes text data and detects negative expressions.

[0876] Conversion function: Converts detected negative expressions into positive expressions.

[0877] Speech synthesis function: Converts positive text data back into speech data.

[0878] 2. Program Processing

[0879] (a) Audio capture and transmission

[0880] The device picks up ambient sound through a microphone, and the acquired sound data is sent to the server in streaming format.

[0881] (b) Converting voice data into text

[0882] The server receives the voice data sent from the device, converts it into text using a speech recognition API, and then prepares the text data for analysis.

[0883] (c) Detection of negative expressions

[0884] The server analyzes the text data using a sentiment analysis library, and based on the analysis results, negative expressions and unpleasant speech patterns are identified.

[0885] (d) Transforming negative expressions into positive ones

[0886] The server converts the detected negative expressions into positive expressions using a generative artificial intelligence model.

[0887] (e) Vocalization and delivery of positive expressions

[0888] The server converts the positively converted text data into voice data using a voice synthesis tool, and the voice data is sent to the terminal.

[0889] (f) Audio playback

[0890] The terminal receives the positive voice data sent from the server and plays it back to the user in real time, so that the user can obtain positive voice information.

[0891] 3. Specific Examples

[0892] Example 1: Converting an angry boss's words

[0893] Boss' statement:

[0894] "Why can't you do something so simple? I can't believe it!"

[0895] The device picks up this speech with a microphone and sends it to the server.

[0896] The server uses a speech recognition API to generate the text "Why can't you do something so simple? I can't believe it!"

[0897] The server uses a sentiment analysis library to determine that this text is negative.

[0898] The server uses a generative artificial intelligence model to convert this into "You made a lot of mistakes today. Let's think together about how to solve the problem."

[0899] The server converts the converted text into voice data using a voice synthesis tool and sends it to the terminal.

[0900] The terminal receives the converted audio data and plays it back to the user.

[0901] Example 2: Converting a colleague's grumpy tone

[0902] Colleagues say:

[0903] "I'm tired. I'm starting to hate this job."

[0904] The device picks up this speech with a microphone and sends it to the server.

[0905] The server uses a speech recognition API to generate the text "I'm tired. I hate this job."

[0906] The server uses a sentiment analysis library to determine that this text is negative.

[0907] The server uses a generative artificial intelligence model to translate this into "I'm a little tired today, but I believe I can get through it."

[0908] The server converts the converted text into voice data using a voice synthesis tool and sends it to the terminal.

[0909] The terminal receives the converted audio data and plays it back to the user.

[0910] In this way, this system converts negative expressions into positive ones, thereby improving the user's productivity.

[0911] The processing flow will be explained below.

[0912] Step 1:

[0913] The device picks up ambient sound through a microphone, and the sound data is recorded in real time and sent to a server via streaming.

[0914] Step 2:

[0915] The server receives the voice data sent from the device. This voice data is converted into text using Google Speech Recognition API or other voice recognition technology. The converted text data is temporarily stored in an internal buffer.

[0916] Step 3:

[0917] The server inputs the converted text data into a sentiment analysis library, which analyzes the text data for negative expressions and detects negative or irritable speech patterns. The results are saved as negative flags.

[0918] Step 4:

[0919] The server checks the analysis results from the sentiment analysis library and extracts negative phrases, which are then converted into positive ones using a generative AI model.

[0920] Step 5:

[0921] The server inputs the generated positive expressions into a text-to-speech (TTS) tool to generate voice data, which is then temporarily stored.

[0922] Step 6:

[0923] The server then transmits the generated audio data to the terminal in streaming format, delivering the audio to the user in real time.

[0924] Step 7:

[0925] The device decodes the voice data received from the server and plays it through a speaker or earphones, allowing the user to receive voice information converted into positive expressions in real time.

[0926] Example 1

[0927] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0928] In conventional systems, negative comments from those around you could be directly conveyed to the user, causing psychological stress and reducing productivity. Furthermore, the process of converting negative comments into positive ones was complicated, making it difficult to respond in real time. This made it difficult for users to work in a comfortable environment.

[0929] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0930] In this invention, the server includes means for converting received speech into text data, means for analyzing the text data to detect negative expressions, and means for converting the detected negative expressions into positive expressions using a generative artificial intelligence model. This makes it possible to acquire surrounding speech, convert the negative expressions into positive expressions in real time, and provide them to the user.

[0931] "Audio capture device" refers to a device for capturing ambient sound. Examples include a microphone.

[0932] A "speech recognition service API" refers to an application programming interface for converting voice data into text, often provided as a cloud-based service.

[0933] "Server" refers to a computer system with data processing capabilities that receives, converts, and analyzes voice data.

[0934] A "generative artificial intelligence model" refers to an algorithm that generates new text based on a given prompt, based on deep learning techniques, such as GPT-3.

[0935] "Real-time" refers to data processing and communication occurring instantaneously with minimal delay.

[0936] "Text data" refers to the result of converting audio data into text information, which serves as the basis for analysis and other processing.

[0937] "Negative expressions" refer to negative content or expressions identified through sentiment analysis, which may cause psychological stress to users.

[0938] "Positive expressions" refer to positive content or expressions that replace negative expressions and are used to maintain the user's psychological stability.

[0939] "Audio data" refers to the digital representation of sound that is used to be heard by a user through a playback device.

[0940] This invention is a system that converts negative comments from people around you into positive expressions and provides positive voice information to the user. Each element and its operation will be described in detail below.

[0941] 1. System Configuration

[0942] (a) Terminal

[0943] The device is equipped with a sound capture device (microphone) for capturing surrounding sounds, and hardware and software for processing and playing back the sound data. The device transmits the collected sound data to a server in real time and plays back the sound data from the server in real time.

[0944] (b) Server

[0945] The server has the following functions:

[0946] Speech recognition function: Converts voice data into text. Specifically, it uses a speech recognition service API (e.g., Google Cloud Speech-to-Text).

[0947] Sentiment analysis function: Analyzes text data and detects negative expressions. Specifically, it uses a sentiment analysis library (e.g., NLTK or Google Cloud Natural Language API).

[0948] Conversion function: Converts detected negative expressions into positive expressions using a generative artificial intelligence model (e.g., GPT-3).

[0949] Speech synthesis function: Positive text data is converted back into voice data, specifically using a speech synthesis tool (e.g., Google Cloud Text-to-Speech).

[0950] 2. Operation overview

[0951] The device captures ambient sounds and sends the audio data in streaming format to the server. The server then converts the received audio data into text data using a speech recognition API. The server then analyzes the text data using a sentiment analysis library to identify negative expressions. The identified negative expressions are converted into positive expressions using a generative artificial intelligence model. The converted positive text data is then reconverted into audio data using a speech synthesis tool and sent to the device. Finally, the device plays the transmitted positive audio data to the user.

[0952] 3. Specific Examples

[0953] Example 1: Converting an angry boss's words

[0954] Boss' statement:

[0955] "Why can't you do something so simple? I can't believe it!"

[0956] 1. The device picks up this speech with a microphone and sends it to the server.

[0957] 2. The server uses a speech recognition API to generate the text "Why can't you do something so simple? I can't believe it!"

[0958] 3. The server uses a sentiment analysis library to determine that the text is negative.

[0959] 4. The server uses a generative artificial intelligence model to convert this into "You made a lot of mistakes today. Let's think together about how to solve the problem."

[0960] 5. The server converts the converted text into voice data using a speech synthesis tool and sends it to the device.

[0961] 6. The device receives the converted audio data and plays it back to the user.

[0962] Example 2: Converting a colleague's grumpy tone

[0963] Colleagues say:

[0964] "I'm tired. I'm starting to hate this job."

[0965] 1. The device picks up this speech with a microphone and sends it to the server.

[0966] 2. The server uses a speech recognition API to generate the text "I'm tired. I hate this job."

[0967] 3. The server uses a sentiment analysis library to determine that the text is negative.

[0968] 4. The server uses a generative artificial intelligence model to convert this to "I'm a little tired today, but I believe I can get it done."

[0969] 5. The server converts the converted text into voice data using a speech synthesis tool and sends it to the device.

[0970] 6. The device receives the converted audio data and plays it back to the user.

[0971] Prompt Sentence Examples

[0972] Change the following negative text into a positive one: "I'm exhausted. I hate this job."

[0973] In this way, this system converts negative expressions from the surrounding environment into positive ones, improving the user's psychological stability and productivity.

[0974] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0975] Step 1:

[0976] The device captures the surrounding sounds using an audio capture device, specifically, a microphone collects audio data and transmits it to a server in real time in a streaming format.

[0977] Input: Ambient audio

[0978] Output: Streaming audio data

[0979] Step 2:

[0980] The server receives the voice data sent from the device. It sends the received voice data to a voice recognition service API and converts the voice data into text data. Specifically, it uses a voice recognition API (e.g., Google Cloud Speech-to-Text) to convert the voice data into text format.

[0981] Input: Streaming audio data

[0982] Output: Text data

[0983] Step 3:

[0984] The server analyzes the text data using a sentiment analysis library to detect negative expressions. Specifically, it uses a sentiment analysis library (e.g., NLTK or Google Cloud Natural Language API) to calculate the sentiment score of the text data. It then identifies any part of the text where the negative score exceeds a certain threshold as a negative expression.

[0985] Input: Text data

[0986] Output: Text data containing negative expressions

[0987] Step 4:

[0988] The server converts the identified negative expressions into positive expressions using a generative artificial intelligence model. Specifically, it uses a generative artificial intelligence model (e.g., GPT-3) and provides a specific prompt. An example of the prompt is "Please convert the following negative text into a positive expression:"

[0989] Input: Text data containing negative expressions

[0990] Output: Text data containing positive expressions

[0991] Step 5:

[0992] The server converts the text data containing positive expressions into voice data using a voice synthesis tool (e.g., Google Cloud Text-to-Speech).

[0993] Input: Text data containing positive expressions

[0994] Output: Positive voice data

[0995] Step 6:

[0996] The server transmits the generated positive voice data to the terminal, and the terminal plays the received voice data to the user through a speaker, allowing the user to receive positive voice information in real time.

[0997] Input: Positive voice data

[0998] Output: Providing positive audio information to the user

[0999] (Application example 1)

[1000] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1001] In the work environment within a factory, workers are often exposed to negative comments and inappropriate language, which can reduce work efficiency and even lead to deterioration of interpersonal relationships. In such an environment, workers experience increased psychological stress and the risk of work errors and accidents increases. The objective of the present invention is to solve these problems and provide a means to provide an environment in which workers can work with a positive attitude.

[1002] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1003] In this invention, the server includes means for converting surrounding voices into text data using a voice recognition function, means for analyzing and detecting negative expressions using a sentiment analysis function, and means for converting the detected negative expressions into positive expressions using a generative artificial intelligence model. This makes it possible to provide negative comments as positive voice information in real time to workers wearing head-mounted displays in factories.

[1004] "Ambient audio" refers to any audio signals emanating within the environment in which the user is present.

[1005] "Capturing means" refers to a device or software that collects audio data and provides it in a form that can be used for further processing.

[1006] "Means for converting into text data" refers to a process or system that analyzes collected voice data and converts it into corresponding text information.

[1007] "Means for analyzing and detecting negative expressions" refers to an algorithm or system for detecting and identifying expressions with negative or downside meanings contained in text data.

[1008] "Means for converting into positive expressions" refers to a generative artificial intelligence model that has the ability to automatically change detected negative expressions into positive or forward-looking expressions.

[1009] "Means for converting into audio data" refers to a speech synthesis system or tool that can convert text data into audio format.

[1010] "Means for playing" refers to a device or system that outputs the converted audio data in an audible form to a user.

[1011] "Head-mounted display" refers to a visual and audio output device that is worn on a user's head and can provide audio information in real time.

[1012] "Worker" refers to a person who performs work in a specific environment such as a factory or workshop.

[1013] "Real-time presentation means" refers to a system or process for immediately conveying captured and processed audio data to a user.

[1014] This invention is a system used in factories in which workers wearing head-mounted displays convert negative comments made by those around them into positive expressions and provide them as voice in real time.

[1015] System configuration and specific operation

[1016] 1. System Configuration

[1017] The system mainly consists of the following components:

[1018] Terminal: A device for workers wearing a head-mounted display (HMD). It has a built-in microphone and speaker, picks up surrounding sounds, and plays back the audio data.

[1019] Server: Converts voice data to text, performs sentiment analysis and expression conversion, and converts the conversion results back into voice data.

[1020] 2. Hardware and Software Used

[1021] Head-mounted display (HMD): A device worn by the worker that inputs and outputs voice.

[1022] Microphone: A sound capture device built into the HMD.

[1023] Server: A powerful computer system.

[1024] Speech Recognition API: Uses the speech_recognition library.

[1025] Sentiment Analysis Library: Uses the sentiment analysis functionality from the transformers library.

[1026] Generative AI models: Generative artificial intelligence models such as GPT-3.

[1027] Speech synthesis tool: Uses the gTTS (Google Text-to-Speech) library.

[1028] 3. Operational Details

[1029] The system operates by having the server receive the voice data sent from the terminal and perform the following processing. First, it converts the voice data into text data using a voice recognition API and analyzes the text data using a sentiment analysis library. Next, it uses a generative artificial intelligence model to convert any negative expressions detected into positive ones. After that, it converts the converted positive text data into voice data using a voice synthesis tool and sends the result to the terminal. The terminal then plays this positive voice data in real time and provides it to the worker.

[1030] 4. Usage example

[1031] Example 1:

[1032] If a worker in a factory hears the following statement:

[1033] "Too many mistakes today!"

[1034] The microphone picks up this speech and sends it to the server, which converts it into text using a speech recognition API and uses a sentiment analysis library to identify negative expressions. The generative AI model is then given the following prompt:

[1035] text

[1036] Turn the following statement into a positive: I made too many mistakes today!

[1037] The generative AI model transforms this negative expression into:

[1038] text

[1039] I made a few mistakes today, but I'm sure I can improve next time.

[1040] This positive text data is converted into voice data using a speech synthesis tool and sent to the terminal. The head-mounted display plays the converted voice data and provides it to the worker.

[1041] This allows workers to obtain positive information in real time, reducing psychological stress and improving work efficiency.

[1042] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1043] Step 1:

[1044] The device captures the surrounding audio through a microphone. This captured audio data is sent to the server in streaming format. The input data is the surrounding audio, and the output data is the audio stream sent to the server.

[1045] Step 2:

[1046] The server converts the received voice data into text data using a voice recognition API. Here, the voice data is converted into text information. The input data is an audio stream, and the output data is text data.

[1047] Step 3:

[1048] The server analyzes the converted text data using a sentiment analysis library to detect negative expressions. The input data is text data, and the output data is the analysis result of negative expressions. Specifically, it determines whether the words and expressions contained in the text are negative.

[1049] Step 4:

[1050] The server uses a generative AI model to convert the detected negative expressions into positive expressions. At this time, it passes a prompt sentence to the generative AI model to instruct the conversion. The input data is negative text data, and the output data is the text data converted into positive. An example of a specific prompt sentence is as follows:

[1051] text

[1052] Turn the following statement into a positive: I made too many mistakes today!

[1053] Step 5:

[1054] The server converts the text data converted into positive data into voice data using a voice synthesis tool. The input data is the positive text data, and the output data is voice data.

[1055] Step 6:

[1056] The server sends the generated audio data to the terminal, where the input data is the audio data and the output data is the audio stream sent to the terminal.

[1057] Step 7:

[1058] The device plays back the received positive voice data and provides it to the user. The input data is the voice data, and the output data is the voice that the user hears. Specifically, the voice is played back through the speaker of the head-mounted display.

[1059] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1060] This invention achieves more advanced emotion management by combining a system that captures the sounds around the user, converts them into text, converts negative expressions into positive ones, and provides them to the user again as audio, with an emotion engine that recognizes the user's emotions. Each element and its operation will be explained in detail below.

[1061] 1. System Configuration

[1062] (a) Terminal

[1063] The device is equipped with a microphone, a speaker or earphones, and an emotion engine (camera, microphone, biometric sensors, etc.) to recognize the user's emotions. The device captures surrounding sounds and sends them to the server. It also plays back the audio data sent from the server.

[1064] (b) Server

[1065] The server has the following functions:

[1066] Speech recognition function: Converts voice data into text.

[1067] Sentiment analysis function: Analyzes text data and detects negative expressions.

[1068] Transformation function: Transform negative expressions into positive expressions.

[1069] Speech synthesis function: Converts positive text data back into speech data.

[1070] Emotion recognition function: Recognizes the user's emotions and reflects the results in other analysis and conversion processes.

[1071] 2. Program Processing

[1072] (a) Audio capture and transmission

[1073] The device picks up ambient sounds through a microphone and transmits the audio data to the server in streaming format. At the same time, the device's emotion engine recognizes the user's emotions (voice tone, facial expressions, biometric signals, etc.) and transmits the data to the server.

[1074] (b) Converting voice data into text

[1075] The server analyzes the received voice data and converts it into text data using a speech recognition API, which is used as input data for sentiment analysis.

[1076] (c) Detection of negative expressions

[1077] The server analyzes the text data using a sentiment analysis library to detect negative expressions. It also takes into account the user's emotional data from the emotion engine to improve the accuracy of the analysis. For example, if the user is feeling stressed, the analysis results will be set to be more sensitive.

[1078] (d) Transforming negative expressions into positive ones

[1079] The server uses a generative artificial intelligence model to convert negative expressions into positive ones, and uses the user's emotional data to make more appropriate conversions. For example, if the user is relaxed, the tone will be adjusted to be calmer.

[1080] (e) Vocalization and delivery of positive expressions

[1081] The server converts the converted positive text data into voice data using a speech synthesis tool, which is then sent to the device and played back in real time.

[1082] (f) Audio playback

[1083] The device receives the audio data sent from the server and plays it back to the user through a speaker or earphone, protecting the user from the negative audio environment around them and allowing them to receive only positive information.

[1084] 3. Specific Examples

[1085] Example 1: Converting an angry boss's words

[1086] Boss' statement:

[1087] "Why can't you do something so simple? I can't believe it!"

[1088] The device picks up this speech with a microphone and sends the audio data to the server. At the same time, the device captures the user's facial expression with a camera and sends the emotional data to the server.

[1089] The server uses a speech recognition API to generate the text "Why can't you do something so simple? I can't believe it!"

[1090] The server analyzes this text using a sentiment analysis library and determines that it is a negative comment. Data from the emotion engine determines that the user is stressed.

[1091] The server uses a generative artificial intelligence model to translate this into "You've made a lot of mistakes today. Let's work together to figure out how to solve the problem." The tone is adjusted to be gentler depending on the user's stress level.

[1092] The server converts the converted text into voice data using a voice synthesis tool and sends it to the terminal.

[1093] The terminal receives the converted audio data and plays it back to the user.

[1094] Example 2: Converting a colleague's grumpy tone

[1095] Colleagues say:

[1096] "I'm tired. I'm starting to hate this job."

[1097] The device picks up this speech with a microphone and sends the voice data to the server, where the emotion engine analyzes the user's voice tone and sends the data to the server.

[1098] The server uses a speech recognition API to generate the text "I'm tired. I hate this job."

[1099] The server analyzes this text using a sentiment analysis library and determines that it is a negative comment. Data from the sentiment engine determines that the user is relaxed.

[1100] The server uses a generative artificial intelligence model to translate this into "I'm a little tired today, but I believe I can get through it," adjusting the tone to match the user's state of relaxation.

[1101] The server converts the converted text into voice data using a voice synthesis tool and sends it to the terminal.

[1102] The terminal receives the converted audio data and plays it back to the user.

[1103] In this way, this system converts negative expressions around the user into positive ones, and by providing optimal voice information taking into account the user's emotional state, it is possible to improve the user's mental health and productivity.

[1104] The processing flow will be explained below.

[1105] Step 1:

[1106] The device captures surrounding sounds through a microphone. At the same time, it uses a camera and biometric sensors to collect the user's facial expressions and biometric signals (e.g., heart rate) to recognize the user's emotions. The collected voice data and emotion data are sent to a server in real time.

[1107] Step 2:

[1108] The server receives the voice data sent from the device, converts it into text data using the Google speech recognition API or other voice recognition technology, and simultaneously receives and analyzes the user's emotion data sent from the device.

[1109] Step 3:

[1110] The server inputs the converted text data into a sentiment analysis library for analysis. The sentiment analysis library detects negative expressions in the text data and identifies negative or unhappy speech. At the same time, it determines the user's current emotional state based on the user's emotional data.

[1111] Step 4:

[1112] The server uses a generative AI model to convert negative expressions in the text into positive ones, taking into account the user's emotional data. For example, if the user is feeling stressed, the text will be converted into a softer, more positive one.

[1113] Step 5:

[1114] The server inputs the generated positive expressions into a text-to-speech (TTS) tool to generate voice data, which is adjusted in tone and speed depending on the user's emotional state.

[1115] Step 6:

[1116] The server then transmits the generated audio data to the terminal in streaming format, delivering the audio to the user in real time.

[1117] Step 7:

[1118] The device decodes the voice data received from the server and plays it through a speaker or earphones, allowing the user to receive voice information converted into positive expressions in real time.

[1119] Specific examples

[1120] Example 1: Converting an angry boss's words

[1121] Step 1:

[1122] The device picks up the boss's statement, "Why can't you do something so simple? I can't believe it!", through a microphone. At the same time, the device captures the user's facial expression with a camera and also obtains emotional data (facial expression, heart rate). This data is then sent to the server.

[1123] Step 2:

[1124] The server receives the voice data and uses a speech recognition API to generate the text "Why can't you do something so simple? I can't believe it!" At the same time, it determines from the emotion data that the user is feeling stressed.

[1125] Step 3:

[1126] The server uses a sentiment analysis library to determine that the text is a negative statement.

[1127] Step 4:

[1128] The server uses a generative artificial intelligence model to translate the text into something like, "You've made a lot of mistakes today. Let's work together to figure out how to solve the problem." The tone is adjusted to calm the user based on their stress level.

[1129] Step 5:

[1130] The server inputs positively-phrased text into a speech synthesis tool to generate speech data, the tone and speed of which are adjusted to help the user relax.

[1131] Step 6:

[1132] The server transmits this audio data to the terminal.

[1133] Step 7:

[1134] The device then plays back the received voice data and delivers it to the user, who can hear a positive, calm voice saying in real time, "You've made a lot of mistakes today. Let's think together about how to solve the problem."

[1135] Example 2: Converting a colleague's grumpy tone

[1136] Step 1:

[1137] The device uses a microphone to capture a colleague's statement, "I'm tired. I'm starting to hate this job." The emotion engine analyzes the user's tone of voice and sends the data to the server.

[1138] Step 2:

[1139] The server receives the voice data and uses a speech recognition API to generate the text "I'm tired. I hate this job." At the same time, it checks the emotion engine to see if the user is relaxed.

[1140] Step 3:

[1141] The server uses a sentiment analysis library to determine that this text is negative.

[1142] Step 4:

[1143] The server uses a generative artificial intelligence model to translate the text into "I'm a little tired today, but I believe I can get through it," adjusting the tone and speed to match the user's level of relaxation.

[1144] Step 5:

[1145] The server inputs the text of the positive expression into a speech synthesis tool to generate speech data.

[1146] Step 6:

[1147] The server transmits this audio data to the terminal.

[1148] Step 7:

[1149] The device plays back the received voice data and delivers it to the user, who can hear a positive, calm tone in real time: "I'm a little tired today, but I believe I can get through it."

[1150] In this way, this system converts negative expressions from the surrounding environment into positive ones, and by providing optimal voice information taking into account the user's emotional state, it is possible to improve the user's mental health and productivity.

[1151] Example 2

[1152] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1153] In recent years, many people have become aware of the psychological impact of negative expressions and stressful environments at work and at home. However, there are no systems that can convert these negative expressions into positive expressions in real time and take the user's emotional state into account. Therefore, there is a need to develop a system that allows users to receive positive information without being exposed to negative environments.

[1154] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for acquiring ambient sound, means for converting the acquired sound into text data, means for analyzing and detecting negative expressions from the text data, means for acquiring user emotion data, means for converting the detected negative expressions into positive expressions, means for converting the converted positive expressions into voice data, and means for playing the voice data. This allows the user to receive voice information in a state where negative expressions have been converted into positive expressions.

[1155] "Means for acquiring ambient sound" refers to a device or method that uses a microphone or other sensors to collect environmental sounds and conversations.

[1156] "Means for converting acquired voice into text data" refers to devices or methods that convert voice signals into text information using technologies such as voice recognition service APIs.

[1157] A "means for analyzing and detecting negative expressions from text data" is an apparatus or method that uses natural language processing techniques to identify negative sentiments and expressions within text data.

[1158] "Means for acquiring user emotional data" refers to devices or methods that use cameras, microphones, biometric sensors, etc. to collect emotional information such as the user's facial expressions, voice tone, and heart rate.

[1159] The "means for converting detected negative expressions into positive expressions" refers to a device or method that uses a generative artificial intelligence model to modify negative parts of text data into semantically positive ones.

[1160] The "means for converting the converted positive expression into voice data" refers to a device or method for converting the positive text data into a voice signal using a voice synthesis tool that generates voice from text.

[1161] "Means for reproducing audio data" refers to a device or method for transmitting audio data to a user in real time through a speaker, earphones, etc.

[1162] This invention is a system that captures the user's surrounding voices, converts them into text, converts negative expressions into positive ones, and provides the text back to the user as voice. Furthermore, by combining it with an emotion engine that recognizes the user's emotions, it provides more advanced emotion management. Each element and its operation procedure are explained in detail below.

[1163] 1. System Configuration

[1164] (a) Terminal

[1165] The device is equipped with a microphone, a speaker or earphones, and an emotion engine (camera, microphone, biometric sensors, etc.) to recognize the user's emotions. The device captures the surrounding sounds and sends them to the server. It also receives the audio data sent from the server and plays it back to the user.

[1166] (b) Server

[1167] The server has the following functions:

[1168] Speech recognition function: Converts voice data into text. Specifically, it uses a speech recognition service API (e.g., a general cloud-based speech recognition API).

[1169] Sentiment analysis function: Analyzes text data and detects negative expressions. Uses a sentiment analysis library.

[1170] Conversion function: Uses a generative artificial intelligence model to convert negative expressions into positive ones.

[1171] Speech synthesis function: Use a speech synthesis tool to convert positive text data back into speech data.

[1172] Emotion recognition function: Recognizes the user's emotions and reflects that emotional data in other analysis and conversion processes.

[1173] 2. Explanation of specific operations

[1174] Audio capture and transmission

[1175] The device picks up ambient sounds through a microphone. This audio data is sent to the server in real time in streaming format. At the same time, the device's emotion engine recognizes the user's emotions and sends data such as facial expressions, voice tone, and biometric signals to the server.

[1176] Converting audio data to text

[1177] The server receives the voice data sent from the device and converts it into text using a voice recognition API.

[1178] Detecting negative expressions

[1179] The server analyzes text data using a sentiment analysis library to detect negative expressions, and also considers user sentiment data from an emotion engine to improve analysis accuracy.

[1180] Transforming negative expressions into positive ones

[1181] The server uses a generative artificial intelligence model to convert negative expressions into positive ones, taking into account the user's emotional data to make the conversion more appropriate.

[1182] Vocalization and delivery of positive expressions

[1183] The server reconverts the converted positive text data into voice data using a voice synthesis tool, and the voice data is sent to the terminal.

[1184] Playing audio

[1185] The terminal receives the audio data sent from the server and plays it back to the user through a speaker or earphones.

[1186] 3. Specific Examples

[1187] Example 1: Converting an angry boss's words

[1188] Boss says: "Why can't you do something so simple? I can't believe it!"

[1189] The device picks up this speech with a microphone and transmits the audio data to the server in real time. At the same time, the device captures the user's facial expression with a camera and transmits the data to the server.

[1190] The server uses a speech recognition API to generate text data of the utterance.

[1191] The server uses a sentiment analysis library to detect that the generated text is negative, and based on data from the emotion engine, recognizes that the user is feeling stressed.

[1192] The server uses a generative artificial intelligence model to translate the utterance into a positive expression, such as "You made a lot of mistakes today. Let's think together about how to solve the problem."

[1193] The server uses a speech synthesis tool to convert the text into voice data and send it to the terminal.

[1194] The device receives the converted audio and plays it back to the user through speakers or earphones.

[1195] Example 2: Converting a colleague's grumpy tone

[1196] A colleague says: "I'm tired. I'm going to hate this job."

[1197] The device picks up this speech with a microphone and sends the voice data to the server in real time. At the same time, the emotion engine analyzes the user's voice tone and sends the data to the server.

[1198] The server uses a speech recognition API to generate text data of the utterance.

[1199] The server uses a sentiment analysis library to detect that the generated text is negative, and based on data from the emotion engine, recognizes that the user is relaxed.

[1200] The server uses a generative artificial intelligence model to translate the utterance into a positive expression, generating the text, "I'm a little tired today, but I believe I can get through it."

[1201] The server uses a speech synthesis tool to convert the text into voice data and send it to the terminal.

[1202] The device receives the converted audio and plays it back to the user through speakers or earphones.

[1203] As described above, this system converts negative comments made by users around them into positive ones, and furthermore, by taking into account the user's emotional state and providing optimal voice information, it is possible to improve the user's mental health and productivity.

[1204] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1205] Step 1: Acquire audio

[1206] The device uses a microphone to capture surrounding sounds in real time. For example, it collects the user's voice while speaking and environmental sounds. The input is audio signals from the external environment, and the output is digital audio data. At the same time, the device's emotion engine (camera, microphone, biometric sensors, etc.) captures the user's facial expressions, voice tone, and biometric signals. The input is the user's emotional state, and the output is the user's emotional data.

[1207] Step 2: Sending audio data

[1208] The terminal transmits the acquired voice data and emotion data to the server in streaming format. The input here is the acquired voice data and the user's emotion data, and the output is data transmitted to the server via a network. Specifically, data transmission is performed using wireless communication.

[1209] Step 3: Converting audio data to text

[1210] The server receives the voice data sent from the device and converts it into text data using a speech recognition API (e.g., a common cloud-based speech recognition service). The input is digital voice data, and the output is text data. For example, the generated text might say, "Why can't you do something so simple?" This process requires network and computer resources.

[1211] Step 4: Detecting negative expressions

[1212] The server uses a sentiment analysis library to analyze the converted text data and detect negative expressions. It also uses the user's emotional data from the emotion engine for analysis. The input is the text data and the user's emotional data, and the output is the identification result of text containing negative expressions. Specifically, the text "Why can't you do something so simple?" is determined to be negative.

[1213] Step 5: Transform negative expressions into positive ones

[1214] The server uses a generative artificial intelligence model (e.g., a generative AI model) to convert the detected negative expressions into positive expressions. The input is text data containing negative expressions and the user's emotional data, and the output is text data converted into positive expressions. For example, "Why can't you do something so simple?" is converted into "You made quite a few mistakes today, but let's find a solution together." This process requires natural language processing technology and a generative AI model.

[1215] Step 6: Vocalize and deliver positive feedback

[1216] The server converts the converted positive text data into voice data using a voice synthesis tool (e.g., a voice synthesis engine). The input is the positive text data, and the output is voice data. For example, a voice saying, "You made quite a few mistakes today, but let's find a solution together" is generated. This voice data is then sent to the device.

[1217] Step 7: Playing Audio

[1218] The terminal receives the voice data sent from the server and plays it back to the user through a speaker or earphone. The input is the voice data sent from the server, and the output is the voice played back to the user. This process allows the user to receive positively converted voice information in real time.

[1219] (Application example 2)

[1220] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1221] In communication between employees and customers in physical stores, responding appropriately to negative comments or complaints from customers can be mentally stressful for employees and can reduce work efficiency. Furthermore, if an employee's emotional state is unstable, their response to customers will be poor. To solve this problem, a system is needed that can convert surrounding negative vocal expressions into positive ones and provide optimal vocal information that takes into account the employee's emotional state.

[1222] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1223] In this invention, the server includes means for converting voice into text data, means for analyzing and detecting negative expressions from the text data, means for converting the detected negative expressions into positive expressions, and means for utilizing an emotion engine that recognizes the user's emotions. This enables employees to receive negative comments from customers as voice that has been converted into positive ones, thereby reducing their mental burden and improving the quality of customer service.

[1224] The "means for acquiring ambient sound" is a function for acquiring sound data in the environment using a microphone, a sound acquisition device, or the like.

[1225] "Means for converting acquired voice data into text data" refers to the process of converting acquired voice data into character information (text data) using voice recognition technology or API.

[1226] "Means for analyzing and detecting negative expressions from text data" refers to a function that uses a sentiment analysis library and natural language processing technology to identify expressions that indicate negative emotions or dissatisfaction from input text data.

[1227] The "means for converting detected negative expressions into positive expressions" is a means for converting detected negative text into positive content using a generative artificial intelligence model.

[1228] "Means of utilizing an emotion engine that recognizes the user's emotions" refers to a technology that uses a camera or biometric sensor to monitor the user's facial expressions and biometric signals, and analyzes them to recognize the user's emotional state.

[1229] The "means for reproducing audio data" is a function for allowing the user to hear the generated audio data using an audio reproduction device such as a speaker or earphones.

[1230] This invention is a system designed to support store employees when dealing with customers. Specifically, the system captures surrounding voices, converts them into text data, converts negative expressions into positive ones, and provides the voice data to the store employees. Furthermore, by combining it with an emotion engine that recognizes the emotions of the user (employee), the system aims to achieve advanced emotion management by providing optimal voice information according to the employee's emotional state.

[1231] System Configuration

[1232] (a) Terminal

[1233] The device is equipped with smart glasses, a microphone, a speaker or earphones, and an emotion engine (camera, biometric sensors, etc.) that recognizes the user's emotions. The device captures the surrounding audio and sends it to the server. It also plays back the audio data sent from the server.

[1234] (b) Server

[1235] The server has the following functions:

[1236] Speech recognition function: Converts voice data into text (e.g., Google Speech-to-Text API).

[1237] Sentiment analysis functions: Analyze text data and detect negative expressions (e.g., NLTK or Transformers libraries).

[1238] Transformation function: Transforming negative expressions into positive expressions (e.g., generative AI models such as GPT-3).

[1239] Speech synthesis function: Converts positive text data back into speech data (e.g. gTTS).

[1240] Emotion recognition: Recognizes the user's emotions and incorporates the results into other analysis and transformation processes (e.g., camera and biometric sensors).

[1241] Program processing

[1242] The device uses a microphone to capture the conversation between the customer and employee and sends the audio data in streaming format to the server. At the same time, the device's emotion engine recognizes the user's emotions (voice tone, facial expressions, biometric signals, etc.) and sends this data to the server. The server analyzes the received audio data and converts it into text data using a speech recognition API. This data is used as input data for emotion analysis.

[1243] The server uses a sentiment analysis library to analyze text data and detect negative expressions. It also takes into account the user's emotional data from the emotion engine to improve analysis accuracy. For example, if the user is feeling stressed, the analysis results will be set more sensitively. The server then uses a generative artificial intelligence model to convert negative expressions into positive ones. The server uses the user's emotional data to perform more appropriate conversions. For example, if the user is relaxed, the tone will be adjusted to be gentler.

[1244] The server then converts the converted positive text data into audio data using a speech synthesis tool. This data is then sent to the device and played in real time. The device then receives the audio data sent from the server and plays it back to the user through a speaker or earphones. This protects the user from the negative audio environment around them and allows them to receive only positive information.

[1245] Specific examples

[1246] Example 1: When a customer makes a complaint

[1247] Customer says: "I can't use this product at all, what's going on?"

[1248] What employees ask: "Are you having trouble with this product? Is there anything I can help you with?"

[1249] Prompt Sentence Examples

[1250] "This product is completely unusable. What's going on?" Please change this to a positive expression.

[1251] By using this system, employees can reduce their mental burden and improve the quality of customer service.

[1252] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1253] Step 1:

[1254] The device captures the surrounding audio through a microphone, and this audio data is sent to the server in real time. The input is the surrounding audio data, and the output is sending the audio data to the server.

[1255] Step 2:

[1256] The device's emotion engine (including the camera and biometric sensors) acquires the user's emotional data (facial expressions, biometric signals, etc.). This data is also sent to the server at the same time as the voice data. The input is facial expressions and biometric signals, and the output is sending emotional data to the server.

[1257] Step 3:

[1258] The server converts the received voice data into text data using a speech recognition API. The input is the voice data, and the output is the corresponding text data.

[1259] Step 4:

[1260] The server analyzes the converted text data using a sentiment analysis library to detect negative expressions. The input is the text data, and the output is the result of whether or not it contains negative expressions.

[1261] Step 5:

[1262] The server uses a generative artificial intelligence model (e.g., GPT-3) to convert text containing negative expressions into positive ones, while also taking into account the user's emotional data to adjust the conversion process. The input is negative text data and emotional data, and the output is the text data converted into positive ones.

[1263] Step 6:

[1264] The server converts the positive text data into speech data using a speech synthesis tool (e.g., gTTS). The input is the positive text data, and the output is speech data.

[1265] Step 7:

[1266] The server sends the generated audio data to the terminal. The input is audio data, and the output is sending audio data to the terminal.

[1267] Step 8:

[1268] The device then plays the received audio data through a speaker or earphone. The input is audio data, and the output is the audio the user hears. This protects the user from the negative audio environment around them and allows them to receive only positive information.

[1269] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1270] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1271] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1272] [Fourth embodiment]

[1273] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1274] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1275] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1276] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1277] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1278] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1279] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1280] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1281] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1282] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1283] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1284] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1285] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1286] This invention is a system that converts negative comments from people around you into positive expressions and provides positive voice information to the user. Each element and its operation will be described in detail below.

[1287] 1. System Configuration

[1288] (a) Terminal

[1289] The device is equipped with a microphone to capture surrounding sounds, and hardware and software for processing and playing back the audio data. The device transmits the collected audio data to a server in real time and plays back the audio data from the server in real time.

[1290] (b) Server

[1291] The server has the following functions:

[1292] Speech recognition function: Converts voice data into text.

[1293] Sentiment analysis function: Analyzes text data and detects negative expressions.

[1294] Conversion function: Converts detected negative expressions into positive expressions.

[1295] Speech synthesis function: Converts positive text data back into speech data.

[1296] 2. Program Processing

[1297] (a) Audio capture and transmission

[1298] The device picks up ambient sound through a microphone, and the acquired sound data is sent to the server in streaming format.

[1299] (b) Converting voice data into text

[1300] The server receives the voice data sent from the device, converts it into text using a speech recognition API, and then prepares the text data for analysis.

[1301] (c) Detection of negative expressions

[1302] The server analyzes the text data using a sentiment analysis library, and based on the analysis results, negative expressions and unpleasant speech patterns are identified.

[1303] (d) Transforming negative expressions into positive ones

[1304] The server converts the detected negative expressions into positive expressions using a generative artificial intelligence model.

[1305] (e) Vocalization and delivery of positive expressions

[1306] The server converts the positively converted text data into voice data using a voice synthesis tool, and the voice data is sent to the terminal.

[1307] (f) Audio playback

[1308] The terminal receives the positive voice data sent from the server and plays it back to the user in real time, so that the user can obtain positive voice information.

[1309] 3. Specific Examples

[1310] Example 1: Converting an angry boss's words

[1311] Boss' statement:

[1312] "Why can't you do something so simple? I can't believe it!"

[1313] The device picks up this speech with a microphone and sends it to the server.

[1314] The server uses a speech recognition API to generate the text "Why can't you do something so simple? I can't believe it!"

[1315] The server uses a sentiment analysis library to determine that this text is negative.

[1316] The server uses a generative artificial intelligence model to convert this into "You made a lot of mistakes today. Let's think together about how to solve the problem."

[1317] The server converts the converted text into voice data using a voice synthesis tool and sends it to the terminal.

[1318] The terminal receives the converted audio data and plays it back to the user.

[1319] Example 2: Converting a colleague's grumpy tone

[1320] Colleagues say:

[1321] "I'm tired. I'm starting to hate this job."

[1322] The device picks up this speech with a microphone and sends it to the server.

[1323] The server uses a speech recognition API to generate the text "I'm tired. I hate this job."

[1324] The server uses a sentiment analysis library to determine that this text is negative.

[1325] The server uses a generative artificial intelligence model to translate this into "I'm a little tired today, but I believe I can get through it."

[1326] The server converts the converted text into voice data using a voice synthesis tool and sends it to the terminal.

[1327] The terminal receives the converted audio data and plays it back to the user.

[1328] In this way, this system converts negative expressions into positive ones, thereby improving the user's productivity.

[1329] The processing flow will be explained below.

[1330] Step 1:

[1331] The device picks up ambient sound through a microphone, and the sound data is recorded in real time and sent to a server via streaming.

[1332] Step 2:

[1333] The server receives the voice data sent from the device. This voice data is converted into text using the Google Speech Recognition API or other voice recognition technology. The converted text data is temporarily stored in an internal buffer.

[1334] Step 3:

[1335] The server inputs the converted text data into a sentiment analysis library, which analyzes the text data for negative expressions and detects negative or irritable speech patterns. The results are saved as negative flags.

[1336] Step 4:

[1337] The server checks the analysis results from the sentiment analysis library and extracts negative phrases, which are then converted into positive ones using a generative AI model.

[1338] Step 5:

[1339] The server inputs the generated positive expressions into a text-to-speech (TTS) tool to generate voice data, which is then temporarily stored.

[1340] Step 6:

[1341] The server then transmits the generated audio data to the terminal in streaming format, delivering the audio to the user in real time.

[1342] Step 7:

[1343] The device decodes the voice data received from the server and plays it through a speaker or earphones, allowing the user to receive voice information converted into positive expressions in real time.

[1344] Example 1

[1345] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1346] In conventional systems, negative comments from those around you could be directly conveyed to the user, causing psychological stress and reducing productivity. Furthermore, the process of converting negative comments into positive ones was complicated, making it difficult to respond in real time. This made it difficult for users to work in a comfortable environment.

[1347] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1348] In this invention, the server includes means for converting received speech into text data, means for analyzing the text data to detect negative expressions, and means for converting the detected negative expressions into positive expressions using a generative artificial intelligence model. This makes it possible to acquire surrounding speech, convert the negative expressions into positive expressions in real time, and provide them to the user.

[1349] "Audio capture device" refers to a device for capturing ambient sound. Examples include a microphone.

[1350] A "speech recognition service API" refers to an application programming interface for converting voice data into text, often provided as a cloud-based service.

[1351] "Server" refers to a computer system with data processing capabilities that receives, converts, and analyzes voice data.

[1352] A "generative artificial intelligence model" refers to an algorithm that generates new text based on a given prompt, based on deep learning techniques, such as GPT-3.

[1353] "Real-time" refers to data processing and communication occurring instantaneously with minimal delay.

[1354] "Text data" refers to the result of converting audio data into text information, which serves as the basis for analysis and other processing.

[1355] "Negative expressions" refer to negative content or expressions identified through sentiment analysis, which may cause psychological stress to users.

[1356] "Positive expressions" refer to positive content or expressions that replace negative expressions and are used to maintain the user's psychological stability.

[1357] "Audio data" refers to the digital representation of sound that is used to be heard by a user through a playback device.

[1358] This invention is a system that converts negative comments from people around you into positive expressions and provides positive voice information to the user. Each element and its operation will be described in detail below.

[1359] 1. System Configuration

[1360] (a) Terminal

[1361] The device is equipped with a sound capture device (microphone) for capturing surrounding sounds, and hardware and software for processing and playing back the sound data. The device transmits the collected sound data to a server in real time and plays back the sound data from the server in real time.

[1362] (b) Server

[1363] The server has the following functions:

[1364] Speech recognition function: Converts voice data into text. Specifically, it uses a speech recognition service API (e.g., Google Cloud Speech-to-Text).

[1365] Sentiment analysis function: Analyzes text data and detects negative expressions. Specifically, it uses a sentiment analysis library (e.g., NLTK or Google Cloud Natural Language API).

[1366] Conversion function: Converts detected negative expressions into positive expressions using a generative artificial intelligence model (e.g., GPT-3).

[1367] Speech synthesis function: Positive text data is converted back into voice data, specifically using a speech synthesis tool (e.g., Google Cloud Text-to-Speech).

[1368] 2. Operation overview

[1369] The device captures ambient sounds and sends the audio data in streaming format to the server. The server then converts the received audio data into text data using a speech recognition API. The server then analyzes the text data using a sentiment analysis library to identify negative expressions. The identified negative expressions are converted into positive expressions using a generative artificial intelligence model. The converted positive text data is then reconverted into audio data using a speech synthesis tool and sent to the device. Finally, the device plays the transmitted positive audio data to the user.

[1370] 3. Specific Examples

[1371] Example 1: Converting an angry boss's words

[1372] Boss' statement:

[1373] "Why can't you do something so simple? I can't believe it!"

[1374] 1. The device picks up this speech with a microphone and sends it to the server.

[1375] 2. The server uses a speech recognition API to generate the text "Why can't you do something so simple? I can't believe it!"

[1376] 3. The server uses a sentiment analysis library to determine that the text is negative.

[1377] 4. The server uses a generative artificial intelligence model to convert this into "You made a lot of mistakes today. Let's think together about how to solve the problem."

[1378] 5. The server converts the converted text into voice data using a speech synthesis tool and sends it to the device.

[1379] 6. The device receives the converted audio data and plays it back to the user.

[1380] Example 2: Converting a colleague's grumpy tone

[1381] Colleagues say:

[1382] "I'm tired. I'm starting to hate this job."

[1383] 1. The device picks up this speech with a microphone and sends it to the server.

[1384] 2. The server uses a speech recognition API to generate the text "I'm tired. I hate this job."

[1385] 3. The server uses a sentiment analysis library to determine that the text is negative.

[1386] 4. The server uses a generative artificial intelligence model to convert this to "I'm a little tired today, but I believe I can get it done."

[1387] 5. The server converts the converted text into voice data using a speech synthesis tool and sends it to the device.

[1388] 6. The device receives the converted audio data and plays it back to the user.

[1389] Prompt Sentence Examples

[1390] Change the following negative text into a positive one: "I'm exhausted. I hate this job."

[1391] In this way, this system converts negative expressions from the surrounding environment into positive ones, improving the user's psychological stability and productivity.

[1392] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1393] Step 1:

[1394] The device captures the surrounding sounds using an audio capture device, specifically, a microphone collects audio data and transmits it to a server in real time in a streaming format.

[1395] Input: Ambient audio

[1396] Output: Streaming audio data

[1397] Step 2:

[1398] The server receives the voice data sent from the device. It sends the received voice data to a voice recognition service API and converts the voice data into text data. Specifically, it uses a voice recognition API (e.g., Google Cloud Speech-to-Text) to convert the voice data into text format.

[1399] Input: Streaming audio data

[1400] Output: Text data

[1401] Step 3:

[1402] The server analyzes the text data using a sentiment analysis library to detect negative expressions. Specifically, it uses a sentiment analysis library (e.g., NLTK or Google Cloud Natural Language API) to calculate the sentiment score of the text data. It then identifies any part of the text where the negative score exceeds a certain threshold as a negative expression.

[1403] Input: Text data

[1404] Output: Text data containing negative expressions

[1405] Step 4:

[1406] The server converts the identified negative expressions into positive expressions using a generative artificial intelligence model. Specifically, it uses a generative artificial intelligence model (e.g., GPT-3) and provides a specific prompt. An example of the prompt is "Please convert the following negative text into a positive expression:"

[1407] Input: Text data containing negative expressions

[1408] Output: Text data containing positive expressions

[1409] Step 5:

[1410] The server converts the text data containing positive expressions into voice data using a voice synthesis tool (e.g., Google Cloud Text-to-Speech).

[1411] Input: Text data containing positive expressions

[1412] Output: Positive voice data

[1413] Step 6:

[1414] The server transmits the generated positive voice data to the terminal, and the terminal plays the received voice data to the user through a speaker, allowing the user to receive positive voice information in real time.

[1415] Input: Positive voice data

[1416] Output: Providing positive audio information to the user

[1417] (Application example 1)

[1418] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1419] In the work environment within a factory, workers are often exposed to negative comments and inappropriate language, which can reduce work efficiency and even lead to deterioration of interpersonal relationships. In such an environment, workers experience increased psychological stress and the risk of work errors and accidents increases. The objective of the present invention is to solve these problems and provide a means to provide an environment in which workers can work with a positive attitude.

[1420] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1421] In this invention, the server includes means for converting surrounding voices into text data using a voice recognition function, means for analyzing and detecting negative expressions using a sentiment analysis function, and means for converting the detected negative expressions into positive expressions using a generative artificial intelligence model. This makes it possible to provide negative comments as positive voice information in real time to workers wearing head-mounted displays in factories.

[1422] "Ambient audio" refers to any audio signals emanating within the environment in which the user is present.

[1423] "Capturing means" refers to a device or software that collects audio data and provides it in a form that can be used for further processing.

[1424] "Means for converting into text data" refers to a process or system that analyzes collected voice data and converts it into corresponding text information.

[1425] "Means for analyzing and detecting negative expressions" refers to an algorithm or system for detecting and identifying expressions with negative or downside meanings contained in text data.

[1426] "Means for converting into positive expressions" refers to a generative artificial intelligence model that has the ability to automatically change detected negative expressions into positive or forward-looking expressions.

[1427] "Means for converting into audio data" refers to a speech synthesis system or tool that can convert text data into audio format.

[1428] "Means for playing" refers to a device or system that outputs the converted audio data in an audible form to a user.

[1429] "Head-mounted display" refers to a visual and audio output device that is worn on a user's head and can provide audio information in real time.

[1430] "Worker" refers to a person who performs work in a specific environment such as a factory or workshop.

[1431] "Real-time presentation means" refers to a system or process for immediately conveying captured and processed audio data to a user.

[1432] This invention is a system used in factories in which workers wearing head-mounted displays convert negative comments made by those around them into positive expressions and provide them as voice in real time.

[1433] System configuration and specific operation

[1434] 1. System Configuration

[1435] The system mainly consists of the following components:

[1436] Terminal: A device for workers wearing a head-mounted display (HMD). It has a built-in microphone and speaker, picks up surrounding sounds, and plays back the audio data.

[1437] Server: Converts voice data to text, performs sentiment analysis and expression conversion, and converts the conversion results back into voice data.

[1438] 2. Hardware and Software Used

[1439] Head-mounted display (HMD): A device worn by the worker that inputs and outputs voice.

[1440] Microphone: A sound capture device built into the HMD.

[1441] Server: A powerful computer system.

[1442] Speech Recognition API: Uses the speech_recognition library.

[1443] Sentiment Analysis Library: Uses the sentiment analysis functionality from the transformers library.

[1444] Generative AI models: Generative artificial intelligence models such as GPT-3.

[1445] Speech synthesis tool: Uses the gTTS (Google Text-to-Speech) library.

[1446] 3. Operational Details

[1447] The system operates by having the server receive the voice data sent from the terminal and perform the following processing. First, it converts the voice data into text data using a voice recognition API and analyzes the text data using a sentiment analysis library. Next, it uses a generative artificial intelligence model to convert any negative expressions detected into positive ones. After that, it converts the converted positive text data into voice data using a voice synthesis tool and sends the result to the terminal. The terminal then plays this positive voice data in real time and provides it to the worker.

[1448] 4. Usage example

[1449] Example 1:

[1450] If a worker in a factory hears the following statement:

[1451] "Too many mistakes today!"

[1452] The microphone picks up this speech and sends it to the server, which converts it into text using a speech recognition API and uses a sentiment analysis library to identify negative expressions. The generative AI model is then given the following prompt:

[1453] text

[1454] Turn the following statement into a positive: I made too many mistakes today!

[1455] The generative AI model transforms this negative expression into:

[1456] text

[1457] I made a few mistakes today, but I'm sure I can improve next time.

[1458] This positive text data is converted into voice data using a speech synthesis tool and sent to the terminal. The head-mounted display plays the converted voice data and provides it to the worker.

[1459] This allows workers to obtain positive information in real time, reducing psychological stress and improving work efficiency.

[1460] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1461] Step 1:

[1462] The device captures the surrounding audio through a microphone. This captured audio data is sent to the server in streaming format. The input data is the surrounding audio, and the output data is the audio stream sent to the server.

[1463] Step 2:

[1464] The server converts the received voice data into text data using a voice recognition API. Here, the voice data is converted into text information. The input data is an audio stream, and the output data is text data.

[1465] Step 3:

[1466] The server analyzes the converted text data using a sentiment analysis library to detect negative expressions. The input data is text data, and the output data is the analysis result of negative expressions. Specifically, it determines whether the words and expressions contained in the text are negative.

[1467] Step 4:

[1468] The server uses a generative AI model to convert the detected negative expressions into positive expressions. At this time, it passes a prompt sentence to the generative AI model to instruct the conversion. The input data is negative text data, and the output data is the text data converted into positive. An example of a specific prompt sentence is as follows:

[1469] text

[1470] Turn the following statement into a positive: I made too many mistakes today!

[1471] Step 5:

[1472] The server converts the text data converted into positive data into voice data using a voice synthesis tool. The input data is the positive text data, and the output data is voice data.

[1473] Step 6:

[1474] The server sends the generated audio data to the terminal, where the input data is the audio data and the output data is the audio stream sent to the terminal.

[1475] Step 7:

[1476] The device plays back the received positive voice data and provides it to the user. The input data is the voice data, and the output data is the voice that the user hears. Specifically, the voice is played back through the speaker of the head-mounted display.

[1477] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1478] This invention achieves more advanced emotion management by combining a system that captures the sounds around the user, converts them into text, converts negative expressions into positive ones, and provides them to the user again as audio, with an emotion engine that recognizes the user's emotions. Each element and its operation will be explained in detail below.

[1479] 1. System Configuration

[1480] (a) Terminal

[1481] The device is equipped with a microphone, a speaker or earphones, and an emotion engine (camera, microphone, biometric sensors, etc.) to recognize the user's emotions. The device captures surrounding sounds and sends them to the server. It also plays back the audio data sent from the server.

[1482] (b) Server

[1483] The server has the following functions:

[1484] Speech recognition function: Converts voice data into text.

[1485] Sentiment analysis function: Analyzes text data and detects negative expressions.

[1486] Transformation function: Transform negative expressions into positive expressions.

[1487] Speech synthesis function: Converts positive text data back into speech data.

[1488] Emotion recognition function: Recognizes the user's emotions and reflects the results in other analysis and conversion processes.

[1489] 2. Program Processing

[1490] (a) Audio capture and transmission

[1491] The device picks up ambient sounds through a microphone and transmits the audio data to the server in streaming format. At the same time, the device's emotion engine recognizes the user's emotions (voice tone, facial expressions, biometric signals, etc.) and transmits the data to the server.

[1492] (b) Converting voice data into text

[1493] The server analyzes the received voice data and converts it into text data using a speech recognition API, which is used as input data for sentiment analysis.

[1494] (c) Detection of negative expressions

[1495] The server analyzes the text data using a sentiment analysis library to detect negative expressions. It also takes into account the user's emotional data from the emotion engine to improve the accuracy of the analysis. For example, if the user is feeling stressed, the analysis results will be set to be more sensitive.

[1496] (d) Transforming negative expressions into positive ones

[1497] The server uses a generative artificial intelligence model to convert negative expressions into positive ones, and uses the user's emotional data to make more appropriate conversions. For example, if the user is relaxed, the tone will be adjusted to be calmer.

[1498] (e) Vocalization and delivery of positive expressions

[1499] The server converts the converted positive text data into voice data using a speech synthesis tool, which is then sent to the device and played back in real time.

[1500] (f) Audio playback

[1501] The device receives the audio data sent from the server and plays it back to the user through a speaker or earphone, protecting the user from the negative audio environment around them and allowing them to receive only positive information.

[1502] 3. Specific Examples

[1503] Example 1: Converting an angry boss's words

[1504] Boss' statement:

[1505] "Why can't you do something so simple? I can't believe it!"

[1506] The device picks up this speech with a microphone and sends the audio data to the server. At the same time, the device captures the user's facial expression with a camera and sends the emotional data to the server.

[1507] The server uses a speech recognition API to generate the text "Why can't you do something so simple? I can't believe it!"

[1508] The server analyzes this text using a sentiment analysis library and determines that it is a negative comment. Data from the emotion engine determines that the user is stressed.

[1509] The server uses a generative artificial intelligence model to translate this into "You've made a lot of mistakes today. Let's work together to figure out how to solve the problem." The tone is adjusted to be gentler depending on the user's stress level.

[1510] The server converts the converted text into voice data using a voice synthesis tool and sends it to the terminal.

[1511] The terminal receives the converted audio data and plays it back to the user.

[1512] Example 2: Converting a colleague's grumpy tone

[1513] Colleagues say:

[1514] "I'm tired. I'm starting to hate this job."

[1515] The device picks up this speech with a microphone and sends the voice data to the server, where the emotion engine analyzes the user's voice tone and sends the data to the server.

[1516] The server uses a speech recognition API to generate the text "I'm tired. I hate this job."

[1517] The server analyzes this text using a sentiment analysis library and determines that it is a negative comment. Data from the sentiment engine determines that the user is relaxed.

[1518] The server uses a generative artificial intelligence model to translate this into "I'm a little tired today, but I believe I can get through it," adjusting the tone to match the user's state of relaxation.

[1519] The server converts the converted text into voice data using a voice synthesis tool and sends it to the terminal.

[1520] The terminal receives the converted audio data and plays it back to the user.

[1521] In this way, this system converts negative expressions around the user into positive ones, and by providing optimal voice information taking into account the user's emotional state, it is possible to improve the user's mental health and productivity.

[1522] The processing flow will be explained below.

[1523] Step 1:

[1524] The device captures surrounding sounds through a microphone. At the same time, it uses a camera and biometric sensors to collect the user's facial expressions and biometric signals (e.g., heart rate) to recognize the user's emotions. The collected voice data and emotion data are sent to a server in real time.

[1525] Step 2:

[1526] The server receives the voice data sent from the device, converts it into text data using the Google speech recognition API or other voice recognition technology, and simultaneously receives and analyzes the user's emotion data sent from the device.

[1527] Step 3:

[1528] The server inputs the converted text data into a sentiment analysis library for analysis. The sentiment analysis library detects negative expressions in the text data and identifies negative or unhappy speech. At the same time, it determines the user's current emotional state based on the user's emotional data.

[1529] Step 4:

[1530] The server uses a generative AI model to convert negative expressions in the text into positive ones, taking into account the user's emotional data. For example, if the user is feeling stressed, the text will be converted into a softer, more positive one.

[1531] Step 5:

[1532] The server inputs the generated positive expressions into a text-to-speech (TTS) tool to generate voice data, which is adjusted in tone and speed depending on the user's emotional state.

[1533] Step 6:

[1534] The server then transmits the generated audio data to the terminal in streaming format, delivering the audio to the user in real time.

[1535] Step 7:

[1536] The device decodes the voice data received from the server and plays it through a speaker or earphones, allowing the user to receive voice information converted into positive expressions in real time.

[1537] Specific examples

[1538] Example 1: Converting an angry boss's words

[1539] Step 1:

[1540] The device picks up the boss's statement, "Why can't you do something so simple? I can't believe it!", through a microphone. At the same time, the device captures the user's facial expression with a camera and also obtains emotional data (facial expression, heart rate). This data is then sent to the server.

[1541] Step 2:

[1542] The server receives the voice data and uses a speech recognition API to generate the text "Why can't you do something so simple? I can't believe it!" At the same time, it determines from the emotion data that the user is feeling stressed.

[1543] Step 3:

[1544] The server uses a sentiment analysis library to determine that the text is a negative statement.

[1545] Step 4:

[1546] The server uses a generative artificial intelligence model to translate the text into something like, "You've made a lot of mistakes today. Let's work together to figure out how to solve the problem." The tone is adjusted to calm the user based on their stress level.

[1547] Step 5:

[1548] The server inputs positively-phrased text into a speech synthesis tool to generate speech data, the tone and speed of which are adjusted to help the user relax.

[1549] Step 6:

[1550] The server transmits this audio data to the terminal.

[1551] Step 7:

[1552] The device then plays back the received voice data and delivers it to the user, who can hear a positive, calm voice saying in real time, "You've made a lot of mistakes today. Let's think together about how to solve the problem."

[1553] Example 2: Converting a colleague's grumpy tone

[1554] Step 1:

[1555] The device uses a microphone to capture a colleague's statement, "I'm tired. I'm starting to hate this job." The emotion engine analyzes the user's tone of voice and sends the data to the server.

[1556] Step 2:

[1557] The server receives the voice data and uses a speech recognition API to generate the text "I'm tired. I hate this job." At the same time, it checks the emotion engine to see if the user is relaxed.

[1558] Step 3:

[1559] The server uses a sentiment analysis library to determine that this text is negative.

[1560] Step 4:

[1561] The server uses a generative artificial intelligence model to translate the text into "I'm a little tired today, but I believe I can get through it," adjusting the tone and speed to match the user's level of relaxation.

[1562] Step 5:

[1563] The server inputs the text of the positive expression into a speech synthesis tool to generate speech data.

[1564] Step 6:

[1565] The server transmits this audio data to the terminal.

[1566] Step 7:

[1567] The device plays back the received voice data and delivers it to the user, who can hear a positive, calm tone in real time: "I'm a little tired today, but I believe I can get through it."

[1568] In this way, this system converts negative expressions from the surrounding environment into positive ones, and by providing optimal voice information taking into account the user's emotional state, it is possible to improve the user's mental health and productivity.

[1569] Example 2

[1570] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1571] In recent years, many people have become aware of the psychological impact of negative expressions and stressful environments at work and at home. However, there are no systems that can convert these negative expressions into positive expressions in real time and take the user's emotional state into account. Therefore, there is a need to develop a system that allows users to receive positive information without being exposed to negative environments.

[1572] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for acquiring ambient sound, means for converting the acquired sound into text data, means for analyzing and detecting negative expressions from the text data, means for acquiring user emotion data, means for converting the detected negative expressions into positive expressions, means for converting the converted positive expressions into voice data, and means for playing the voice data. This allows the user to receive voice information in a state where negative expressions have been converted into positive expressions.

[1573] "Means for acquiring ambient sound" refers to a device or method that uses a microphone or other sensors to collect environmental sounds and conversations.

[1574] "Means for converting acquired voice into text data" refers to devices or methods that convert voice signals into text information using technologies such as voice recognition service APIs.

[1575] A "means for analyzing and detecting negative expressions from text data" is an apparatus or method that uses natural language processing techniques to identify negative sentiments and expressions within text data.

[1576] "Means for acquiring user emotional data" refers to devices or methods that use cameras, microphones, biometric sensors, etc. to collect emotional information such as the user's facial expressions, voice tone, and heart rate.

[1577] The "means for converting detected negative expressions into positive expressions" refers to a device or method that uses a generative artificial intelligence model to modify negative parts of text data into semantically positive ones.

[1578] The "means for converting the converted positive expression into voice data" refers to a device or method for converting the positive text data into a voice signal using a voice synthesis tool that generates voice from text.

[1579] "Means for reproducing audio data" refers to a device or method for transmitting audio data to a user in real time through a speaker, earphones, etc.

[1580] This invention is a system that captures the user's surrounding voices, converts them into text, converts negative expressions into positive ones, and provides the text back to the user as voice. Furthermore, by combining it with an emotion engine that recognizes the user's emotions, it provides more advanced emotion management. Each element and its operation procedure are explained in detail below.

[1581] 1. System Configuration

[1582] (a) Terminal

[1583] The device is equipped with a microphone, a speaker or earphones, and an emotion engine (camera, microphone, biometric sensors, etc.) to recognize the user's emotions. The device captures the surrounding sounds and sends them to the server. It also receives the audio data sent from the server and plays it back to the user.

[1584] (b) Server

[1585] The server has the following functions:

[1586] Speech recognition function: Converts voice data into text. Specifically, it uses a speech recognition service API (e.g., a general cloud-based speech recognition API).

[1587] Sentiment analysis function: Analyzes text data and detects negative expressions. Uses a sentiment analysis library.

[1588] Conversion function: Uses a generative artificial intelligence model to convert negative expressions into positive ones.

[1589] Speech synthesis function: Use a speech synthesis tool to convert positive text data back into speech data.

[1590] Emotion recognition function: Recognizes the user's emotions and reflects that emotional data in other analysis and conversion processes.

[1591] 2. Explanation of specific operations

[1592] Audio capture and transmission

[1593] The device picks up ambient sounds through a microphone. This audio data is sent to the server in real time in streaming format. At the same time, the device's emotion engine recognizes the user's emotions and sends data such as facial expressions, voice tone, and biometric signals to the server.

[1594] Converting audio data to text

[1595] The server receives the voice data sent from the device and converts it into text using a voice recognition API.

[1596] Detecting negative expressions

[1597] The server analyzes text data using a sentiment analysis library to detect negative expressions, and also considers user sentiment data from an emotion engine to improve analysis accuracy.

[1598] Transforming negative expressions into positive ones

[1599] The server uses a generative artificial intelligence model to convert negative expressions into positive ones, taking into account the user's emotional data to make the conversion more appropriate.

[1600] Vocalization and delivery of positive expressions

[1601] The server reconverts the converted positive text data into voice data using a voice synthesis tool, and the voice data is sent to the terminal.

[1602] Playing audio

[1603] The terminal receives the audio data sent from the server and plays it back to the user through a speaker or earphones.

[1604] 3. Specific Examples

[1605] Example 1: Converting an angry boss's words

[1606] Boss says: "Why can't you do something so simple? I can't believe it!"

[1607] The device picks up this speech with a microphone and transmits the audio data to the server in real time. At the same time, the device captures the user's facial expression with a camera and transmits the data to the server.

[1608] The server uses a speech recognition API to generate text data of the utterance.

[1609] The server uses a sentiment analysis library to detect that the generated text is negative, and based on data from the emotion engine, recognizes that the user is feeling stressed.

[1610] The server uses a generative artificial intelligence model to translate the utterance into a positive expression, such as "You made a lot of mistakes today. Let's think together about how to solve the problem."

[1611] The server uses a speech synthesis tool to convert the text into voice data and send it to the terminal.

[1612] The device receives the converted audio and plays it back to the user through speakers or earphones.

[1613] Example 2: Converting a colleague's grumpy tone

[1614] A colleague says: "I'm tired. I'm going to hate this job."

[1615] The device picks up this speech with a microphone and sends the voice data to the server in real time. At the same time, the emotion engine analyzes the user's voice tone and sends the data to the server.

[1616] The server uses a speech recognition API to generate text data of the utterance.

[1617] The server uses a sentiment analysis library to detect that the generated text is negative, and based on data from the emotion engine, recognizes that the user is relaxed.

[1618] The server uses a generative artificial intelligence model to translate the utterance into a positive expression, generating the text, "I'm a little tired today, but I believe I can get through it."

[1619] The server uses a speech synthesis tool to convert the text into voice data and send it to the terminal.

[1620] The device receives the converted audio and plays it back to the user through speakers or earphones.

[1621] As described above, this system converts negative comments made by users around them into positive ones, and furthermore, by taking into account the user's emotional state and providing optimal voice information, it is possible to improve the user's mental health and productivity.

[1622] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1623] Step 1: Acquire audio

[1624] The device uses a microphone to capture surrounding sounds in real time. For example, it collects the user's voice while speaking and environmental sounds. The input is audio signals from the external environment, and the output is digital audio data. At the same time, the device's emotion engine (camera, microphone, biometric sensors, etc.) captures the user's facial expressions, voice tone, and biometric signals. The input is the user's emotional state, and the output is the user's emotional data.

[1625] Step 2: Sending audio data

[1626] The terminal transmits the acquired voice data and emotion data to the server in streaming format. The input here is the acquired voice data and the user's emotion data, and the output is data transmitted to the server via a network. Specifically, data transmission is performed using wireless communication.

[1627] Step 3: Converting audio data to text

[1628] The server receives the voice data sent from the device and converts it into text data using a speech recognition API (e.g., a common cloud-based speech recognition service). The input is digital voice data, and the output is text data. For example, the generated text might say, "Why can't you do something so simple?" This process requires network and computer resources.

[1629] Step 4: Detecting negative expressions

[1630] The server uses a sentiment analysis library to analyze the converted text data and detect negative expressions. It also uses the user's emotional data from the emotion engine for analysis. The input is the text data and the user's emotional data, and the output is the identification result of text containing negative expressions. Specifically, the text "Why can't you do something so simple?" is determined to be negative.

[1631] Step 5: Transform negative expressions into positive ones

[1632] The server uses a generative artificial intelligence model (e.g., a generative AI model) to convert the detected negative expressions into positive expressions. The input is text data containing negative expressions and the user's emotional data, and the output is text data converted into positive expressions. For example, "Why can't you do something so simple?" is converted into "You made quite a few mistakes today, but let's find a solution together." This process requires natural language processing technology and a generative AI model.

[1633] Step 6: Vocalize and deliver positive feedback

[1634] The server converts the converted positive text data into voice data using a voice synthesis tool (e.g., a voice synthesis engine). The input is the positive text data, and the output is voice data. For example, a voice saying, "You made quite a few mistakes today, but let's find a solution together" is generated. This voice data is then sent to the device.

[1635] Step 7: Playing Audio

[1636] The terminal receives the voice data sent from the server and plays it back to the user through a speaker or earphone. The input is the voice data sent from the server, and the output is the voice played back to the user. This process allows the user to receive positively converted voice information in real time.

[1637] (Application example 2)

[1638] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1639] In communication between employees and customers in physical stores, responding appropriately to negative comments or complaints from customers can be mentally stressful for employees and can reduce work efficiency. Furthermore, if an employee's emotional state is unstable, their response to customers will be poor. To solve this problem, a system is needed that can convert surrounding negative vocal expressions into positive ones and provide optimal vocal information that takes into account the employee's emotional state.

[1640] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1641] In this invention, the server includes means for converting voice into text data, means for analyzing and detecting negative expressions from the text data, means for converting the detected negative expressions into positive expressions, and means for utilizing an emotion engine that recognizes the user's emotions. This enables employees to receive negative comments from customers as voice that has been converted into positive ones, thereby reducing their mental burden and improving the quality of customer service.

[1642] The "means for acquiring ambient sound" is a function for acquiring sound data in the environment using a microphone, a sound acquisition device, or the like.

[1643] "Means for converting acquired voice data into text data" refers to the process of converting acquired voice data into character information (text data) using voice recognition technology or API.

[1644] "Means for analyzing and detecting negative expressions from text data" refers to a function that uses a sentiment analysis library and natural language processing technology to identify expressions that indicate negative emotions or dissatisfaction from input text data.

[1645] The "means for converting detected negative expressions into positive expressions" is a means for converting detected negative text into positive content using a generative artificial intelligence model.

[1646] "Means of utilizing an emotion engine that recognizes the user's emotions" refers to a technology that uses a camera or biometric sensor to monitor the user's facial expressions and biometric signals, and analyzes them to recognize the user's emotional state.

[1647] The "means for reproducing audio data" is a function for allowing the user to hear the generated audio data using an audio reproduction device such as a speaker or earphones.

[1648] This invention is a system designed to support store employees when dealing with customers. Specifically, the system captures surrounding voices, converts them into text data, converts negative expressions into positive ones, and provides the voice data to the store employees. Furthermore, by combining it with an emotion engine that recognizes the emotions of the user (employee), the system aims to achieve advanced emotion management by providing optimal voice information according to the employee's emotional state.

[1649] System Configuration

[1650] (a) Terminal

[1651] The device is equipped with smart glasses, a microphone, a speaker or earphones, and an emotion engine (camera, biometric sensors, etc.) that recognizes the user's emotions. The device captures the surrounding audio and sends it to the server. It also plays back the audio data sent from the server.

[1652] (b) Server

[1653] The server has the following functions:

[1654] Speech recognition function: Converts voice data into text (e.g., Google Speech-to-Text API).

[1655] Sentiment analysis functions: Analyze text data and detect negative expressions (e.g., NLTK or Transformers libraries).

[1656] Transformation function: Transforming negative expressions into positive expressions (e.g., generative AI models such as GPT-3).

[1657] Speech synthesis function: Converts positive text data back into speech data (e.g. gTTS).

[1658] Emotion recognition: Recognizes the user's emotions and incorporates the results into other analysis and transformation processes (e.g., camera and biometric sensors).

[1659] Program processing

[1660] The device uses a microphone to capture the conversation between the customer and employee and sends the audio data in streaming format to the server. At the same time, the device's emotion engine recognizes the user's emotions (voice tone, facial expressions, biometric signals, etc.) and sends this data to the server. The server analyzes the received audio data and converts it into text data using a speech recognition API. This data is used as input data for emotion analysis.

[1661] The server uses a sentiment analysis library to analyze text data and detect negative expressions. It also takes into account the user's emotional data from the emotion engine to improve analysis accuracy. For example, if the user is feeling stressed, the analysis results will be set more sensitively. The server then uses a generative artificial intelligence model to convert negative expressions into positive ones. The server uses the user's emotional data to perform more appropriate conversions. For example, if the user is relaxed, the tone will be adjusted to be gentler.

[1662] The server then converts the converted positive text data into audio data using a speech synthesis tool. This data is then sent to the device and played in real time. The device then receives the audio data sent from the server and plays it back to the user through a speaker or earphones. This protects the user from the negative audio environment around them and allows them to receive only positive information.

[1663] Specific examples

[1664] Example 1: When a customer makes a complaint

[1665] Customer says: "I can't use this product at all, what's going on?"

[1666] What employees ask: "Are you having trouble with this product? Is there anything I can help you with?"

[1667] Prompt Sentence Examples

[1668] "This product is completely unusable. What's going on?" Please change this to a positive expression.

[1669] By using this system, employees can reduce their mental burden and improve the quality of customer service.

[1670] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1671] Step 1:

[1672] The device captures the surrounding audio through a microphone, and this audio data is sent to the server in real time. The input is the surrounding audio data, and the output is sending the audio data to the server.

[1673] Step 2:

[1674] The device's emotion engine (including the camera and biometric sensors) acquires the user's emotional data (facial expressions, biometric signals, etc.). This data is also sent to the server at the same time as the voice data. The input is facial expressions and biometric signals, and the output is sending emotional data to the server.

[1675] Step 3:

[1676] The server converts the received voice data into text data using a speech recognition API. The input is the voice data, and the output is the corresponding text data.

[1677] Step 4:

[1678] The server analyzes the converted text data using a sentiment analysis library to detect negative expressions. The input is the text data, and the output is the result of whether or not it contains negative expressions.

[1679] Step 5:

[1680] The server uses a generative artificial intelligence model (e.g., GPT-3) to convert text containing negative expressions into positive ones, while also taking into account the user's emotional data to adjust the conversion process. The input is negative text data and emotional data, and the output is the text data converted into positive ones.

[1681] Step 6:

[1682] The server converts the positive text data into speech data using a speech synthesis tool (e.g., gTTS). The input is the positive text data, and the output is speech data.

[1683] Step 7:

[1684] The server sends the generated audio data to the terminal. The input is audio data, and the output is sending audio data to the terminal.

[1685] Step 8:

[1686] The device then plays the received audio data through a speaker or earphone. The input is audio data, and the output is the audio the user hears. This protects the user from the negative audio environment around them and allows them to receive only positive information.

[1687] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1688] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1689] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1690] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1691] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1692] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1693] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1694] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1695] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1696] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1697] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1698] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1699] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1700] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1701] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1702] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1703] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1704] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1705] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1706] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1707] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1708] The following is further disclosed regarding the above embodiment.

[1709] (Claim 1)

[1710] a means for acquiring ambient audio;

[1711] A means for converting the acquired voice into text data;

[1712] A means for analyzing and detecting negative expressions from text data;

[1713] A means of converting detected negative expressions into positive expressions;

[1714] a means for converting the converted positive expression into audio data;

[1715] means for playing audio data;

[1716] A system including:

[1717] (Claim 2)

[1718] 10. The system of claim 1, wherein the means for capturing ambient sound comprises a microphone.

[1719] (Claim 3)

[1720] 10. The system of claim 1, wherein the means for converting to text data comprises a speech recognition service API.

[1721] (Claim 4)

[1722] 10. The system of claim 1, wherein the means for analyzing and detecting negative language comprises a sentiment analysis library.

[1723] (Claim 5)

[1724] 10. The system of claim 1, wherein the means for converting negative expressions into positive expressions comprises a generative artificial intelligence model.

[1725] (Claim 6)

[1726] 10. The system of claim 1, wherein the means for converting the positive expressions into audio data comprises a text-to-speech conversion tool.

[1727] (Claim 7)

[1728] 10. The system of claim 1, wherein the means for playing audio data comprises a speaker or earphones.

[1729] "Example 1"

[1730] (Claim 1)

[1731] a means for acquiring ambient audio;

[1732] A means for transmitting the captured audio to a server in real time;

[1733] means for converting the received voice into text data;

[1734] A means for analyzing text data to detect negative expressions;

[1735] A means for converting the detected negative expressions into positive expressions using a generative artificial intelligence model;

[1736] A means for converting positive expressions into audio data;

[1737] means for playing the converted audio data in real time;

[1738] A system including:

[1739] (Claim 2)

[1740] 10. The system of claim 1, wherein the means for capturing ambient audio comprises an audio capture device.

[1741] (Claim 3)

[1742] 10. The system of claim 1, wherein the means for converting to text data comprises a speech recognition service API.

[1743] "Application Example 1"

[1744] (Claim 1)

[1745] a means for acquiring ambient audio;

[1746] A means for converting the acquired voice into text data;

[1747] A means for analyzing and detecting negative expressions from text data;

[1748] A means of converting detected negative expressions into positive expressions;

[1749] a means for converting the converted positive expression into audio data;

[1750] means for playing audio data;

[1751] a means for providing positive audio information in real time to a worker wearing a head-mounted display;

[1752] A system including:

[1753] (Claim 2)

[1754] 10. The system of claim 1, wherein the means for capturing ambient sound comprises a microphone.

[1755] (Claim 3)

[1756] 10. The system of claim 1, wherein the means for converting to text data comprises a speech recognition service API.

[1757] "Example 2: Combining Emotion Engines"

[1758] (Claim 1)

[1759] a means for acquiring ambient audio;

[1760] A means for converting the acquired voice into text data;

[1761] A means for analyzing and detecting negative expressions from text data;

[1762] A means for acquiring user emotion data;

[1763] A means of converting detected negative expressions into positive expressions;

[1764] a means for converting the converted positive expression into audio data;

[1765] means for playing audio data;

[1766] A system including:

[1767] (Claim 2)

[1768] 10. The system of claim 1, wherein the means for capturing ambient sound comprises a microphone.

[1769] (Claim 3)

[1770] 10. The system of claim 1, wherein the means for converting to text data comprises a speech recognition service API.

[1771] (Claim 4)

[1772] 10. The system of claim 1, further comprising an emotion engine for analyzing emotion data of a user.

[1773] (Claim 5)

[1774] 10. The system of claim 1, wherein the means for converting detected negative expressions into positive expressions comprises a generative AI model.

[1775] "Application example 2 when combining emotion engines"

[1776] (Claim 1)

[1777] a means for acquiring ambient audio;

[1778] A means for converting the acquired voice into text data;

[1779] A means for analyzing and detecting negative expressions from text data;

[1780] A means of converting detected negative expressions into positive expressions;

[1781] a means for converting the converted positive expression into audio data;

[1782] a means for utilizing an emotion engine for recognizing the emotion of a user;

[1783] means for playing audio data;

[1784] A system including:

[1785] (Claim 2)

[1786] 10. The system of claim 1, wherein the means for capturing ambient sound comprises a microphone.

[1787] (Claim 3)

[1788] 10. The system of claim 1, wherein the means for converting to text data comprises a speech recognition service API. [Explanation of symbols]

[1789] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. a means for acquiring ambient audio; A means for converting the acquired voice into text data; A means for analyzing and detecting negative expressions from text data; A means of converting detected negative expressions into positive expressions; a means for converting the converted positive expression into audio data; means for playing audio data; A system including:

2. The system of claim 1 , wherein the means for capturing ambient sound includes a microphone.

3. The system of claim 1 , wherein the means for converting to text data includes a speech recognition service API.

4. The system of claim 1 , wherein the means for analyzing and detecting negative language comprises a sentiment analysis library.

5. 10. The system of claim 1, wherein the means for converting negative expressions into positive expressions comprises a generative artificial intelligence model.

6. 10. The system of claim 1, wherein the means for converting the positive expressions into audio data comprises a text-to-speech conversion tool.

7. 2. The system of claim 1, wherein the means for playing audio data includes a speaker or an earphone.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A