System
The system addresses communication errors in aircraft operations by converting voice data to text, translating languages, analyzing emotions, and controlling aircraft operations, thereby improving safety and response times.
Patent Information
- Application Number
- JP2024133625
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-08-08
- Publication Date
- 2026-02-20
AI Technical Summary
Communication errors between aircraft pilots and air traffic controllers due to language differences and emotional states during emergencies can lead to safety risks and delayed responses.
A system that acquires voice data, converts it into text, translates languages, performs emotion analysis, controls automatic or manual operations, and generates real-time advice and suggestions based on the analysis results, while storing data for later analysis.
Minimizes communication errors by providing accurate and timely responses, enhancing aircraft safety and operational efficiency.
Smart Images

Figure 2026030641000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] In conversations between aircraft pilots and air traffic controllers, differences in language, English proficiency, and individual mental states during emergencies or dangers can lead to communication errors, which can result in aircraft accidents. Furthermore, if communication between pilots and controllers is not smooth, it may be difficult to respond immediately and appropriately. To reduce these risks and improve aircraft operation safety, a system is needed that supports accurate information exchange and appropriate responses in real time, regardless of language or emotional state. [Means for solving the problem]
[0005] The present invention provides a system including means for acquiring voice data, means for converting the acquired voice data into text data, means for translating the text data into multiple languages, means for performing emotion analysis on the text data and voice data, means for controlling automatic piloting or manual operation based on the emotion analysis results, means for generating advice and suggestions in real time based on the analysis results and the situation, and means for storing all data for later analysis. This system supports the safe operation of aircraft by analyzing conversations between aircraft pilots and air traffic controllers in real time and determining appropriate responses taking into account differences in language and emotional states.
[0006] "Means for acquiring voice data" means means for capturing voice information in real time using microphones on board aircraft or voice systems in air traffic control rooms and acquiring it in electronic data format.
[0007] The "means for converting acquired voice data into text data" refers to a means for analyzing acquired voice data in real time using a voice recognition engine and converting it into corresponding text data.
[0008] "Means for translating text data into multiple languages" refers to means that has the ability to translate text data from a specified language into other languages using a translation engine.
[0009] "Means for performing emotion analysis of text data and voice data" refers to means for analyzing the emotional aspects and tone of text data and voice data using natural language processing technology and voice analysis technology to evaluate the emotional state of the speaker.
[0010] "Means for controlling automatic operation or manual operation based on emotion analysis results" refers to a control system that issues appropriate automatic operation instructions or assists manual operation while referring to emotion analysis results.
[0011] "Means for generating advice and suggestions in real time based on analysis results and situations" refers to means for generating and providing optimal advice and suggestions to pilots and air traffic controllers, taking into account analysis results and real-time situations.
[0012] "Means for storing all data and using it for later analysis" refers to means for storing all data, such as voice data, text data, emotion analysis results, and control history, for a long period of time and using it for later analysis and system improvement. [Brief explanation of the drawings]
[0013] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0014] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0015] First, the terms used in the following description will be explained.
[0016] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0017] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0018] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0019] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0020] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0021] [First embodiment]
[0022] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0023] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0024] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0025] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0026] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0027] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0028] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0029] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0030] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0031] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0032] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0033] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0034] The present invention provides a system for analyzing conversations between aircraft pilots and air traffic controllers in real time, controlling aircraft operations, and ensuring safety. The following is a detailed description of an embodiment of the system.
[0035] System configuration
[0036] The system consists of the following main components:
[0037] 1. Device that acquires audio data
[0038] 2. A server that converts the acquired voice data into text data
[0039] 3. Server that translates text data into multiple languages
[0040] 4. Server that performs emotion analysis of text data and voice data
[0041] 5. Server that controls automatic or manual operation based on the results of emotion analysis
[0042] 6. Server that generates advice and suggestions in real time based on the analysis results and the situation
[0043] 7. A server that stores all data and uses it for later analysis
[0044] System Operation
[0045] Acquiring audio data
[0046] Voice information from the user (pilot or air traffic controller) is acquired through a microphone on board the aircraft or through the voice system in the air traffic control room. The terminal captures this voice data and transmits it.
[0047] Voice Recognition
[0048] When the voice data arrives at the server, the voice recognition engine activates and converts it into text data. For example, a voice saying "There is an engine problem, please check!" is converted into the corresponding text.
[0049] Text data translation
[0050] The server then identifies the language of this text data and, if necessary, translates it into the appropriate language using a built-in translation engine. For example, a Japanese sentence like "Engine has a malfunction, please check!" is translated into English as "Engine has a malfunction, please check!"
[0051] Emotion analysis
[0052] The translated text data and the original audio data are analyzed by a server, which uses natural language processing and speech analysis techniques to assess the speaker's emotional state, for example, determining whether they are stressed or calm in an emergency.
[0053] Providing operational control and advice
[0054] The server controls autopilot or manual operation based on the emotion analysis results and text content. If it determines that the risk is high, it sends a signal to the aircraft's autopilot system to perform appropriate operations. It also generates advice and suggestions based on the situation in real time and provides them to the user via the terminal. For example, if an engine abnormality is detected, it immediately provides a voice message advising, "Please perform the engine restart procedure."
[0055] Specific examples
[0056] Emergency Scenarios
[0057] 1. The user (pilot) says, "There's an engine malfunction, please check!"
[0058] 2. The device captures this audio and sends it to the server.
[0059] 3. The server performs speech recognition and converts it into text.
[0060] 4. If necessary, the server translates the text into English.
[0061] 5. The server performs emotion analysis and detects that the pilot is in a high stress state.
[0062] 6. The server signals the autopilot system to execute a safety maneuver.
[0063] 7. The server generates an advice saying, "Autopilot has been strengthened due to an engine abnormality. Please check the engine instruments and perform the engine restart procedure." and notifies the user via the terminal.
[0064] 8. All processed data is stored on the server and used for later analysis.
[0065] This system can significantly improve aircraft safety and minimize communication errors caused by language differences or mental state during an emergency.
[0066] The processing flow will be explained below.
[0067] Step 1:
[0068] The device captures the voice data of the pilot and air traffic controller through microphones on the plane or the audio system in the air traffic control room, and transmits the captured voice data to a server in real time.
[0069] Step 2:
[0070] The server receives the voice data. It converts the voice data into text data using a speech recognition engine. For example, a voice message such as "There is an engine error. Please check!" is converted into text format.
[0071] Step 3:
[0072] The server identifies the language of the converted text data, using a built-in language identification engine to determine in which language the text is written.
[0073] Step 4:
[0074] The server translates the text data into the specified language as needed. For example, the Japanese text "Engine has a malfunction, please check!" is translated into English as "Engine has a malfunction, please check!".
[0075] Step 5:
[0076] The server performs sentiment analysis based on the translated text data and the original audio data, using natural language processing and audio analysis techniques to assess the speaker's emotional state (e.g., stressed, overwhelmed, calm, etc.).
[0077] Step 6:
[0078] The server assesses the risk level based on the results of emotion analysis and the text content. If the risk level is deemed high, the server sends a signal to the autopilot system to temporarily take control of the aircraft.
[0079] Step 7:
[0080] The server generates optimal countermeasures in real time based on the analysis results and the situation. This includes specific procedures and advice. For example, advice such as "Perform the engine restart procedure" is generated.
[0081] Step 8:
[0082] The terminal then notifies the pilot or controller of the generated advice in real time as a voice or text message.
[0083] Step 9:
[0084] The server stores all voice data, text data, translation data, sentiment analysis results, and response actions as logs, which will be used later for further analysis and system improvement.
[0085] Example 1
[0086] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0087] In communication between aircraft pilots and air traffic controllers, errors in information transmission can occur due to language differences or the pilot's mental state during an emergency. Such errors can have a significant impact on aircraft operation and pose a risk of compromising safety. It is also necessary to accurately grasp the pilot's stress level and emotions and provide appropriate responses in real time, but current systems are unable to adequately achieve this.
[0088] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0089] In this invention, the server includes means for acquiring voice data, means for converting the acquired voice data into text data, means for translating the text data into multiple languages, means for performing emotion analysis of the text data and voice data, means for controlling automatic or manual operation based on the emotion analysis results, means for generating advice and suggestions in real time based on the analysis results and the situation, means for storing all data for later analysis, means for transmitting the acquired voice data in digital form, means for converting the voice data into text data using a voice recognition engine, means for passing the voice data and text data to an emotion analysis engine, means for evaluating the emotional state using the emotion analysis engine, means for sending signals to an automatic operation system, and means for generating and transmitting voice messages based on the analysis results. This makes it possible to minimize communication errors due to language differences or mental states during an emergency, thereby significantly improving aircraft safety.
[0090] "Voice data" refers to data in which voice information uttered by a user is recorded in digital format.
[0091] "Text data" is character string information converted from voice data using a voice recognition engine.
[0092] A "translation engine" is a system that converts text data from an identified language to another language.
[0093] An "emotion analysis engine" is a system that analyzes text and audio data to evaluate the speaker's emotional state.
[0094] An "automatic operation system" is a system that performs automatic piloting and other automatic control of an aircraft based on the results of analysis.
[0095] "Digital format" refers to a format in which analog data is converted into numerical data.
[0096] A "voice recognition engine" is software or hardware for converting voice data into text data.
[0097] A "signal" is an electronic signal for transmitting control information.
[0098] "Real-time" refers to data being processed and analyzed as soon as it is acquired.
[0099] "Analysis results" are the analysis results derived by the sentiment analysis engine and other analysis means.
[0100] "Advice and suggestions" are measures and instructions generated based on the analysis results and the situation.
[0101] "Storage" refers to keeping acquired data and analysis results for a certain period of time.
[0102] "Later analysis" refers to re-analysis at a later date using the stored data.
[0103] This invention is a system for analyzing conversations between aircraft pilots and air traffic controllers in real time to control aircraft operations and ensure safety. The system captures voice data, converts it into text data, translates it into multiple languages, performs sentiment analysis, controls automatic or manual operations based on the analysis results, and generates advice and suggestions in real time as needed. Furthermore, all data is stored for later analysis.
[0104] System configuration and operation
[0105] Acquiring audio data
[0106] The user (pilot or air traffic controller) reports the emergency situation by voice, and this voice data is picked up by the terminal in digital format through the microphone on board or the voice system in the air traffic control room, and then transmitted to the server.
[0107] Voice Recognition
[0108] The server passes the received voice data to a voice recognition engine. This engine generally uses high-performance voice recognition technology such as Google Speech-to-Text API or IBM Watson Speech to Text. The voice data is converted into text data. For example, a voice saying "There is an engine problem, please check!" is converted into text "There is an engine problem, please check!"
[0109] Text data translation
[0110] The server then automatically identifies the language of the text data. If the identified language is other than English, the server uses a built-in translation engine (e.g., Microsoft Translator or Google Translate API) to translate the text data into English. For example, the Japanese input "Engine has a malfunction, please check!" is translated into "Engine has a malfunction, please check!"
[0111] Emotion analysis
[0112] The server passes the translated text data and the original audio data to a sentiment analysis engine, which uses natural language processing and speech analysis technologies such as IBM Watson Natural Language Understanding and Microsoft Azure Text Analytics to assess the speaker's emotional state, for example, determining whether they are highly stressed or calm in an emergency.
[0113] Providing operational control and advice
[0114] The server sends a signal to the aircraft's automatic operation system based on the results of the emotion analysis and the content of the text data. If a high level of danger is detected, the server immediately executes safety operations. It also generates advice and suggestions based on the situation in real time and provides them to the user via the terminal. For example, advice such as "Perform engine restart procedures" is immediately provided as a voice message.
[0115] Data storage
[0116] All processed data is stored on a server as a detailed log, which is used for future analysis, system improvement, aircraft accident investigations, and operational optimization.
[0117] Specific examples
[0118] Emergency Scenarios
[0119] 1. The user (pilot) reports audibly, "An engine malfunction has occurred. Please check!"
[0120] 2. The device captures this audio and immediately transmits it digitally to the server.
[0121] 3. The server passes the received voice data to a voice recognition engine and converts it into text data.
[0122] 4. The server identifies the language of the text data and translates it into English using a translation engine if necessary.
[0123] 5. The server analyzes the translated text data and the original audio data using an emotion analysis engine to detect high stress levels.
[0124] 6. The server sends a signal to the automated operation system to immediately execute safety operations, and also creates a situation-specific advice message, such as "Perform engine restart procedure," as a voice message and provides it to the user via the terminal.
[0125] 7. All processing data will be stored on the server and used for future analysis and improvement.
[0126] Prompt Sentence Examples
[0127] Prompt: "Analyze the audio of an aircraft pilot reporting an emergency and provide a detailed description of how the automated flight control system responds and provides advice to the pilot."
[0128] The use of this system can significantly improve aviation safety and minimize communication errors caused by language differences or mental state during an emergency.
[0129] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0130] The flow of this system's program processing
[0131] Step 1:
[0132] The user (pilot or air traffic controller) reports information verbally. For example, in an emergency situation, they might say, "There's an engine malfunction, please check!" This is the input.
[0133] Step 2:
[0134] The terminal captures the user's voice in digital form through an onboard microphone or an audio system in the air traffic control room, and transmits this digital voice data to a server.
[0135] Input: User's voice
[0136] Output: Digital audio data
[0137] Step 3:
[0138] The server passes the received digital voice data to a voice recognition engine (e.g., Google Speech-to-Text API or IBM Watson Speech to Text) and converts it into text data.
[0139] Input: Digital audio data
[0140] Output: Text data: "An engine error has occurred. Please check!"
[0141] Step 4:
[0142] Identifying the language of server-generated text data, for example automatically recognizing that the text is in Japanese.
[0143] Input: Text data
[0144] Output: Language identification result (Japanese)
[0145] Step 5:
[0146] The server uses a built-in translation engine (e.g., Microsoft Translator or Google Translate API) to translate from the identified language to English. For example, "Engine has a malfunction, please check!" is translated to "Engine has a malfunction, please check!"
[0147] Input: Japanese text data
[0148] Output: English text data
[0149] Step 6:
[0150] The server passes the translated English text data and the original audio data to a sentiment analysis engine (e.g., IBM Watson Natural Language Understanding or Microsoft Azure Text Analytics) to evaluate the speaker's emotional state, e.g., to determine whether the speaker has a high stress level.
[0151] Input: English text and audio data
[0152] Output: Emotion analysis results (high stress)
[0153] Step 7:
[0154] Based on the results of the emotion analysis and the content of the text data, the server sends a signal to the aircraft's automatic operation system to perform appropriate safety operations, such as sending a signal to "strengthen autopilot."
[0155] Input: Sentiment analysis results and text data
[0156] Output: Control signal to automatic operation system
[0157] Step 8:
[0158] The server generates advice and suggestions based on the situation in real time and provides them to the user through the terminal, for example, generating a voice message saying "Please perform the engine restart procedure."
[0159] Input: Sentiment analysis results and situation information
[0160] Output: Advice and suggestions (voice message)
[0161] Step 9:
[0162] The server stores all transaction data in a detailed log, which is used for future analysis and system improvement.
[0163] Input: All analysis results and process logs
[0164] Output: Saved processed data
[0165] (Application example 1)
[0166] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0167] In on-site security work, security personnel often experience high levels of stress during emergencies, which can result in delayed responses and mistakes. Furthermore, communication barriers in multilingual environments and difficulty assessing situations in real time can also be a factor, potentially worsening dangerous situations.
[0168] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0169] In this invention, the server includes means for acquiring voice data, means for converting the acquired voice data into text data, means for translating the text data into multiple languages, means for performing emotion analysis of the text data and voice data, means for controlling automatic or manual operation based on the emotion analysis results, means for generating advice and suggestions in real time based on the analysis results and the situation, means for alerting the user when a high-stress state is detected, and means for saving all data for later analysis, thereby enabling quick and appropriate decisions to be made even in emergencies or high-stress situations.
[0170] The "means for acquiring voice data" refers to a device or software that captures the voice uttered by the user in real time and transmits the voice data to the system.
[0171] The "means for converting acquired voice data into text data" refers to a device or software that uses voice recognition technology to convert acquired voice data into a corresponding string of characters.
[0172] A "means for translating text data into multiple languages" is a translation engine or algorithm for converting text data into another language.
[0173] "Means for analyzing emotions in text data and voice data" refers to devices or software that use natural language processing and voice analysis techniques to evaluate the emotions of a speaker from text data or voice data.
[0174] "Means for controlling automatic or manual operations based on the results of emotion analysis" refers to devices or software that appropriately manage or correct automatic operations or manual instructions within the system based on the results of emotion analysis.
[0175] "Means for generating advice and suggestions in real time based on analysis results and the current situation" refers to devices or software that generate appropriate advice and suggestions in a timely manner based on data analysis results and the current situation.
[0176] "Means for alerting the user when a high stress state is detected" refers to devices or software that alert the user by issuing appropriate alerts or notifications when it is determined that the user's stress level is high.
[0177] "Means for storing all data and making it available for later analysis" refers to devices or software that store all data processed by the system and make it available for later analysis, performance evaluation, etc.
[0178] The present invention is a real-time assistant system for supporting security operations. This system has the functions of capturing voice data, converting it into text data, translating it into multiple languages, performing emotion analysis, providing advice in real time, and detecting high-stress situations.
[0179] System configuration
[0180] The system consists of the following main components:
[0181] 1. A means of acquiring voice data: Capture the user's voice in real time, specifically using a smartphone or microphone.
[0182] 2. Means for converting acquired voice data into text data: Convert the voice into text using voice recognition technology (for example, the SpeechRecognition library).
[0183] 3. Means for translating text data into multiple languages: Translate text data into other languages using a translation engine (e.g., Google Translate API).
[0184] 4. Means for sentiment analysis of text and voice data: Sentiment analysis is performed using natural language processing technology and voice analysis technology (e.g., Hugging Face Transformers).
[0185] 5. Means for controlling automatic or manual operations based on the results of sentiment analysis: Based on the results of sentiment analysis, operations within the system are appropriately managed or corrected.
[0186] 6. Means of generating advice and suggestions in real time based on analysis results and the current situation: Generate appropriate advice and suggestions based on the results of data analysis and the current situation.
[0187] 7. A method to alert users when high stress levels are detected: When stress levels are determined to be high, appropriate alerts and notifications will be provided.
[0188] 8. Means of storing all data for later analysis: All processed data will be stored and used for later analysis and performance evaluation.
[0189] System Operation
[0190] The specific operation of the system can be explained as follows:
[0191] Acquiring voice data and converting it to text data
[0192] The voice spoken by the user (security officer) is picked up through the smartphone's microphone. This voice data is converted into text using voice recognition technology. For example, a voice saying "Someone is breaking in! Call for help!" is converted into text data.
[0193] Text data translation
[0194] This text data is translated from Japanese to English using a language recognition function. Using a translation engine, "Someone is trespassing, please call for help immediately!" is converted to "Someone is trespassing, please call for help immediately!"
[0195] Emotion analysis
[0196] The translated text data and the original audio data are analyzed by a sentiment analysis model to determine the user's emergency state or high stress level. If a specific emotional state is detected, appropriate advice is generated.
[0197] Real-time advice and suggestions
[0198] Based on the results of sentiment analysis and the text content, the system generates real-time advice. For example, if a high stress state is detected, the system will provide advice such as "Be careful. Stress is increasing."
[0199] Data storage
[0200] All data is logged and made available for later analysis or performance evaluation.
[0201] Specific examples
[0202] For example, if a security worker says, "Someone is trespassing, please call for help immediately!", the system converts the speech into text and translates it into "Someone is trespassing, please call for help immediately!". If the emotion analysis determines that this is an "emergency situation," the system provides a real-time voice message with advice such as, "Be careful. Stress is building."
[0203] Prompt Sentence Examples
[0204] Voice: "Someone's breaking in, hurry up and call for help!"
[0205] Translation: "Someone is trespassing, please call for help immediately!"
[0206] Sentiment Analysis: "Emergency"
[0207] Advice: "Pay attention. Stress is building."
[0208] This allows security personnel to respond quickly and appropriately.
[0209] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0210] Step 1:
[0211] The device captures the user's voice in real time through a microphone. The input is the user's speech, and the output is digital voice data. This voice data is sent directly to the server.
[0212] Step 2:
[0213] The server converts the acquired voice data into text data using a voice recognition engine (for example, the SpeechRecognition library). The input is digital voice data, and the output is the corresponding text data. This conversion makes it easier to handle voice information as text information.
[0214] Step 3:
[0215] The server sends the generated text data to a translation engine (for example, Google Translate API) to translate it into another language. The input is text data in the original language, and the output is the translated text data. This translation allows for smooth support in multilingual environments.
[0216] Step 4:
[0217] The server performs sentiment analysis using the translated text data and the original voice data. Natural language processing techniques (e.g., Hugging Face Transformers) are used for sentiment analysis. The input is text and voice data, and the output is the analysis result of the emotional state. Based on the analysis result, the user's state of urgency and stress level are evaluated.
[0218] Step 5:
[0219] The server controls automatic or manual operation based on the results of emotion analysis. Specifically, if a high-stress state is detected, it automatically changes the system operation or issues instructions to the user. The input to this step is the emotion analysis result, and the output is appropriate operation control instructions.
[0220] Step 6:
[0221] The server generates advice and suggestions in real time based on the analysis results and the current situation. For example, if a high stress state is detected, advice such as "Be careful. Your stress is increasing" is generated. The input is the analysis results and situation information, and the output is the generated advice.
[0222] Step 7:
[0223] The server sends the generated advice to the terminal and notifies the user. The input is the generated advice, and the output is the notification received by the user.
[0224] Step 8:
[0225] All data is stored by the server. The stored data is used for later analysis and performance evaluation. The input is all the data during and after processing, and the output is the stored data. This storage allows for future optimization and improvement.
[0226] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0227] This invention is a system for analyzing conversations between aircraft pilots and air traffic controllers in real time, controlling aircraft operations, and ensuring safety. The system incorporates an emotion engine for recognizing the user's emotional state, thereby further improving safety. The following is a detailed description of an embodiment of the system.
[0228] System configuration
[0229] The system consists of the following main components:
[0230] 1. Device that acquires audio data
[0231] 2. A server that converts the acquired voice data into text data
[0232] 3. Server that translates text data into multiple languages
[0233] 4. Server that performs emotion analysis of text data and voice data
[0234] 5. Emotion Engine
[0235] 6. Server that controls automatic or manual operation based on the results of emotion analysis
[0236] 7. Server that generates advice and suggestions in real time based on the analysis results and the situation
[0237] 8. A server that stores all data and uses it for later analysis
[0238] System Operation
[0239] Acquiring audio data
[0240] Voice information from the user (pilot or air traffic controller) is acquired through a microphone on board the aircraft or through the voice system in the air traffic control room. The terminal captures this voice data and transmits it.
[0241] Voice Recognition
[0242] When the voice data arrives at the server, the voice recognition engine activates and converts it into text data. For example, a voice message such as "There is an engine error. Please check!" is converted into text format.
[0243] Text data translation
[0244] The server then identifies the language of this text data and, if necessary, translates it into the appropriate language using a built-in translation engine. For example, a Japanese sentence like "Engine has a malfunction, please check!" is translated into English as "Engine has a malfunction, please check!"
[0245] Emotion analysis
[0246] The translated text data and the original audio data are analyzed by a server, which uses natural language processing and speech analysis techniques to assess the speaker's emotional state, for example, determining whether they are highly stressed or calm in an emergency.
[0247] Use of emotion engine
[0248] The server is equipped with an emotion engine that can recognize the user's emotional state in more detail. This emotion engine analyzes both text and voice data and has the ability to highly evaluate the user's emotional state. This allows the system to determine the level of urgency corresponding to a specific emotional state and propose appropriate countermeasures accordingly.
[0249] Providing operational control and advice
[0250] The server controls autopilot or manual operation based on the emotion analysis results and text content. If it determines that the risk is high, it sends a signal to the aircraft's autopilot system to perform appropriate operations. It also generates advice and suggestions based on the situation in real time and provides them to the user via the terminal. For example, if an engine abnormality is detected, it immediately provides a voice message advising, "Please perform the engine restart procedure."
[0251] Data storage
[0252] The server stores all processing data, analysis results, audio data, and response measures as logs, which will be used later for detailed analysis and system improvement.
[0253] Specific examples
[0254] Emergency Scenarios
[0255] 1. The user (pilot) says, "There's an engine malfunction, please check!"
[0256] 2. The device captures this audio and sends it to the server.
[0257] 3. The server performs speech recognition and converts it into text.
[0258] 4. If necessary, the server translates the text into English.
[0259] 5. The server performs emotion analysis and detects that the pilot is in a high stress state.
[0260] 6. The emotion engine analyzes the user's stress level in detail and determines that the level of urgency is high.
[0261] 7. The server signals the autopilot system to execute a safety maneuver.
[0262] 8. The server generates an advice saying, "Autopilot has been strengthened due to an engine abnormality. Please check the engine instruments and perform the engine restart procedure." and notifies the user via the terminal.
[0263] 9. All processed data is stored on the server and used for later analysis.
[0264] The system can significantly improve aviation safety and minimize communication errors by providing countermeasures based on the user's emotional state.
[0265] The processing flow will be explained below.
[0266] Step 1:
[0267] The device captures the voice data of the pilot and air traffic controller through microphones on the plane or the audio system in the air traffic control room, and transmits the captured voice data to a server in real time.
[0268] Step 2:
[0269] The server receives the voice data. It converts the voice data into text data using a speech recognition engine. For example, a voice message such as "There is an engine error. Please check!" is converted into text format.
[0270] Step 3:
[0271] The server identifies the language of the converted text data, using a built-in language identification engine to determine in which language the text is written.
[0272] Step 4:
[0273] The server translates the text data into the specified language as needed. For example, the Japanese text "Engine has a malfunction, please check!" is translated into English as "Engine has a malfunction, please check!".
[0274] Step 5:
[0275] The server performs sentiment analysis based on the translated text data and the original audio data, using natural language processing and audio analysis technologies to assess the speaker's emotional state (e.g., stressed, overwhelmed, calm, etc.).
[0276] Step 6:
[0277] The emotion engine performs detailed analysis of text and voice data to intelligently assess the user's emotional state, for example detecting when a pilot is under high stress.
[0278] Step 7:
[0279] The server assesses the level of danger based on the results of sentiment analysis and the text content. If the level of danger is deemed high, the server sends a signal to the autopilot system to temporarily control the aircraft's operation. For example, if a pilot is under high stress, the autopilot system will automatically maintain the aircraft's altitude and speed.
[0280] Step 8:
[0281] The server generates optimal countermeasures in real time based on the analysis results and the situation. This includes specific procedures and advice. For example, advice such as "Perform the engine restart procedure" is generated.
[0282] Step 9:
[0283] The terminal then notifies the pilot or controller of the generated advice in real time as a voice or text message.
[0284] Step 10:
[0285] The server stores all voice data, text data, translation data, sentiment analysis results, and response actions as logs, which will be used later for further analysis and system improvement.
[0286] Example 2
[0287] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0288] In communications between aircraft pilots and air traffic controllers, differences in emotional state and language during emergencies can make it difficult to understand instructions, delaying appropriate responses. To solve this problem, real-time voice analysis, evaluation of emotional state, and proposal of appropriate countermeasures are required, but previous systems have found it difficult to perform these tasks in an integrated manner.
[0289] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for acquiring voice data, means for converting the acquired voice data into text data, means for translating the text data into multiple languages, means for performing emotion analysis of the text data and voice data, means for controlling automatic operation or manual operation based on the emotion analysis results, means for generating advice and suggestions in real time based on the analysis results and the situation, means for saving all data and using it for later analysis, and means for incorporating an emotion engine that highly evaluates the user's emotional state. This makes it possible to provide prompt measures based on accurate voice recognition and emotion analysis in real time.
[0290] "Voice data" refers to voices uttered by a user captured as digital signals.
[0291] "Means for acquiring" refers to the technology or device for acquiring voice data using a terminal or sensor within the domain.
[0292] "Text data" is voice data that has been analyzed and expressed as text information.
[0293] "Means for converting into text data" refers to technology or devices that convert acquired voice data into text format using natural language processing technology.
[0294] "Translation means" refers to the technology and software used to convert text data into other languages.
[0295] "Sentiment analysis" is the process of analyzing audio and text data to assess the emotional state of a speaker.
[0296] "Means for performing emotion analysis" refers to technologies or devices that use the algorithms or software required for emotion analysis.
[0297] An "emotion engine" is specialized software or algorithms that intelligently assess a user's emotional state.
[0298] "Means for controlling autopilot or manual operation" refers to technology or devices that automatically control aircraft operation based on the results of emotion analysis or in accordance with the pilot's instructions.
[0299] The "means for generating advice or suggestions" refers to technology or software that generates suggestions in real time to help the user take appropriate actions or make appropriate decisions based on the analysis results and the situation.
[0300] "Means for storing data and using it for later analysis" refers to the technology and devices that record all processing data and use it for later data analysis and system improvement.
[0301] This invention is a system for analyzing conversations between aircraft pilots and air traffic controllers in real time to control aircraft operations and ensure safety. The system uses multiple servers, terminals, and software to achieve functions such as voice data acquisition, voice recognition, text data translation, emotion analysis, autopilot or manual operation control, advice and suggestion generation, and data storage.
[0302] System configuration
[0303] The system consists of the following main components:
[0304] 1. Device that acquires audio data
[0305] 2. Server that converts voice data into text data
[0306] 3. Server that translates text data into multiple languages
[0307] 4. Server that performs emotion analysis of text data and voice data
[0308] 5. Emotion Engine
[0309] 6. Server that controls automatic or manual operation based on the results of emotion analysis
[0310] 7. Server that generates advice and suggestions in real time based on the analysis results and the situation
[0311] 8. A server that stores all data and uses it for later analysis
[0312] Hardware and software used
[0313] 1. Terminal: The aircraft's microphone and communication system are used to acquire voice data. Standard data communication protocols are used to transmit the voice data.
[0314] 2. Server: To analyze the voice data, a speech recognition engine such as Google Cloud Speech-to-Text or Amazon Transcribe is used.
[0315] 3. Text data translation: Translate text data into multiple languages using the Google Translate API, etc.
[0316] 4. Sentiment Analysis: Evaluate emotional state using IBM Watson Natural Language Understanding and Amazon's speech analysis technology.
[0317] 5. Emotional Engine: Utilizing specialized engines such as Affectiva to provide a sophisticated assessment of the user's emotional state.
[0318] 6. Autopilot and manual operation control: Works in conjunction with the aircraft's autopilot system (e.g., Honeywell system) to ensure safe operation.
[0319] 7. Advice and suggestion generation: Use dedicated advice generation algorithms to generate context-sensitive advice in real time.
[0320] 8. Data storage: All processed data, analysis results, audio data, and response actions will be stored using cloud storage for later analysis.
[0321] Specific examples
[0322] Emergency scenarios:
[0323] 1. The user (pilot) says, "There's an engine malfunction, please check!"
[0324] 2. The onboard microphone captures the audio and sends the data to the device.
[0325] 3. The device sends the audio data to the server.
[0326] 4. The server receives the audio data and calls the Google Cloud Speech-to-Text service.
[0327] 5. The server converts the voice data into text and generates the text "An engine error has occurred, please check!"
[0328] 6. The server translates the text data into English using the Google Translate API. "Engine has a malfunction, please check!" becomes "Engine has a malfunction, please check!"
[0329] 7. The server sends the text data to IBM Watson Natural Language Understanding for sentiment analysis.
[0330] 8. The server sends the audio data to Amazon Transcribe, which analyzes the emotional state of the audio.
[0331] 9. The emotion engine on the server integrates the text and voice data to perform a detailed emotion assessment.
[0332] 10. The emotion engine determines that the user's stress level is "high."
[0333] 11. The server sends a signal to the autopilot system, generating the message "Perform engine restart procedure."
[0334] 12. The device will announce this message to the user audibly.
[0335] 13. The server records the processed data, analysis results, audio data, and response measures as logs and stores them in a database.
[0336] 14. The stored data can be later analyzed to help improve the system.
[0337] Prompt Sentence Examples
[0338] "Please explain the steps your system takes when a user says, 'There's an engine malfunction, please check!'"
[0339] In this way, the system significantly improves aircraft safety, can provide countermeasures based on the user's emotional state, and minimizes communication errors.
[0340] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0341] Step 1:
[0342] The voice of the user (pilot or air traffic controller) is acquired through an on-board microphone or the voice system in the air traffic control room. This captures the user's real-time voice information as digital voice data. The input is the user's voice, and the output is digital voice data.
[0343] Step 2:
[0344] The device sends the captured audio data to a server. The audio data is transferred to the server via the Internet or a dedicated communication line. The input is the captured digital audio data, and the output is the transferred audio data.
[0345] Step 3:
[0346] The server receives the voice data and calls the Google Cloud Speech-to-Text service to perform speech recognition. The voice data is analyzed and converted into text data. For example, a speech saying "There is an engine problem, please check!" is converted into text "There is an engine problem, please check!" The input is voice data, and the output is text data.
[0347] Step 4:
[0348] The server identifies the language of the generated text data. If necessary, it uses the Google Translate API to translate the text into other languages. For example, text entered in Japanese is translated into English. The input is the text data, and the output is the translated text data.
[0349] Step 5:
[0350] The server sends the text data to IBM Watson Natural Language Understanding for sentiment analysis. The text data is used to evaluate the speaker's emotional state. For example, the text content can be used to analyze a pilot's stress level and urgency. The input is text data, and the output is an evaluation of the emotional state.
[0351] Step 6:
[0352] The server sends the audio data to Amazon Transcribe, which analyzes the tone and pitch of the voice. This allows for a detailed assessment of the speaker's emotional state and urgency from the audio data. The input is the audio data, and the output is the emotional assessment result from the audio analysis.
[0353] Step 7:
[0354] The emotion engine in the server integrates the results of the analysis of the text data and voice data to perform a detailed emotion evaluation. The emotion engine highly evaluates the user's emotional state and determines, for example, that the user is in a high-stress state. The inputs are the emotion analysis results and the voice analysis results, and the output is the integrated emotion evaluation result.
[0355] Step 8:
[0356] The server generates a control signal to an autopilot system or manual operation based on the emotion analysis result, for example, to signal an aircraft's autopilot system to execute emergency measures. The input is the integrated emotion evaluation result, and the output is a control signal.
[0357] Step 9:
[0358] The server generates appropriate advice or suggestions in real time based on the analysis results and the situation, and notifies the user through the terminal. For example, it generates advice such as "Please perform the engine restart procedure." The input is the emotion analysis results and situation information, and the output is the generated advice.
[0359] Step 10:
[0360] The server records and stores all processing data, analysis results, audio data, and response measures as logs. The stored data is used for later detailed analysis and system improvement. The input is the processing data and analysis results, and the output is the stored log data.
[0361] (Application example 2)
[0362] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0363] Conventional call center systems have the problem of being unable to grasp the emotional state of the operator and customer in real time and provide immediate solutions. This has led to problems such as customers' dissatisfaction not being resolved properly and operators feeling high levels of stress. This leads to problems such as lower customer satisfaction and a decline in the work efficiency of operators.
[0364] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0365] In this invention, the server includes means for acquiring voice data, means for converting the acquired voice data into text data, means for translating the text data into multiple languages, means for performing emotion analysis of the text data and voice data, means for controlling automatic or manual operations based on the emotion analysis results, means for generating advice and suggestions in real time based on the analysis results and the situation, means for triage and proposing countermeasures based on the emotional state, and means for saving all data and using it for later analysis. This makes it possible to grasp the emotional states of customers and operators in real time and immediately provide appropriate countermeasures.
[0366] "Audio data" refers to sound information in digital or analog form collected through audio equipment such as a microphone.
[0367] "Text data" is information obtained by converting voice data into a character string format.
[0368] A "multilingual translation method" is a technique or process for converting text data into different languages.
[0369] "Emotion analysis" is a technique that analyzes audio and text data to assess a speaker's emotional state.
[0370] "Autonomous operation" means that the system performs a specific action without human intervention.
[0371] "Manual operation" means direct human involvement in operating a system.
[0372] "Means for generating advice and suggestions in real time" refers to techniques and methods that provide immediate advice to operators or users based on the current situation.
[0373] "Triage" is the prioritization of response measures based on urgency and importance.
[0374] A "solution" is an action or suggestion to be taken in response to a particular situation.
[0375] "Storing data" means recording collected data on a storage medium so that it can be used later.
[0376] "Use for later analysis" means to conduct analysis based on the stored data for future consideration and improvement.
[0377] MODE FOR CARRYING OUT THE INVENTION
[0378] The present invention is a system for supporting appropriate responses by analyzing the emotional states of customers and agents in a call center in real time. An embodiment of this system will be described in detail below.
[0379] System configuration
[0380] The system consists of the following main components:
[0381] 1. Means of capturing voice data: microphones and telephone systems to capture conversations between customers and operators.
[0382] 2. A means of converting the acquired voice data into text data: a speech recognition engine that converts the voice into a string format (e.g., Google Speech-to-Text API).
[0383] 3. A means of translating text data into multiple languages: A system that translates text into the required languages (e.g., Microsoft Azure Cognitive Services).
[0384] 4. Means of sentiment analysis of text and voice data: Software that analyzes the emotional state of customers and operators (e.g., IBM Watson Tone Analyzer).
[0385] 5. Means for controlling automatic or manual operations based on the results of emotion analysis: A control system for performing processing according to emotional states.
[0386] 6. Means of generating advice and suggestions in real time based on analysis results and situations: An engine that generates real time advice and suggestions to operators.
[0387] 7. Triage and suggest responses based on emotional state: A mechanism for prioritizing responses based on urgency and importance.
[0388] 8. A means of storing all data and making it available for later analysis: a database to store the collected and analyzed data.
[0389] System Operation
[0390] Audio Acquisition
[0391] Voice information from users (customers and operators) is acquired through a microphone or telephone system. The terminal captures this voice data and sends it to the server.
[0392] Voice Recognition
[0393] The voice data received by the server is converted into text data by a voice recognition engine (Google Speech-to-Text API).
[0394] Text data translation
[0395] The server then converts this text data into the appropriate language if it needs to be translated using a multi-language translation means.
[0396] Emotion analysis
[0397] The translated text data and the original audio data are then subjected to natural language processing and speech analysis techniques to assess the user's emotional state, using analysis engines such as IBM Watson Tone Analyzer.
[0398] Detailed analysis by Emotion Engine
[0399] The emotion engine performs advanced analysis of text and voice data to determine specific emotional states and their urgency.
[0400] Providing operational control and advice
[0401] Based on the results of the emotion analysis, the server controls automatic or manual operations and generates and notifies advice and suggestions in real time. Appropriate advice is provided to the operator in voice or text format.
[0402] Data storage
[0403] The server stores all audio data, text data, and analysis results in a database for later analysis.
[0404] Specific examples
[0405] Emergency Scenarios
[0406] 1. A customer says, "I recently purchased a broken item and I'm very unhappy with it. I'd like to process the return immediately."
[0407] 2. The device captures this audio and sends it to the server.
[0408] 3. The server performs speech recognition and converts it into text.
[0409] 4. If necessary, the server translates the text.
[0410] 5. The server performs sentiment analysis and detects that the customer is in a high stress state.
[0411] 6. The emotion engine analyzes the customer's stress level in detail and determines that the situation is urgent.
[0412] 7. The server advises the operator, "The customer is unhappy. Please stay calm and explain with specific examples."
[0413] 8. All processed data is stored on the server and used for later analysis.
[0414] This system will significantly improve the quality and efficiency of customer service at call centers, increasing customer satisfaction. It will also reduce stress on operators and contribute to improved work efficiency.
[0415] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0416] Step 1:
[0417] The user (customer and operator) starts a conversation. The terminal captures the voice data through a microphone or telephone system. This captured voice data is recorded in digital format and sent to the server. The input here is the raw voice data, and the output is the audio file sent to the server.
[0418] Step 2:
[0419] The server sends the received voice data to the Google Speech-to-Text API for speech recognition. In this process, the voice data is converted into text data in string format. The input is the voice data, and the output is the converted text data. Specifically, the voice data undergoes signal processing and pattern matching.
[0420] Step 3:
[0421] The server uses Microsoft Azure Cognitive Services to translate the acquired text data into multiple languages as needed. In this step, the language of the text data is identified and converted into the appropriate language. The input is text data, and the output is translated text data. The specific operation utilizes natural language processing algorithms.
[0422] Step 4:
[0423] The server then sends the translated text data and the original audio data to the IBM Watson Tone Analyzer for sentiment analysis. The input is text data and audio data, and the output is the sentiment analysis results. This function evaluates the speaker's emotional state based on the tone and content of the audio.
[0424] Step 5:
[0425] The server further evaluates the emotion analysis results to determine the specific emotional state and its urgency. This is done by the emotion engine, which understands the user's emotional state. The input is the emotion analysis results, and the output is detailed emotional state and urgency information. Specific operations include integrating and evaluating the analysis results.
[0426] Step 6:
[0427] The server generates and notifies the operator in real time with advice and suggestions based on the results of emotion analysis. The input is the detailed emotional state and the customer's comments, and the output is advice and suggestions provided to the operator. Specifically, the system generates appropriate countermeasures and notifies the operator via voice or text.
[0428] Step 7:
[0429] The server stores all data in a database for later analysis. The input is all processed data (audio, text, analysis results, etc.), and the output is the stored data. Specific operations include recording and storing data.
[0430] Prompt Sentence Examples
[0431] Speech input: "I'm very unhappy with the product I recently purchased, which is broken. I'd like to process the return immediately."
[0432] Expected output: "The customer is unhappy. Please stay calm and provide specific examples. We'll expedite the return process."
[0433] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0434] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0435] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0436] [Second embodiment]
[0437] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0438] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0439] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0440] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0441] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0442] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0443] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0444] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0445] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0446] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0447] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0448] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0449] The present invention provides a system for analyzing conversations between aircraft pilots and air traffic controllers in real time, controlling aircraft operations, and ensuring safety. The following is a detailed description of an embodiment of the system.
[0450] System configuration
[0451] The system consists of the following main components:
[0452] 1. Device that acquires audio data
[0453] 2. A server that converts the acquired voice data into text data
[0454] 3. Server that translates text data into multiple languages
[0455] 4. Server that performs emotion analysis of text data and voice data
[0456] 5. Server that controls automatic or manual operation based on the results of emotion analysis
[0457] 6. Server that generates advice and suggestions in real time based on the analysis results and the situation
[0458] 7. A server that stores all data and uses it for later analysis
[0459] System Operation
[0460] Acquiring audio data
[0461] Voice information from the user (pilot or air traffic controller) is acquired through a microphone on board the aircraft or through the voice system in the air traffic control room. The terminal captures this voice data and transmits it.
[0462] Voice Recognition
[0463] When the voice data arrives at the server, the voice recognition engine activates and converts it into text data. For example, a voice saying "There is an engine problem, please check!" is converted into the corresponding text.
[0464] Text data translation
[0465] The server then identifies the language of this text data and, if necessary, translates it into the appropriate language using a built-in translation engine. For example, a Japanese sentence like "Engine has a malfunction, please check!" is translated into English as "Engine has a malfunction, please check!"
[0466] Emotion analysis
[0467] The translated text data and the original audio data are analyzed by a server, which uses natural language processing and speech analysis techniques to assess the speaker's emotional state, for example, determining whether they are stressed or calm in an emergency.
[0468] Providing operational control and advice
[0469] The server controls autopilot or manual operation based on the emotion analysis results and text content. If it determines that the risk is high, it sends a signal to the aircraft's autopilot system to perform appropriate operations. It also generates advice and suggestions based on the situation in real time and provides them to the user via the terminal. For example, if an engine abnormality is detected, it immediately provides a voice message advising, "Please perform the engine restart procedure."
[0470] Specific examples
[0471] Emergency Scenarios
[0472] 1. The user (pilot) says, "There's an engine malfunction, please check!"
[0473] 2. The device captures this audio and sends it to the server.
[0474] 3. The server performs speech recognition and converts it into text.
[0475] 4. If necessary, the server translates the text into English.
[0476] 5. The server performs emotion analysis and detects that the pilot is in a high stress state.
[0477] 6. The server signals the autopilot system to execute a safety maneuver.
[0478] 7. The server generates an advice saying, "Autopilot has been strengthened due to an engine abnormality. Please check the engine instruments and perform the engine restart procedure." and notifies the user via the terminal.
[0479] 8. All processed data is stored on the server and used for later analysis.
[0480] This system can significantly improve aircraft safety and minimize communication errors caused by language differences or mental state during an emergency.
[0481] The processing flow will be explained below.
[0482] Step 1:
[0483] The device captures the voice data of the pilot and air traffic controller through microphones on the plane or the audio system in the air traffic control room, and transmits the captured voice data to a server in real time.
[0484] Step 2:
[0485] The server receives the voice data. It converts the voice data into text data using a speech recognition engine. For example, a voice message such as "There is an engine error. Please check!" is converted into text format.
[0486] Step 3:
[0487] The server identifies the language of the converted text data, using a built-in language identification engine to determine in which language the text is written.
[0488] Step 4:
[0489] The server translates the text data into the specified language as needed. For example, the Japanese text "Engine has a malfunction, please check!" is translated into English as "Engine has a malfunction, please check!".
[0490] Step 5:
[0491] The server performs sentiment analysis based on the translated text data and the original audio data, using natural language processing and audio analysis techniques to assess the speaker's emotional state (e.g., stressed, overwhelmed, calm, etc.).
[0492] Step 6:
[0493] The server assesses the risk level based on the results of emotion analysis and the text content. If the risk level is deemed high, the server sends a signal to the autopilot system to temporarily take control of the aircraft.
[0494] Step 7:
[0495] The server generates optimal countermeasures in real time based on the analysis results and the situation. This includes specific procedures and advice. For example, advice such as "Perform the engine restart procedure" is generated.
[0496] Step 8:
[0497] The terminal then notifies the pilot or controller of the generated advice in real time as a voice or text message.
[0498] Step 9:
[0499] The server stores all voice data, text data, translation data, sentiment analysis results, and response actions as logs, which will be used later for further analysis and system improvement.
[0500] Example 1
[0501] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0502] In communication between aircraft pilots and air traffic controllers, errors in information transmission can occur due to language differences or the pilot's mental state during an emergency. Such errors can have a significant impact on aircraft operation and pose a risk of compromising safety. It is also necessary to accurately grasp the pilot's stress level and emotions and provide appropriate responses in real time, but current systems are unable to adequately achieve this.
[0503] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0504] In this invention, the server includes means for acquiring voice data, means for converting the acquired voice data into text data, means for translating the text data into multiple languages, means for performing emotion analysis of the text data and voice data, means for controlling automatic or manual operation based on the emotion analysis results, means for generating advice and suggestions in real time based on the analysis results and the situation, means for storing all data for later analysis, means for transmitting the acquired voice data in digital form, means for converting the voice data into text data using a voice recognition engine, means for passing the voice data and text data to an emotion analysis engine, means for evaluating the emotional state using the emotion analysis engine, means for sending signals to an automatic operation system, and means for generating and transmitting voice messages based on the analysis results. This makes it possible to minimize communication errors due to language differences or mental states during an emergency, thereby significantly improving aircraft safety.
[0505] "Voice data" refers to data in which voice information uttered by a user is recorded in digital format.
[0506] "Text data" is character string information converted from voice data using a voice recognition engine.
[0507] A "translation engine" is a system that converts text data from an identified language to another language.
[0508] An "emotion analysis engine" is a system that analyzes text and audio data to evaluate the speaker's emotional state.
[0509] An "automatic operation system" is a system that performs automatic piloting and other automatic control of an aircraft based on the results of analysis.
[0510] "Digital format" refers to a format in which analog data is converted into numerical data.
[0511] A "voice recognition engine" is software or hardware for converting voice data into text data.
[0512] A "signal" is an electronic signal for transmitting control information.
[0513] "Real-time" refers to data being processed and analyzed as soon as it is acquired.
[0514] "Analysis results" are the analysis results derived by the sentiment analysis engine and other analysis means.
[0515] "Advice and suggestions" are measures and instructions generated based on the analysis results and the situation.
[0516] "Storage" refers to keeping acquired data and analysis results for a certain period of time.
[0517] "Later analysis" refers to re-analysis at a later date using the stored data.
[0518] This invention is a system for analyzing conversations between aircraft pilots and air traffic controllers in real time to control aircraft operations and ensure safety. The system captures voice data, converts it into text data, translates it into multiple languages, performs sentiment analysis, controls automatic or manual operations based on the analysis results, and generates advice and suggestions in real time as needed. Furthermore, all data is stored for later analysis.
[0519] System configuration and operation
[0520] Acquiring audio data
[0521] The user (pilot or air traffic controller) reports the emergency situation by voice, and this voice data is picked up by the terminal in digital format through the microphone on board or the voice system in the air traffic control room, and then transmitted to the server.
[0522] Voice Recognition
[0523] The server passes the received voice data to a voice recognition engine. This engine generally uses high-performance voice recognition technology such as Google Speech-to-Text API or IBM Watson Speech to Text. The voice data is converted into text data. For example, a voice saying "There is an engine problem, please check!" is converted into text "There is an engine problem, please check!"
[0524] Text data translation
[0525] The server then automatically identifies the language of the text data. If the identified language is other than English, the server uses a built-in translation engine (e.g., Microsoft Translator or Google Translate API) to translate the text data into English. For example, the Japanese input "Engine has a malfunction, please check!" is translated into "Engine has a malfunction, please check!"
[0526] Emotion analysis
[0527] The server passes the translated text data and the original audio data to a sentiment analysis engine, which uses natural language processing and speech analysis technologies such as IBM Watson Natural Language Understanding and Microsoft Azure Text Analytics to assess the speaker's emotional state, for example, determining whether they are highly stressed or calm in an emergency.
[0528] Providing operational control and advice
[0529] The server sends a signal to the aircraft's automatic operation system based on the results of the emotion analysis and the content of the text data. If a high level of danger is detected, the server immediately executes safety operations. It also generates advice and suggestions based on the situation in real time and provides them to the user via the terminal. For example, advice such as "Perform engine restart procedures" is immediately provided as a voice message.
[0530] Data storage
[0531] All processed data is stored on a server as a detailed log, which is used for future analysis, system improvement, aircraft accident investigations, and operational optimization.
[0532] Specific examples
[0533] Emergency Scenarios
[0534] 1. The user (pilot) reports audibly, "An engine malfunction has occurred. Please check!"
[0535] 2. The device captures this audio and immediately transmits it digitally to the server.
[0536] 3. The server passes the received voice data to a voice recognition engine and converts it into text data.
[0537] 4. The server identifies the language of the text data and translates it into English using a translation engine if necessary.
[0538] 5. The server analyzes the translated text data and the original audio data using an emotion analysis engine to detect high stress levels.
[0539] 6. The server sends a signal to the automated operation system to immediately execute safety operations, and also creates a situation-specific advice message, such as "Perform engine restart procedure," as a voice message and provides it to the user via the terminal.
[0540] 7. All processing data will be stored on the server and used for future analysis and improvement.
[0541] Prompt Sentence Examples
[0542] Prompt: "Analyze the audio of an aircraft pilot reporting an emergency and provide a detailed description of how the automated flight control system responds and provides advice to the pilot."
[0543] The use of this system can significantly improve aviation safety and minimize communication errors caused by language differences or mental state during an emergency.
[0544] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0545] The flow of this system's program processing
[0546] Step 1:
[0547] The user (pilot or air traffic controller) reports information verbally. For example, in an emergency situation, they might say, "There's an engine malfunction, please check!" This is the input.
[0548] Step 2:
[0549] The terminal captures the user's voice in digital form through an onboard microphone or an audio system in the air traffic control room, and transmits this digital voice data to a server.
[0550] Input: User's voice
[0551] Output: Digital audio data
[0552] Step 3:
[0553] The server passes the received digital voice data to a voice recognition engine (e.g., Google Speech-to-Text API or IBM Watson Speech to Text) and converts it into text data.
[0554] Input: Digital audio data
[0555] Output: Text data: "An engine error has occurred. Please check!"
[0556] Step 4:
[0557] Identifying the language of server-generated text data, for example automatically recognizing that the text is in Japanese.
[0558] Input: Text data
[0559] Output: Language identification result (Japanese)
[0560] Step 5:
[0561] The server uses a built-in translation engine (e.g., Microsoft Translator or Google Translate API) to translate from the identified language to English. For example, "Engine has a malfunction, please check!" is translated to "Engine has a malfunction, please check!"
[0562] Input: Japanese text data
[0563] Output: English text data
[0564] Step 6:
[0565] The server passes the translated English text data and the original audio data to a sentiment analysis engine (e.g., IBM Watson Natural Language Understanding or Microsoft Azure Text Analytics) to evaluate the speaker's emotional state, e.g., to determine whether the speaker has a high stress level.
[0566] Input: English text and audio data
[0567] Output: Emotion analysis results (high stress)
[0568] Step 7:
[0569] Based on the results of the emotion analysis and the content of the text data, the server sends a signal to the aircraft's automatic operation system to perform appropriate safety operations, such as sending a signal to "strengthen autopilot."
[0570] Input: Sentiment analysis results and text data
[0571] Output: Control signal to automatic operation system
[0572] Step 8:
[0573] The server generates advice and suggestions based on the situation in real time and provides them to the user through the terminal, for example, generating a voice message saying "Please perform the engine restart procedure."
[0574] Input: Sentiment analysis results and situation information
[0575] Output: Advice and suggestions (voice message)
[0576] Step 9:
[0577] The server stores all transaction data in a detailed log, which is used for future analysis and system improvement.
[0578] Input: All analysis results and process logs
[0579] Output: Saved processed data
[0580] (Application example 1)
[0581] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0582] In on-site security work, security personnel often experience high levels of stress during emergencies, which can result in delayed responses and mistakes. Furthermore, communication barriers in multilingual environments and difficulty assessing situations in real time can also be a factor, potentially worsening dangerous situations.
[0583] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0584] In this invention, the server includes means for acquiring voice data, means for converting the acquired voice data into text data, means for translating the text data into multiple languages, means for performing emotion analysis of the text data and voice data, means for controlling automatic or manual operation based on the emotion analysis results, means for generating advice and suggestions in real time based on the analysis results and the situation, means for alerting the user when a high-stress state is detected, and means for saving all data for later analysis, thereby enabling quick and appropriate decisions to be made even in emergencies or high-stress situations.
[0585] The "means for acquiring voice data" refers to a device or software that captures the voice uttered by the user in real time and transmits the voice data to the system.
[0586] The "means for converting acquired voice data into text data" refers to a device or software that uses voice recognition technology to convert acquired voice data into a corresponding string of characters.
[0587] A "means for translating text data into multiple languages" is a translation engine or algorithm for converting text data into another language.
[0588] "Means for analyzing emotions in text data and voice data" refers to devices or software that use natural language processing and voice analysis techniques to evaluate the emotions of a speaker from text data or voice data.
[0589] "Means for controlling automatic or manual operations based on the results of emotion analysis" refers to devices or software that appropriately manage or correct automatic operations or manual instructions within the system based on the results of emotion analysis.
[0590] "Means for generating advice and suggestions in real time based on analysis results and the current situation" refers to devices or software that generate appropriate advice and suggestions in a timely manner based on data analysis results and the current situation.
[0591] "Means for alerting the user when a high stress state is detected" refers to devices or software that alert the user by issuing appropriate alerts or notifications when it is determined that the user's stress level is high.
[0592] "Means for storing all data and making it available for later analysis" refers to devices or software that store all data processed by the system and make it available for later analysis, performance evaluation, etc.
[0593] The present invention is a real-time assistant system for supporting security operations. This system has the functions of capturing voice data, converting it into text data, translating it into multiple languages, performing emotion analysis, providing advice in real time, and detecting high-stress situations.
[0594] System configuration
[0595] The system consists of the following main components:
[0596] 1. A means of acquiring voice data: Capture the user's voice in real time, specifically using a smartphone or microphone.
[0597] 2. Means for converting acquired voice data into text data: Convert the voice into text using voice recognition technology (for example, the SpeechRecognition library).
[0598] 3. Means for translating text data into multiple languages: Translate text data into other languages using a translation engine (e.g., Google Translate API).
[0599] 4. Means for sentiment analysis of text and voice data: Sentiment analysis is performed using natural language processing technology and voice analysis technology (e.g., Hugging Face Transformers).
[0600] 5. Means for controlling automatic or manual operations based on the results of sentiment analysis: Based on the results of sentiment analysis, operations within the system are appropriately managed or corrected.
[0601] 6. Means of generating advice and suggestions in real time based on analysis results and the current situation: Generate appropriate advice and suggestions based on the results of data analysis and the current situation.
[0602] 7. A method to alert users when high stress levels are detected: When stress levels are determined to be high, appropriate alerts and notifications will be provided.
[0603] 8. Means of storing all data for later analysis: All processed data will be stored and used for later analysis and performance evaluation.
[0604] System Operation
[0605] The specific operation of the system can be explained as follows:
[0606] Acquiring voice data and converting it to text data
[0607] The voice spoken by the user (security officer) is picked up through the smartphone's microphone. This voice data is converted into text using voice recognition technology. For example, a voice saying "Someone is breaking in! Call for help!" is converted into text data.
[0608] Text data translation
[0609] This text data is translated from Japanese to English using a language recognition function. Using a translation engine, "Someone is trespassing, please call for help immediately!" is converted to "Someone is trespassing, please call for help immediately!"
[0610] Emotion analysis
[0611] The translated text data and the original audio data are analyzed by a sentiment analysis model to determine the user's emergency state or high stress level. If a specific emotional state is detected, appropriate advice is generated.
[0612] Real-time advice and suggestions
[0613] Based on the results of sentiment analysis and the text content, the system generates real-time advice. For example, if a high stress state is detected, the system will provide advice such as "Be careful. Stress is increasing."
[0614] Data storage
[0615] All data is logged and made available for later analysis or performance evaluation.
[0616] Specific examples
[0617] For example, if a security worker says, "Someone is trespassing, please call for help immediately!", the system converts the speech into text and translates it into "Someone is trespassing, please call for help immediately!". If the emotion analysis determines that this is an "emergency situation," the system provides a real-time voice message with advice such as, "Be careful. Stress is building."
[0618] Prompt Sentence Examples
[0619] Voice: "Someone's breaking in, hurry up and call for help!"
[0620] Translation: "Someone is trespassing, please call for help immediately!"
[0621] Sentiment Analysis: "Emergency"
[0622] Advice: "Pay attention. Stress is building."
[0623] This allows security personnel to respond quickly and appropriately.
[0624] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0625] Step 1:
[0626] The device captures the user's voice in real time through a microphone. The input is the user's speech, and the output is digital voice data. This voice data is sent directly to the server.
[0627] Step 2:
[0628] The server converts the acquired voice data into text data using a voice recognition engine (for example, the SpeechRecognition library). The input is digital voice data, and the output is the corresponding text data. This conversion makes it easier to handle voice information as text information.
[0629] Step 3:
[0630] The server sends the generated text data to a translation engine (for example, Google Translate API) to translate it into another language. The input is text data in the original language, and the output is the translated text data. This translation allows for smooth support in multilingual environments.
[0631] Step 4:
[0632] The server performs sentiment analysis using the translated text data and the original voice data. Natural language processing techniques (e.g., Hugging Face Transformers) are used for sentiment analysis. The input is text and voice data, and the output is the analysis result of the emotional state. Based on the analysis result, the user's state of urgency and stress level are evaluated.
[0633] Step 5:
[0634] The server controls automatic or manual operation based on the results of emotion analysis. Specifically, if a high-stress state is detected, it automatically changes the system operation or issues instructions to the user. The input to this step is the emotion analysis result, and the output is appropriate operation control instructions.
[0635] Step 6:
[0636] The server generates advice and suggestions in real time based on the analysis results and the current situation. For example, if a high stress state is detected, advice such as "Be careful. Your stress is increasing" is generated. The input is the analysis results and situation information, and the output is the generated advice.
[0637] Step 7:
[0638] The server sends the generated advice to the terminal and notifies the user. The input is the generated advice, and the output is the notification received by the user.
[0639] Step 8:
[0640] All data is stored by the server. The stored data is used for later analysis and performance evaluation. The input is all the data during and after processing, and the output is the stored data. This storage allows for future optimization and improvement.
[0641] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0642] This invention is a system for analyzing conversations between aircraft pilots and air traffic controllers in real time, controlling aircraft operations, and ensuring safety. The system incorporates an emotion engine for recognizing the user's emotional state, thereby further improving safety. The following is a detailed description of an embodiment of the system.
[0643] System configuration
[0644] The system consists of the following main components:
[0645] 1. Device that acquires audio data
[0646] 2. A server that converts the acquired voice data into text data
[0647] 3. Server that translates text data into multiple languages
[0648] 4. Server that performs emotion analysis of text data and voice data
[0649] 5. Emotion Engine
[0650] 6. Server that controls automatic or manual operation based on the results of emotion analysis
[0651] 7. Server that generates advice and suggestions in real time based on the analysis results and the situation
[0652] 8. A server that stores all data and uses it for later analysis
[0653] System Operation
[0654] Acquiring audio data
[0655] Voice information from the user (pilot or air traffic controller) is acquired through a microphone on board the aircraft or through the voice system in the air traffic control room. The terminal captures this voice data and transmits it.
[0656] Voice Recognition
[0657] When the voice data arrives at the server, the voice recognition engine activates and converts it into text data. For example, a voice message such as "There is an engine error. Please check!" is converted into text format.
[0658] Text data translation
[0659] The server then identifies the language of this text data and, if necessary, translates it into the appropriate language using a built-in translation engine. For example, a Japanese sentence like "Engine has a malfunction, please check!" is translated into English as "Engine has a malfunction, please check!"
[0660] Emotion analysis
[0661] The translated text data and the original audio data are analyzed by a server, which uses natural language processing and speech analysis techniques to assess the speaker's emotional state, for example, determining whether they are highly stressed or calm in an emergency.
[0662] Use of emotion engine
[0663] The server is equipped with an emotion engine that can recognize the user's emotional state in more detail. This emotion engine analyzes both text and voice data and has the ability to highly evaluate the user's emotional state. This allows the system to determine the level of urgency corresponding to a specific emotional state and propose appropriate countermeasures accordingly.
[0664] Providing operational control and advice
[0665] The server controls autopilot or manual operation based on the emotion analysis results and text content. If it determines that the risk is high, it sends a signal to the aircraft's autopilot system to perform appropriate operations. It also generates advice and suggestions based on the situation in real time and provides them to the user via the terminal. For example, if an engine abnormality is detected, it immediately provides a voice message advising, "Please perform the engine restart procedure."
[0666] Data storage
[0667] The server stores all processing data, analysis results, audio data, and response measures as logs, which will be used later for detailed analysis and system improvement.
[0668] Specific examples
[0669] Emergency Scenarios
[0670] 1. The user (pilot) says, "There's an engine malfunction, please check!"
[0671] 2. The device captures this audio and sends it to the server.
[0672] 3. The server performs speech recognition and converts it into text.
[0673] 4. If necessary, the server translates the text into English.
[0674] 5. The server performs emotion analysis and detects that the pilot is in a high stress state.
[0675] 6. The emotion engine analyzes the user's stress level in detail and determines that the level of urgency is high.
[0676] 7. The server signals the autopilot system to execute a safety maneuver.
[0677] 8. The server generates an advice saying, "Autopilot has been strengthened due to an engine abnormality. Please check the engine instruments and perform the engine restart procedure." and notifies the user via the terminal.
[0678] 9. All processed data is stored on the server and used for later analysis.
[0679] The system can significantly improve aviation safety and minimize communication errors by providing countermeasures based on the user's emotional state.
[0680] The processing flow will be explained below.
[0681] Step 1:
[0682] The device captures the voice data of the pilot and air traffic controller through microphones on the plane or the audio system in the air traffic control room, and transmits the captured voice data to a server in real time.
[0683] Step 2:
[0684] The server receives the voice data. It converts the voice data into text data using a speech recognition engine. For example, a voice message such as "There is an engine error. Please check!" is converted into text format.
[0685] Step 3:
[0686] The server identifies the language of the converted text data, using a built-in language identification engine to determine in which language the text is written.
[0687] Step 4:
[0688] The server translates the text data into the specified language as needed. For example, the Japanese text "Engine has a malfunction, please check!" is translated into English as "Engine has a malfunction, please check!".
[0689] Step 5:
[0690] The server performs sentiment analysis based on the translated text data and the original audio data, using natural language processing and audio analysis technologies to assess the speaker's emotional state (e.g., stressed, overwhelmed, calm, etc.).
[0691] Step 6:
[0692] The emotion engine performs detailed analysis of text and voice data to intelligently assess the user's emotional state, for example detecting when a pilot is under high stress.
[0693] Step 7:
[0694] The server assesses the level of danger based on the results of sentiment analysis and the text content. If the level of danger is deemed high, the server sends a signal to the autopilot system to temporarily control the aircraft's operation. For example, if a pilot is under high stress, the autopilot system will automatically maintain the aircraft's altitude and speed.
[0695] Step 8:
[0696] The server generates optimal countermeasures in real time based on the analysis results and the situation. This includes specific procedures and advice. For example, advice such as "Perform the engine restart procedure" is generated.
[0697] Step 9:
[0698] The terminal then notifies the pilot or controller of the generated advice in real time as a voice or text message.
[0699] Step 10:
[0700] The server stores all voice data, text data, translation data, sentiment analysis results, and response actions as logs, which will be used later for further analysis and system improvement.
[0701] Example 2
[0702] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0703] In communications between aircraft pilots and air traffic controllers, differences in emotional state and language during emergencies can make it difficult to understand instructions, delaying appropriate responses. To solve this problem, real-time voice analysis, evaluation of emotional state, and proposal of appropriate countermeasures are required, but previous systems have found it difficult to perform these tasks in an integrated manner.
[0704] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for acquiring voice data, means for converting the acquired voice data into text data, means for translating the text data into multiple languages, means for performing emotion analysis of the text data and voice data, means for controlling automatic operation or manual operation based on the emotion analysis results, means for generating advice and suggestions in real time based on the analysis results and the situation, means for saving all data and using it for later analysis, and means for incorporating an emotion engine that highly evaluates the user's emotional state. This makes it possible to provide prompt measures based on accurate voice recognition and emotion analysis in real time.
[0705] "Voice data" refers to voices uttered by a user captured as digital signals.
[0706] "Means for acquiring" refers to the technology or device for acquiring voice data using a terminal or sensor within the domain.
[0707] "Text data" is voice data that has been analyzed and expressed as text information.
[0708] "Means for converting into text data" refers to technology or devices that convert acquired voice data into text format using natural language processing technology.
[0709] "Translation means" refers to the technology and software used to convert text data into other languages.
[0710] "Sentiment analysis" is the process of analyzing audio and text data to assess the emotional state of a speaker.
[0711] "Means for performing emotion analysis" refers to technologies or devices that use the algorithms or software required for emotion analysis.
[0712] An "emotion engine" is specialized software or algorithms that intelligently assess a user's emotional state.
[0713] "Means for controlling autopilot or manual operation" refers to technology or devices that automatically control aircraft operation based on the results of emotion analysis or in accordance with the pilot's instructions.
[0714] The "means for generating advice or suggestions" refers to technology or software that generates suggestions in real time to help the user take appropriate actions or make appropriate decisions based on the analysis results and the situation.
[0715] "Means for storing data and using it for later analysis" refers to the technology and devices that record all processing data and use it for later data analysis and system improvement.
[0716] This invention is a system for analyzing conversations between aircraft pilots and air traffic controllers in real time to control aircraft operations and ensure safety. The system uses multiple servers, terminals, and software to achieve functions such as voice data acquisition, voice recognition, text data translation, emotion analysis, autopilot or manual operation control, advice and suggestion generation, and data storage.
[0717] System configuration
[0718] The system consists of the following main components:
[0719] 1. Device that acquires audio data
[0720] 2. Server that converts voice data into text data
[0721] 3. Server that translates text data into multiple languages
[0722] 4. Server that performs emotion analysis of text data and voice data
[0723] 5. Emotion Engine
[0724] 6. Server that controls automatic or manual operation based on the results of emotion analysis
[0725] 7. Server that generates advice and suggestions in real time based on the analysis results and the situation
[0726] 8. A server that stores all data and uses it for later analysis
[0727] Hardware and software used
[0728] 1. Terminal: The aircraft's microphone and communication system are used to acquire voice data. Standard data communication protocols are used to transmit the voice data.
[0729] 2. Server: To analyze the voice data, a speech recognition engine such as Google Cloud Speech-to-Text or Amazon Transcribe is used.
[0730] 3. Text data translation: Translate text data into multiple languages using the Google Translate API, etc.
[0731] 4. Sentiment Analysis: Evaluate emotional state using IBM Watson Natural Language Understanding and Amazon's speech analysis technology.
[0732] 5. Emotional Engine: Utilizing specialized engines such as Affectiva to provide a sophisticated assessment of the user's emotional state.
[0733] 6. Autopilot and manual operation control: Works in conjunction with the aircraft's autopilot system (e.g., Honeywell system) to ensure safe operation.
[0734] 7. Advice and suggestion generation: Use dedicated advice generation algorithms to generate context-sensitive advice in real time.
[0735] 8. Data storage: All processed data, analysis results, audio data, and response actions will be stored using cloud storage for later analysis.
[0736] Specific examples
[0737] Emergency scenarios:
[0738] 1. The user (pilot) says, "There's an engine malfunction, please check!"
[0739] 2. The onboard microphone captures the audio and sends the data to the device.
[0740] 3. The device sends the audio data to the server.
[0741] 4. The server receives the audio data and calls the Google Cloud Speech-to-Text service.
[0742] 5. The server converts the voice data into text and generates the text "An engine error has occurred, please check!"
[0743] 6. The server translates the text data into English using the Google Translate API. "Engine has a malfunction, please check!" becomes "Engine has a malfunction, please check!"
[0744] 7. The server sends the text data to IBM Watson Natural Language Understanding for sentiment analysis.
[0745] 8. The server sends the audio data to Amazon Transcribe, which analyzes the emotional state of the audio.
[0746] 9. The emotion engine on the server integrates the text and voice data to perform a detailed emotion assessment.
[0747] 10. The emotion engine determines that the user's stress level is "high."
[0748] 11. The server sends a signal to the autopilot system, generating the message "Perform engine restart procedure."
[0749] 12. The device will announce this message to the user audibly.
[0750] 13. The server records the processed data, analysis results, audio data, and response measures as logs and stores them in a database.
[0751] 14. The stored data can be later analyzed to help improve the system.
[0752] Prompt Sentence Examples
[0753] "Please explain the steps your system takes when a user says, 'There's an engine malfunction, please check!'"
[0754] In this way, the system significantly improves aircraft safety, can provide countermeasures based on the user's emotional state, and minimizes communication errors.
[0755] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0756] Step 1:
[0757] The voice of the user (pilot or air traffic controller) is acquired through an on-board microphone or the voice system in the air traffic control room. This captures the user's real-time voice information as digital voice data. The input is the user's voice, and the output is digital voice data.
[0758] Step 2:
[0759] The device sends the captured audio data to a server. The audio data is transferred to the server via the Internet or a dedicated communication line. The input is the captured digital audio data, and the output is the transferred audio data.
[0760] Step 3:
[0761] The server receives the voice data and calls the Google Cloud Speech-to-Text service to perform speech recognition. The voice data is analyzed and converted into text data. For example, a speech saying "There is an engine problem, please check!" is converted into text "There is an engine problem, please check!" The input is voice data, and the output is text data.
[0762] Step 4:
[0763] The server identifies the language of the generated text data. If necessary, it uses the Google Translate API to translate the text into other languages. For example, text entered in Japanese is translated into English. The input is the text data, and the output is the translated text data.
[0764] Step 5:
[0765] The server sends the text data to IBM Watson Natural Language Understanding for sentiment analysis. The text data is used to evaluate the speaker's emotional state. For example, the text content can be used to analyze a pilot's stress level and urgency. The input is text data, and the output is an evaluation of the emotional state.
[0766] Step 6:
[0767] The server sends the audio data to Amazon Transcribe, which analyzes the tone and pitch of the voice. This allows for a detailed assessment of the speaker's emotional state and urgency from the audio data. The input is the audio data, and the output is the emotional assessment result from the audio analysis.
[0768] Step 7:
[0769] The emotion engine in the server integrates the results of the analysis of the text data and voice data to perform a detailed emotion evaluation. The emotion engine highly evaluates the user's emotional state and determines, for example, that the user is in a high-stress state. The inputs are the emotion analysis results and the voice analysis results, and the output is the integrated emotion evaluation result.
[0770] Step 8:
[0771] The server generates a control signal to an autopilot system or manual operation based on the emotion analysis result, for example, to signal an aircraft's autopilot system to execute emergency measures. The input is the integrated emotion evaluation result, and the output is a control signal.
[0772] Step 9:
[0773] The server generates appropriate advice or suggestions in real time based on the analysis results and the situation, and notifies the user through the terminal. For example, it generates advice such as "Please perform the engine restart procedure." The input is the emotion analysis results and situation information, and the output is the generated advice.
[0774] Step 10:
[0775] The server records and stores all processing data, analysis results, audio data, and response measures as logs. The stored data is used for later detailed analysis and system improvement. The input is the processing data and analysis results, and the output is the stored log data.
[0776] (Application example 2)
[0777] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0778] Conventional call center systems have the problem of being unable to grasp the emotional state of the operator and customer in real time and provide immediate solutions. This has led to problems such as customers' dissatisfaction not being resolved properly and operators feeling high levels of stress. This leads to problems such as lower customer satisfaction and a decline in the work efficiency of operators.
[0779] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0780] In this invention, the server includes means for acquiring voice data, means for converting the acquired voice data into text data, means for translating the text data into multiple languages, means for performing emotion analysis of the text data and voice data, means for controlling automatic or manual operations based on the emotion analysis results, means for generating advice and suggestions in real time based on the analysis results and the situation, means for triage and proposing countermeasures based on the emotional state, and means for saving all data and using it for later analysis. This makes it possible to grasp the emotional states of customers and operators in real time and immediately provide appropriate countermeasures.
[0781] "Audio data" refers to sound information in digital or analog form collected through audio equipment such as a microphone.
[0782] "Text data" is information obtained by converting voice data into a character string format.
[0783] A "multilingual translation method" is a technique or process for converting text data into different languages.
[0784] "Emotion analysis" is a technique that analyzes audio and text data to assess a speaker's emotional state.
[0785] "Autonomous operation" means that the system performs a specific action without human intervention.
[0786] "Manual operation" means direct human involvement in operating a system.
[0787] "Means for generating advice and suggestions in real time" refers to techniques and methods that provide immediate advice to operators or users based on the current situation.
[0788] "Triage" is the prioritization of response measures based on urgency and importance.
[0789] A "solution" is an action or suggestion to be taken in response to a particular situation.
[0790] "Storing data" means recording collected data on a storage medium so that it can be used later.
[0791] "Use for later analysis" means to conduct analysis based on the stored data for future consideration and improvement.
[0792] MODE FOR CARRYING OUT THE INVENTION
[0793] The present invention is a system for supporting appropriate responses by analyzing the emotional states of customers and agents in a call center in real time. An embodiment of this system will be described in detail below.
[0794] System configuration
[0795] The system consists of the following main components:
[0796] 1. Means of capturing voice data: microphones and telephone systems to capture conversations between customers and operators.
[0797] 2. A means of converting the acquired voice data into text data: a speech recognition engine that converts the voice into a string format (e.g., Google Speech-to-Text API).
[0798] 3. A means of translating text data into multiple languages: A system that translates text into the required languages (e.g., Microsoft Azure Cognitive Services).
[0799] 4. Means of sentiment analysis of text and voice data: Software that analyzes the emotional state of customers and operators (e.g., IBM Watson Tone Analyzer).
[0800] 5. Means for controlling automatic or manual operations based on the results of emotion analysis: A control system for performing processing according to emotional states.
[0801] 6. Means of generating advice and suggestions in real time based on analysis results and situations: An engine that generates real time advice and suggestions to operators.
[0802] 7. Triage and suggest responses based on emotional state: A mechanism for prioritizing responses based on urgency and importance.
[0803] 8. A means of storing all data and making it available for later analysis: a database to store the collected and analyzed data.
[0804] System Operation
[0805] Audio Acquisition
[0806] Voice information from users (customers and operators) is acquired through a microphone or telephone system. The terminal captures this voice data and sends it to the server.
[0807] Voice Recognition
[0808] The voice data received by the server is converted into text data by a voice recognition engine (Google Speech-to-Text API).
[0809] Text data translation
[0810] The server then converts this text data into the appropriate language if it needs to be translated using a multi-language translation means.
[0811] Emotion analysis
[0812] The translated text data and the original audio data are then subjected to natural language processing and speech analysis techniques to assess the user's emotional state, using analysis engines such as IBM Watson Tone Analyzer.
[0813] Detailed analysis by Emotion Engine
[0814] The emotion engine performs advanced analysis of text and voice data to determine specific emotional states and their urgency.
[0815] Providing operational control and advice
[0816] Based on the results of the emotion analysis, the server controls automatic or manual operations and generates and notifies advice and suggestions in real time. Appropriate advice is provided to the operator in voice or text format.
[0817] Data storage
[0818] The server stores all audio data, text data, and analysis results in a database for later analysis.
[0819] Specific examples
[0820] Emergency Scenarios
[0821] 1. A customer says, "I recently purchased a broken item and I'm very unhappy with it. I'd like to process the return immediately."
[0822] 2. The device captures this audio and sends it to the server.
[0823] 3. The server performs speech recognition and converts it into text.
[0824] 4. If necessary, the server translates the text.
[0825] 5. The server performs sentiment analysis and detects that the customer is in a high stress state.
[0826] 6. The emotion engine analyzes the customer's stress level in detail and determines that the situation is urgent.
[0827] 7. The server advises the operator, "The customer is unhappy. Please stay calm and explain with specific examples."
[0828] 8. All processed data is stored on the server and used for later analysis.
[0829] This system will significantly improve the quality and efficiency of customer service at call centers, increasing customer satisfaction. It will also reduce stress on operators and contribute to improved work efficiency.
[0830] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0831] Step 1:
[0832] The user (customer and operator) starts a conversation. The terminal captures the voice data through a microphone or telephone system. This captured voice data is recorded in digital format and sent to the server. The input here is the raw voice data, and the output is the audio file sent to the server.
[0833] Step 2:
[0834] The server sends the received voice data to the Google Speech-to-Text API for speech recognition. In this process, the voice data is converted into text data in string format. The input is the voice data, and the output is the converted text data. Specifically, the voice data undergoes signal processing and pattern matching.
[0835] Step 3:
[0836] The server uses Microsoft Azure Cognitive Services to translate the acquired text data into multiple languages as needed. In this step, the language of the text data is identified and converted into the appropriate language. The input is text data, and the output is translated text data. The specific operation utilizes natural language processing algorithms.
[0837] Step 4:
[0838] The server then sends the translated text data and the original audio data to the IBM Watson Tone Analyzer for sentiment analysis. The input is text data and audio data, and the output is the sentiment analysis results. This function evaluates the speaker's emotional state based on the tone and content of the audio.
[0839] Step 5:
[0840] The server further evaluates the emotion analysis results to determine the specific emotional state and its urgency. This is done by the emotion engine, which understands the user's emotional state. The input is the emotion analysis results, and the output is detailed emotional state and urgency information. Specific operations include integrating and evaluating the analysis results.
[0841] Step 6:
[0842] The server generates and notifies the operator in real time with advice and suggestions based on the results of emotion analysis. The input is the detailed emotional state and the customer's comments, and the output is advice and suggestions provided to the operator. Specifically, the system generates appropriate countermeasures and notifies the operator via voice or text.
[0843] Step 7:
[0844] The server stores all data in a database for later analysis. The input is all processed data (audio, text, analysis results, etc.), and the output is the stored data. Specific operations include recording and storing data.
[0845] Prompt Sentence Examples
[0846] Speech input: "I'm very unhappy with the product I recently purchased, which is broken. I'd like to process the return immediately."
[0847] Expected output: "The customer is unhappy. Please stay calm and provide specific examples. We'll expedite the return process."
[0848] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0849] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0850] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0851] [Third embodiment]
[0852] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0853] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[0854] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0855] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0856] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0857] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0858] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0859] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0860] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0861] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0862] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0863] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0864] The present invention provides a system for analyzing conversations between aircraft pilots and air traffic controllers in real time, controlling aircraft operations, and ensuring safety. The following is a detailed description of an embodiment of the system.
[0865] System configuration
[0866] The system consists of the following main components:
[0867] 1. Device that acquires audio data
[0868] 2. A server that converts the acquired voice data into text data
[0869] 3. Server that translates text data into multiple languages
[0870] 4. Server that performs emotion analysis of text data and voice data
[0871] 5. Server that controls automatic or manual operation based on the results of emotion analysis
[0872] 6. Server that generates advice and suggestions in real time based on the analysis results and the situation
[0873] 7. A server that stores all data and uses it for later analysis
[0874] System Operation
[0875] Acquiring audio data
[0876] Voice information from the user (pilot or air traffic controller) is acquired through a microphone on board the aircraft or through the voice system in the air traffic control room. The terminal captures this voice data and transmits it.
[0877] Voice Recognition
[0878] When the voice data arrives at the server, the voice recognition engine activates and converts it into text data. For example, a voice saying "There is an engine problem, please check!" is converted into the corresponding text.
[0879] Text data translation
[0880] The server then identifies the language of this text data and, if necessary, translates it into the appropriate language using a built-in translation engine. For example, a Japanese sentence like "Engine has a malfunction, please check!" is translated into English as "Engine has a malfunction, please check!"
[0881] Emotion analysis
[0882] The translated text data and the original audio data are analyzed by a server, which uses natural language processing and speech analysis techniques to assess the speaker's emotional state, for example, determining whether they are stressed or calm in an emergency.
[0883] Providing operational control and advice
[0884] The server controls autopilot or manual operation based on the emotion analysis results and text content. If it determines that the risk is high, it sends a signal to the aircraft's autopilot system to perform appropriate operations. It also generates advice and suggestions based on the situation in real time and provides them to the user via the terminal. For example, if an engine abnormality is detected, it immediately provides a voice message advising, "Please perform the engine restart procedure."
[0885] Specific examples
[0886] Emergency Scenarios
[0887] 1. The user (pilot) says, "There's an engine malfunction, please check!"
[0888] 2. The device captures this audio and sends it to the server.
[0889] 3. The server performs speech recognition and converts it into text.
[0890] 4. If necessary, the server translates the text into English.
[0891] 5. The server performs emotion analysis and detects that the pilot is in a high stress state.
[0892] 6. The server signals the autopilot system to execute a safety maneuver.
[0893] 7. The server generates an advice saying, "Autopilot has been strengthened due to an engine abnormality. Please check the engine instruments and perform the engine restart procedure." and notifies the user via the terminal.
[0894] 8. All processed data is stored on the server and used for later analysis.
[0895] This system can significantly improve aircraft safety and minimize communication errors caused by language differences or mental state during an emergency.
[0896] The processing flow will be explained below.
[0897] Step 1:
[0898] The device captures the voice data of the pilot and air traffic controller through microphones on the plane or the audio system in the air traffic control room, and transmits the captured voice data to a server in real time.
[0899] Step 2:
[0900] The server receives the voice data. It converts the voice data into text data using a speech recognition engine. For example, a voice message such as "There is an engine error. Please check!" is converted into text format.
[0901] Step 3:
[0902] The server identifies the language of the converted text data, using a built-in language identification engine to determine in which language the text is written.
[0903] Step 4:
[0904] The server translates the text data into the specified language as needed. For example, the Japanese text "Engine has a malfunction, please check!" is translated into English as "Engine has a malfunction, please check!".
[0905] Step 5:
[0906] The server performs sentiment analysis based on the translated text data and the original audio data, using natural language processing and audio analysis techniques to assess the speaker's emotional state (e.g., stressed, overwhelmed, calm, etc.).
[0907] Step 6:
[0908] The server assesses the risk level based on the results of emotion analysis and the text content. If the risk level is deemed high, the server sends a signal to the autopilot system to temporarily take control of the aircraft.
[0909] Step 7:
[0910] The server generates optimal countermeasures in real time based on the analysis results and the situation. This includes specific procedures and advice. For example, advice such as "Perform the engine restart procedure" is generated.
[0911] Step 8:
[0912] The terminal then notifies the pilot or controller of the generated advice in real time as a voice or text message.
[0913] Step 9:
[0914] The server stores all voice data, text data, translation data, sentiment analysis results, and response actions as logs, which will be used later for further analysis and system improvement.
[0915] Example 1
[0916] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0917] In communication between aircraft pilots and air traffic controllers, errors in information transmission can occur due to language differences or the pilot's mental state during an emergency. Such errors can have a significant impact on aircraft operation and pose a risk of compromising safety. It is also necessary to accurately grasp the pilot's stress level and emotions and provide appropriate responses in real time, but current systems are unable to adequately achieve this.
[0918] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0919] In this invention, the server includes means for acquiring voice data, means for converting the acquired voice data into text data, means for translating the text data into multiple languages, means for performing emotion analysis of the text data and voice data, means for controlling automatic or manual operation based on the emotion analysis results, means for generating advice and suggestions in real time based on the analysis results and the situation, means for storing all data for later analysis, means for transmitting the acquired voice data in digital form, means for converting the voice data into text data using a voice recognition engine, means for passing the voice data and text data to an emotion analysis engine, means for evaluating the emotional state using the emotion analysis engine, means for sending signals to an automatic operation system, and means for generating and transmitting voice messages based on the analysis results. This makes it possible to minimize communication errors due to language differences or mental states during an emergency, thereby significantly improving aircraft safety.
[0920] "Voice data" refers to data in which voice information uttered by a user is recorded in digital format.
[0921] "Text data" is character string information converted from voice data using a voice recognition engine.
[0922] A "translation engine" is a system that converts text data from an identified language to another language.
[0923] An "emotion analysis engine" is a system that analyzes text and audio data to evaluate the speaker's emotional state.
[0924] An "automatic operation system" is a system that performs automatic piloting and other automatic control of an aircraft based on the results of analysis.
[0925] "Digital format" refers to a format in which analog data is converted into numerical data.
[0926] A "voice recognition engine" is software or hardware for converting voice data into text data.
[0927] A "signal" is an electronic signal for transmitting control information.
[0928] "Real-time" refers to data being processed and analyzed as soon as it is acquired.
[0929] "Analysis results" are the analysis results derived by the sentiment analysis engine and other analysis means.
[0930] "Advice and suggestions" are measures and instructions generated based on the analysis results and the situation.
[0931] "Storage" refers to keeping acquired data and analysis results for a certain period of time.
[0932] "Later analysis" refers to re-analysis at a later date using the stored data.
[0933] This invention is a system for analyzing conversations between aircraft pilots and air traffic controllers in real time to control aircraft operations and ensure safety. The system captures voice data, converts it into text data, translates it into multiple languages, performs sentiment analysis, controls automatic or manual operations based on the analysis results, and generates advice and suggestions in real time as needed. Furthermore, all data is stored for later analysis.
[0934] System configuration and operation
[0935] Acquiring audio data
[0936] The user (pilot or air traffic controller) reports the emergency situation by voice, and this voice data is picked up by the terminal in digital format through the microphone on board or the voice system in the air traffic control room, and then transmitted to the server.
[0937] Voice Recognition
[0938] The server passes the received voice data to a voice recognition engine. This engine generally uses high-performance voice recognition technology such as Google Speech-to-Text API or IBM Watson Speech to Text. The voice data is converted into text data. For example, a voice saying "There is an engine problem, please check!" is converted into text "There is an engine problem, please check!"
[0939] Text data translation
[0940] The server then automatically identifies the language of the text data. If the identified language is other than English, the server uses a built-in translation engine (e.g., Microsoft Translator or Google Translate API) to translate the text data into English. For example, the Japanese input "Engine has a malfunction, please check!" is translated into "Engine has a malfunction, please check!"
[0941] Emotion analysis
[0942] The server passes the translated text data and the original audio data to a sentiment analysis engine, which uses natural language processing and speech analysis technologies such as IBM Watson Natural Language Understanding and Microsoft Azure Text Analytics to assess the speaker's emotional state, for example, determining whether they are highly stressed or calm in an emergency.
[0943] Providing operational control and advice
[0944] The server sends a signal to the aircraft's automatic operation system based on the results of the emotion analysis and the content of the text data. If a high level of danger is detected, the server immediately executes safety operations. It also generates advice and suggestions based on the situation in real time and provides them to the user via the terminal. For example, advice such as "Perform engine restart procedures" is immediately provided as a voice message.
[0945] Data storage
[0946] All processed data is stored on a server as a detailed log, which is used for future analysis, system improvement, aircraft accident investigations, and operational optimization.
[0947] Specific examples
[0948] Emergency Scenarios
[0949] 1. The user (pilot) reports audibly, "An engine malfunction has occurred. Please check!"
[0950] 2. The device captures this audio and immediately transmits it digitally to the server.
[0951] 3. The server passes the received voice data to a voice recognition engine and converts it into text data.
[0952] 4. The server identifies the language of the text data and translates it into English using a translation engine if necessary.
[0953] 5. The server analyzes the translated text data and the original audio data using an emotion analysis engine to detect high stress levels.
[0954] 6. The server sends a signal to the automated operation system to immediately execute safety operations, and also creates a situation-specific advice message, such as "Perform engine restart procedure," as a voice message and provides it to the user via the terminal.
[0955] 7. All processing data will be stored on the server and used for future analysis and improvement.
[0956] Prompt Sentence Examples
[0957] Prompt: "Analyze the audio of an aircraft pilot reporting an emergency and provide a detailed description of how the automated flight control system responds and provides advice to the pilot."
[0958] The use of this system can significantly improve aviation safety and minimize communication errors caused by language differences or mental state during an emergency.
[0959] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0960] The flow of this system's program processing
[0961] Step 1:
[0962] The user (pilot or air traffic controller) reports information verbally. For example, in an emergency situation, they might say, "There's an engine malfunction, please check!" This is the input.
[0963] Step 2:
[0964] The terminal captures the user's voice in digital form through an onboard microphone or an audio system in the air traffic control room, and transmits this digital voice data to a server.
[0965] Input: User's voice
[0966] Output: Digital audio data
[0967] Step 3:
[0968] The server passes the received digital voice data to a voice recognition engine (e.g., Google Speech-to-Text API or IBM Watson Speech to Text) and converts it into text data.
[0969] Input: Digital audio data
[0970] Output: Text data: "An engine error has occurred. Please check!"
[0971] Step 4:
[0972] Identifying the language of server-generated text data, for example automatically recognizing that the text is in Japanese.
[0973] Input: Text data
[0974] Output: Language identification result (Japanese)
[0975] Step 5:
[0976] The server uses a built-in translation engine (e.g., Microsoft Translator or Google Translate API) to translate from the identified language to English. For example, "Engine has a malfunction, please check!" is translated to "Engine has a malfunction, please check!"
[0977] Input: Japanese text data
[0978] Output: English text data
[0979] Step 6:
[0980] The server passes the translated English text data and the original audio data to a sentiment analysis engine (e.g., IBM Watson Natural Language Understanding or Microsoft Azure Text Analytics) to evaluate the speaker's emotional state, e.g., to determine whether the speaker has a high stress level.
[0981] Input: English text and audio data
[0982] Output: Emotion analysis results (high stress)
[0983] Step 7:
[0984] Based on the results of the emotion analysis and the content of the text data, the server sends a signal to the aircraft's automatic operation system to perform appropriate safety operations, such as sending a signal to "strengthen autopilot."
[0985] Input: Sentiment analysis results and text data
[0986] Output: Control signal to automatic operation system
[0987] Step 8:
[0988] The server generates advice and suggestions based on the situation in real time and provides them to the user through the terminal, for example, generating a voice message saying "Please perform the engine restart procedure."
[0989] Input: Sentiment analysis results and situation information
[0990] Output: Advice and suggestions (voice message)
[0991] Step 9:
[0992] The server stores all transaction data in a detailed log, which is used for future analysis and system improvement.
[0993] Input: All analysis results and process logs
[0994] Output: Saved processed data
[0995] (Application example 1)
[0996] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0997] In on-site security work, security personnel often experience high levels of stress during emergencies, which can result in delayed responses and mistakes. Furthermore, communication barriers in multilingual environments and difficulty assessing situations in real time can also be a factor, potentially worsening dangerous situations.
[0998] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0999] In this invention, the server includes means for acquiring voice data, means for converting the acquired voice data into text data, means for translating the text data into multiple languages, means for performing emotion analysis of the text data and voice data, means for controlling automatic or manual operation based on the emotion analysis results, means for generating advice and suggestions in real time based on the analysis results and the situation, means for alerting the user when a high-stress state is detected, and means for saving all data for later analysis, thereby enabling quick and appropriate decisions to be made even in emergencies or high-stress situations.
[1000] The "means for acquiring voice data" refers to a device or software that captures the voice uttered by the user in real time and transmits the voice data to the system.
[1001] The "means for converting acquired voice data into text data" refers to a device or software that uses voice recognition technology to convert acquired voice data into a corresponding string of characters.
[1002] A "means for translating text data into multiple languages" is a translation engine or algorithm for converting text data into another language.
[1003] "Means for analyzing emotions in text data and voice data" refers to devices or software that use natural language processing and voice analysis techniques to evaluate the emotions of a speaker from text data or voice data.
[1004] "Means for controlling automatic or manual operations based on the results of emotion analysis" refers to devices or software that appropriately manage or correct automatic operations or manual instructions within the system based on the results of emotion analysis.
[1005] "Means for generating advice and suggestions in real time based on analysis results and the current situation" refers to devices or software that generate appropriate advice and suggestions in a timely manner based on data analysis results and the current situation.
[1006] "Means for alerting the user when a high stress state is detected" refers to devices or software that alert the user by issuing appropriate alerts or notifications when it is determined that the user's stress level is high.
[1007] "Means for storing all data and making it available for later analysis" refers to devices or software that store all data processed by the system and make it available for later analysis, performance evaluation, etc.
[1008] The present invention is a real-time assistant system for supporting security operations. This system has the functions of capturing voice data, converting it into text data, translating it into multiple languages, performing emotion analysis, providing advice in real time, and detecting high-stress situations.
[1009] System configuration
[1010] The system consists of the following main components:
[1011] 1. A means of acquiring voice data: Capture the user's voice in real time, specifically using a smartphone or microphone.
[1012] 2. Means for converting acquired voice data into text data: Convert the voice into text using voice recognition technology (for example, the SpeechRecognition library).
[1013] 3. Means for translating text data into multiple languages: Translate text data into other languages using a translation engine (e.g., Google Translate API).
[1014] 4. Means for sentiment analysis of text and voice data: Sentiment analysis is performed using natural language processing technology and voice analysis technology (e.g., Hugging Face Transformers).
[1015] 5. Means for controlling automatic or manual operations based on the results of sentiment analysis: Based on the results of sentiment analysis, operations within the system are appropriately managed or corrected.
[1016] 6. Means of generating advice and suggestions in real time based on analysis results and the current situation: Generate appropriate advice and suggestions based on the results of data analysis and the current situation.
[1017] 7. A method to alert users when high stress levels are detected: When stress levels are determined to be high, appropriate alerts and notifications will be provided.
[1018] 8. Means of storing all data for later analysis: All processed data will be stored and used for later analysis and performance evaluation.
[1019] System Operation
[1020] The specific operation of the system can be explained as follows:
[1021] Acquiring voice data and converting it to text data
[1022] The voice spoken by the user (security officer) is picked up through the smartphone's microphone. This voice data is converted into text using voice recognition technology. For example, a voice saying "Someone is breaking in! Call for help!" is converted into text data.
[1023] Text data translation
[1024] This text data is translated from Japanese to English using a language recognition function. Using a translation engine, "Someone is trespassing, please call for help immediately!" is converted to "Someone is trespassing, please call for help immediately!"
[1025] Emotion analysis
[1026] The translated text data and the original audio data are analyzed by a sentiment analysis model to determine the user's emergency state or high stress level. If a specific emotional state is detected, appropriate advice is generated.
[1027] Real-time advice and suggestions
[1028] Based on the results of sentiment analysis and the text content, the system generates real-time advice. For example, if a high stress state is detected, the system will provide advice such as "Be careful. Stress is increasing."
[1029] Data storage
[1030] All data is logged and made available for later analysis or performance evaluation.
[1031] Specific examples
[1032] For example, if a security worker says, "Someone is trespassing, please call for help immediately!", the system converts the speech into text and translates it into "Someone is trespassing, please call for help immediately!". If the emotion analysis determines that this is an "emergency situation," the system provides a real-time voice message with advice such as, "Be careful. Stress is building."
[1033] Prompt Sentence Examples
[1034] Voice: "Someone's breaking in, hurry up and call for help!"
[1035] Translation: "Someone is trespassing, please call for help immediately!"
[1036] Sentiment Analysis: "Emergency"
[1037] Advice: "Pay attention. Stress is building."
[1038] This allows security personnel to respond quickly and appropriately.
[1039] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1040] Step 1:
[1041] The device captures the user's voice in real time through a microphone. The input is the user's speech, and the output is digital voice data. This voice data is sent directly to the server.
[1042] Step 2:
[1043] The server converts the acquired voice data into text data using a voice recognition engine (for example, the SpeechRecognition library). The input is digital voice data, and the output is the corresponding text data. This conversion makes it easier to handle voice information as text information.
[1044] Step 3:
[1045] The server sends the generated text data to a translation engine (for example, Google Translate API) to translate it into another language. The input is text data in the original language, and the output is the translated text data. This translation allows for smooth support in multilingual environments.
[1046] Step 4:
[1047] The server performs sentiment analysis using the translated text data and the original voice data. Natural language processing techniques (e.g., Hugging Face Transformers) are used for sentiment analysis. The input is text and voice data, and the output is the analysis result of the emotional state. Based on the analysis result, the user's state of urgency and stress level are evaluated.
[1048] Step 5:
[1049] The server controls automatic or manual operation based on the results of emotion analysis. Specifically, if a high-stress state is detected, it automatically changes the system operation or issues instructions to the user. The input to this step is the emotion analysis result, and the output is appropriate operation control instructions.
[1050] Step 6:
[1051] The server generates advice and suggestions in real time based on the analysis results and the current situation. For example, if a high stress state is detected, advice such as "Be careful. Your stress is increasing" is generated. The input is the analysis results and situation information, and the output is the generated advice.
[1052] Step 7:
[1053] The server sends the generated advice to the terminal and notifies the user. The input is the generated advice, and the output is the notification received by the user.
[1054] Step 8:
[1055] All data is stored by the server. The stored data is used for later analysis and performance evaluation. The input is all the data during and after processing, and the output is the stored data. This storage allows for future optimization and improvement.
[1056] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1057] This invention is a system for analyzing conversations between aircraft pilots and air traffic controllers in real time, controlling aircraft operations, and ensuring safety. The system incorporates an emotion engine for recognizing the user's emotional state, thereby further improving safety. The following is a detailed description of an embodiment of the system.
[1058] System configuration
[1059] The system consists of the following main components:
[1060] 1. Device that acquires audio data
[1061] 2. A server that converts the acquired voice data into text data
[1062] 3. Server that translates text data into multiple languages
[1063] 4. Server that performs emotion analysis of text data and voice data
[1064] 5. Emotion Engine
[1065] 6. Server that controls automatic or manual operation based on the results of emotion analysis
[1066] 7. Server that generates advice and suggestions in real time based on the analysis results and the situation
[1067] 8. A server that stores all data and uses it for later analysis
[1068] System Operation
[1069] Acquiring audio data
[1070] Voice information from the user (pilot or air traffic controller) is acquired through a microphone on board the aircraft or through the voice system in the air traffic control room. The terminal captures this voice data and transmits it.
[1071] Voice Recognition
[1072] When the voice data arrives at the server, the voice recognition engine activates and converts it into text data. For example, a voice message such as "There is an engine error. Please check!" is converted into text format.
[1073] Text data translation
[1074] The server then identifies the language of this text data and, if necessary, translates it into the appropriate language using a built-in translation engine. For example, a Japanese sentence like "Engine has a malfunction, please check!" is translated into English as "Engine has a malfunction, please check!"
[1075] Emotion analysis
[1076] The translated text data and the original audio data are analyzed by a server, which uses natural language processing and speech analysis techniques to assess the speaker's emotional state, for example, determining whether they are highly stressed or calm in an emergency.
[1077] Use of emotion engine
[1078] The server is equipped with an emotion engine that can recognize the user's emotional state in more detail. This emotion engine analyzes both text and voice data and has the ability to highly evaluate the user's emotional state. This allows the system to determine the level of urgency corresponding to a specific emotional state and propose appropriate countermeasures accordingly.
[1079] Providing operational control and advice
[1080] The server controls autopilot or manual operation based on the emotion analysis results and text content. If it determines that the risk is high, it sends a signal to the aircraft's autopilot system to perform appropriate operations. It also generates advice and suggestions based on the situation in real time and provides them to the user via the terminal. For example, if an engine abnormality is detected, it immediately provides a voice message advising, "Please perform the engine restart procedure."
[1081] Data storage
[1082] The server stores all processing data, analysis results, audio data, and response measures as logs, which will be used later for detailed analysis and system improvement.
[1083] Specific examples
[1084] Emergency Scenarios
[1085] 1. The user (pilot) says, "There's an engine malfunction, please check!"
[1086] 2. The device captures this audio and sends it to the server.
[1087] 3. The server performs speech recognition and converts it into text.
[1088] 4. If necessary, the server translates the text into English.
[1089] 5. The server performs emotion analysis and detects that the pilot is in a high stress state.
[1090] 6. The emotion engine analyzes the user's stress level in detail and determines that the level of urgency is high.
[1091] 7. The server signals the autopilot system to execute a safety maneuver.
[1092] 8. The server generates an advice saying, "Autopilot has been strengthened due to an engine abnormality. Please check the engine instruments and perform the engine restart procedure." and notifies the user via the terminal.
[1093] 9. All processed data is stored on the server and used for later analysis.
[1094] The system can significantly improve aviation safety and minimize communication errors by providing countermeasures based on the user's emotional state.
[1095] The processing flow will be explained below.
[1096] Step 1:
[1097] The device captures the voice data of the pilot and air traffic controller through microphones on the plane or the audio system in the air traffic control room, and transmits the captured voice data to a server in real time.
[1098] Step 2:
[1099] The server receives the voice data. It converts the voice data into text data using a speech recognition engine. For example, a voice message such as "There is an engine error. Please check!" is converted into text format.
[1100] Step 3:
[1101] The server identifies the language of the converted text data, using a built-in language identification engine to determine in which language the text is written.
[1102] Step 4:
[1103] The server translates the text data into the specified language as needed. For example, the Japanese text "Engine has a malfunction, please check!" is translated into English as "Engine has a malfunction, please check!".
[1104] Step 5:
[1105] The server performs sentiment analysis based on the translated text data and the original audio data, using natural language processing and audio analysis technologies to assess the speaker's emotional state (e.g., stressed, overwhelmed, calm, etc.).
[1106] Step 6:
[1107] The emotion engine performs detailed analysis of text and voice data to intelligently assess the user's emotional state, for example detecting when a pilot is under high stress.
[1108] Step 7:
[1109] The server assesses the level of danger based on the results of sentiment analysis and the text content. If the level of danger is deemed high, the server sends a signal to the autopilot system to temporarily control the aircraft's operation. For example, if a pilot is under high stress, the autopilot system will automatically maintain the aircraft's altitude and speed.
[1110] Step 8:
[1111] The server generates optimal countermeasures in real time based on the analysis results and the situation. This includes specific procedures and advice. For example, advice such as "Perform the engine restart procedure" is generated.
[1112] Step 9:
[1113] The terminal then notifies the pilot or controller of the generated advice in real time as a voice or text message.
[1114] Step 10:
[1115] The server stores all voice data, text data, translation data, sentiment analysis results, and response actions as logs, which will be used later for further analysis and system improvement.
[1116] Example 2
[1117] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1118] In communications between aircraft pilots and air traffic controllers, differences in emotional state and language during emergencies can make it difficult to understand instructions, delaying appropriate responses. To solve this problem, real-time voice analysis, evaluation of emotional state, and proposal of appropriate countermeasures are required, but previous systems have found it difficult to perform these tasks in an integrated manner.
[1119] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for acquiring voice data, means for converting the acquired voice data into text data, means for translating the text data into multiple languages, means for performing emotion analysis of the text data and voice data, means for controlling automatic operation or manual operation based on the emotion analysis results, means for generating advice and suggestions in real time based on the analysis results and the situation, means for saving all data and using it for later analysis, and means for incorporating an emotion engine that highly evaluates the user's emotional state. This makes it possible to provide prompt measures based on accurate voice recognition and emotion analysis in real time.
[1120] "Voice data" refers to voices uttered by a user captured as digital signals.
[1121] "Means for acquiring" refers to the technology or device for acquiring voice data using a terminal or sensor within the domain.
[1122] "Text data" is voice data that has been analyzed and expressed as text information.
[1123] "Means for converting into text data" refers to technology or devices that convert acquired voice data into text format using natural language processing technology.
[1124] "Translation means" refers to the technology and software used to convert text data into other languages.
[1125] "Sentiment analysis" is the process of analyzing audio and text data to assess the emotional state of a speaker.
[1126] "Means for performing emotion analysis" refers to technologies or devices that use the algorithms or software required for emotion analysis.
[1127] An "emotion engine" is specialized software or algorithms that intelligently assess a user's emotional state.
[1128] "Means for controlling autopilot or manual operation" refers to technology or devices that automatically control aircraft operation based on the results of emotion analysis or in accordance with the pilot's instructions.
[1129] The "means for generating advice or suggestions" refers to technology or software that generates suggestions in real time to help the user take appropriate actions or make appropriate decisions based on the analysis results and the situation.
[1130] "Means for storing data and using it for later analysis" refers to the technology and devices that record all processing data and use it for later data analysis and system improvement.
[1131] This invention is a system for analyzing conversations between aircraft pilots and air traffic controllers in real time to control aircraft operations and ensure safety. The system uses multiple servers, terminals, and software to achieve functions such as voice data acquisition, voice recognition, text data translation, emotion analysis, autopilot or manual operation control, advice and suggestion generation, and data storage.
[1132] System configuration
[1133] The system consists of the following main components:
[1134] 1. Device that acquires audio data
[1135] 2. Server that converts voice data into text data
[1136] 3. Server that translates text data into multiple languages
[1137] 4. Server that performs emotion analysis of text data and voice data
[1138] 5. Emotion Engine
[1139] 6. Server that controls automatic or manual operation based on the results of emotion analysis
[1140] 7. Server that generates advice and suggestions in real time based on the analysis results and the situation
[1141] 8. A server that stores all data and uses it for later analysis
[1142] Hardware and software used
[1143] 1. Terminal: The aircraft's microphone and communication system are used to acquire voice data. Standard data communication protocols are used to transmit the voice data.
[1144] 2. Server: To analyze the voice data, a speech recognition engine such as Google Cloud Speech-to-Text or Amazon Transcribe is used.
[1145] 3. Text data translation: Translate text data into multiple languages using the Google Translate API, etc.
[1146] 4. Sentiment Analysis: Evaluate emotional state using IBM Watson Natural Language Understanding and Amazon's speech analysis technology.
[1147] 5. Emotional Engine: Utilizing specialized engines such as Affectiva to provide a sophisticated assessment of the user's emotional state.
[1148] 6. Autopilot and manual operation control: Works in conjunction with the aircraft's autopilot system (e.g., Honeywell system) to ensure safe operation.
[1149] 7. Advice and suggestion generation: Use dedicated advice generation algorithms to generate context-sensitive advice in real time.
[1150] 8. Data storage: All processed data, analysis results, audio data, and response actions will be stored using cloud storage for later analysis.
[1151] Specific examples
[1152] Emergency scenarios:
[1153] 1. The user (pilot) says, "There's an engine malfunction, please check!"
[1154] 2. The onboard microphone captures the audio and sends the data to the device.
[1155] 3. The device sends the audio data to the server.
[1156] 4. The server receives the audio data and calls the Google Cloud Speech-to-Text service.
[1157] 5. The server converts the voice data into text and generates the text "An engine error has occurred, please check!"
[1158] 6. The server translates the text data into English using the Google Translate API. "Engine has a malfunction, please check!" becomes "Engine has a malfunction, please check!"
[1159] 7. The server sends the text data to IBM Watson Natural Language Understanding for sentiment analysis.
[1160] 8. The server sends the audio data to Amazon Transcribe, which analyzes the emotional state of the audio.
[1161] 9. The emotion engine on the server integrates the text and voice data to perform a detailed emotion assessment.
[1162] 10. The emotion engine determines that the user's stress level is "high."
[1163] 11. The server sends a signal to the autopilot system, generating the message "Perform engine restart procedure."
[1164] 12. The device will announce this message to the user audibly.
[1165] 13. The server records the processed data, analysis results, audio data, and response measures as logs and stores them in a database.
[1166] 14. The stored data can be later analyzed to help improve the system.
[1167] Prompt Sentence Examples
[1168] "Please explain the steps your system takes when a user says, 'There's an engine malfunction, please check!'"
[1169] In this way, the system significantly improves aircraft safety, can provide countermeasures based on the user's emotional state, and minimizes communication errors.
[1170] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1171] Step 1:
[1172] The voice of the user (pilot or air traffic controller) is acquired through an on-board microphone or the voice system in the air traffic control room. This captures the user's real-time voice information as digital voice data. The input is the user's voice, and the output is digital voice data.
[1173] Step 2:
[1174] The device sends the captured audio data to a server. The audio data is transferred to the server via the Internet or a dedicated communication line. The input is the captured digital audio data, and the output is the transferred audio data.
[1175] Step 3:
[1176] The server receives the voice data and calls the Google Cloud Speech-to-Text service to perform speech recognition. The voice data is analyzed and converted into text data. For example, a speech saying "There is an engine problem, please check!" is converted into text "There is an engine problem, please check!" The input is voice data, and the output is text data.
[1177] Step 4:
[1178] The server identifies the language of the generated text data. If necessary, it uses the Google Translate API to translate the text into other languages. For example, text entered in Japanese is translated into English. The input is the text data, and the output is the translated text data.
[1179] Step 5:
[1180] The server sends the text data to IBM Watson Natural Language Understanding for sentiment analysis. The text data is used to evaluate the speaker's emotional state. For example, the text content can be used to analyze a pilot's stress level and urgency. The input is text data, and the output is an evaluation of the emotional state.
[1181] Step 6:
[1182] The server sends the audio data to Amazon Transcribe, which analyzes the tone and pitch of the voice. This allows for a detailed assessment of the speaker's emotional state and urgency from the audio data. The input is the audio data, and the output is the emotional assessment result from the audio analysis.
[1183] Step 7:
[1184] The emotion engine in the server integrates the results of the analysis of the text data and voice data to perform a detailed emotion evaluation. The emotion engine highly evaluates the user's emotional state and determines, for example, that the user is in a high-stress state. The inputs are the emotion analysis results and the voice analysis results, and the output is the integrated emotion evaluation result.
[1185] Step 8:
[1186] The server generates a control signal to an autopilot system or manual operation based on the emotion analysis result, for example, to signal an aircraft's autopilot system to execute emergency measures. The input is the integrated emotion evaluation result, and the output is a control signal.
[1187] Step 9:
[1188] The server generates appropriate advice or suggestions in real time based on the analysis results and the situation, and notifies the user through the terminal. For example, it generates advice such as "Please perform the engine restart procedure." The input is the emotion analysis results and situation information, and the output is the generated advice.
[1189] Step 10:
[1190] The server records and stores all processing data, analysis results, audio data, and response measures as logs. The stored data is used for later detailed analysis and system improvement. The input is the processing data and analysis results, and the output is the stored log data.
[1191] (Application example 2)
[1192] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1193] Conventional call center systems have the problem of being unable to grasp the emotional state of the operator and customer in real time and provide immediate solutions. This has led to problems such as customers' dissatisfaction not being resolved properly and operators feeling high levels of stress. This leads to problems such as lower customer satisfaction and a decline in the work efficiency of operators.
[1194] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1195] In this invention, the server includes means for acquiring voice data, means for converting the acquired voice data into text data, means for translating the text data into multiple languages, means for performing emotion analysis of the text data and voice data, means for controlling automatic or manual operations based on the emotion analysis results, means for generating advice and suggestions in real time based on the analysis results and the situation, means for triage and proposing countermeasures based on the emotional state, and means for saving all data and using it for later analysis. This makes it possible to grasp the emotional states of customers and operators in real time and immediately provide appropriate countermeasures.
[1196] "Audio data" refers to sound information in digital or analog form collected through audio equipment such as a microphone.
[1197] "Text data" is information obtained by converting voice data into a character string format.
[1198] A "multilingual translation method" is a technique or process for converting text data into different languages.
[1199] "Emotion analysis" is a technique that analyzes audio and text data to assess a speaker's emotional state.
[1200] "Autonomous operation" means that the system performs a specific action without human intervention.
[1201] "Manual operation" means direct human involvement in operating a system.
[1202] "Means for generating advice and suggestions in real time" refers to techniques and methods that provide immediate advice to operators or users based on the current situation.
[1203] "Triage" is the prioritization of response measures based on urgency and importance.
[1204] A "solution" is an action or suggestion to be taken in response to a particular situation.
[1205] "Storing data" means recording collected data on a storage medium so that it can be used later.
[1206] "Use for later analysis" means to conduct analysis based on the stored data for future consideration and improvement.
[1207] MODE FOR CARRYING OUT THE INVENTION
[1208] The present invention is a system for supporting appropriate responses by analyzing the emotional states of customers and agents in a call center in real time. An embodiment of this system will be described in detail below.
[1209] System configuration
[1210] The system consists of the following main components:
[1211] 1. Means of capturing voice data: microphones and telephone systems to capture conversations between customers and operators.
[1212] 2. A means of converting the acquired voice data into text data: a speech recognition engine that converts the voice into a string format (e.g., Google Speech-to-Text API).
[1213] 3. A means of translating text data into multiple languages: A system that translates text into the required languages (e.g., Microsoft Azure Cognitive Services).
[1214] 4. Means of sentiment analysis of text and voice data: Software that analyzes the emotional state of customers and operators (e.g., IBM Watson Tone Analyzer).
[1215] 5. Means for controlling automatic or manual operations based on the results of emotion analysis: A control system for performing processing according to emotional states.
[1216] 6. Means of generating advice and suggestions in real time based on analysis results and situations: An engine that generates real time advice and suggestions to operators.
[1217] 7. Triage and suggest responses based on emotional state: A mechanism for prioritizing responses based on urgency and importance.
[1218] 8. A means of storing all data and making it available for later analysis: a database to store the collected and analyzed data.
[1219] System Operation
[1220] Audio Acquisition
[1221] Voice information from users (customers and operators) is acquired through a microphone or telephone system. The terminal captures this voice data and sends it to the server.
[1222] Voice Recognition
[1223] The voice data received by the server is converted into text data by a voice recognition engine (Google Speech-to-Text API).
[1224] Text data translation
[1225] The server then converts this text data into the appropriate language if it needs to be translated using a multi-language translation means.
[1226] Emotion analysis
[1227] The translated text data and the original audio data are then subjected to natural language processing and speech analysis techniques to assess the user's emotional state, using analysis engines such as IBM Watson Tone Analyzer.
[1228] Detailed analysis by Emotion Engine
[1229] The emotion engine performs advanced analysis of text and voice data to determine specific emotional states and their urgency.
[1230] Providing operational control and advice
[1231] Based on the results of the emotion analysis, the server controls automatic or manual operations and generates and notifies advice and suggestions in real time. Appropriate advice is provided to the operator in voice or text format.
[1232] Data storage
[1233] The server stores all audio data, text data, and analysis results in a database for later analysis.
[1234] Specific examples
[1235] Emergency Scenarios
[1236] 1. A customer says, "I recently purchased a broken item and I'm very unhappy with it. I'd like to process the return immediately."
[1237] 2. The device captures this audio and sends it to the server.
[1238] 3. The server performs speech recognition and converts it into text.
[1239] 4. If necessary, the server translates the text.
[1240] 5. The server performs sentiment analysis and detects that the customer is in a high stress state.
[1241] 6. The emotion engine analyzes the customer's stress level in detail and determines that the situation is urgent.
[1242] 7. The server advises the operator, "The customer is unhappy. Please stay calm and explain with specific examples."
[1243] 8. All processed data is stored on the server and used for later analysis.
[1244] This system will significantly improve the quality and efficiency of customer service at call centers, increasing customer satisfaction. It will also reduce stress on operators and contribute to improved work efficiency.
[1245] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1246] Step 1:
[1247] The user (customer and operator) starts a conversation. The terminal captures the voice data through a microphone or telephone system. This captured voice data is recorded in digital format and sent to the server. The input here is the raw voice data, and the output is the audio file sent to the server.
[1248] Step 2:
[1249] The server sends the received voice data to the Google Speech-to-Text API for speech recognition. In this process, the voice data is converted into text data in string format. The input is the voice data, and the output is the converted text data. Specifically, the voice data undergoes signal processing and pattern matching.
[1250] Step 3:
[1251] The server uses Microsoft Azure Cognitive Services to translate the acquired text data into multiple languages as needed. In this step, the language of the text data is identified and converted into the appropriate language. The input is text data, and the output is translated text data. The specific operation utilizes natural language processing algorithms.
[1252] Step 4:
[1253] The server then sends the translated text data and the original audio data to the IBM Watson Tone Analyzer for sentiment analysis. The input is text data and audio data, and the output is the sentiment analysis results. This function evaluates the speaker's emotional state based on the tone and content of the audio.
[1254] Step 5:
[1255] The server further evaluates the emotion analysis results to determine the specific emotional state and its urgency. This is done by the emotion engine, which understands the user's emotional state. The input is the emotion analysis results, and the output is detailed emotional state and urgency information. Specific operations include integrating and evaluating the analysis results.
[1256] Step 6:
[1257] The server generates and notifies the operator in real time with advice and suggestions based on the results of emotion analysis. The input is the detailed emotional state and the customer's comments, and the output is advice and suggestions provided to the operator. Specifically, the system generates appropriate countermeasures and notifies the operator via voice or text.
[1258] Step 7:
[1259] The server stores all data in a database for later analysis. The input is all processed data (audio, text, analysis results, etc.), and the output is the stored data. Specific operations include recording and storing data.
[1260] Prompt Sentence Examples
[1261] Speech input: "I'm very unhappy with the product I recently purchased, which is broken. I'd like to process the return immediately."
[1262] Expected output: "The customer is unhappy. Please stay calm and provide specific examples. We'll expedite the return process."
[1263] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1264] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1265] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1266] [Fourth embodiment]
[1267] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1268] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1269] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1270] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1271] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1272] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1273] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1274] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1275] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1276] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1277] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1278] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1279] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1280] The present invention provides a system for analyzing conversations between aircraft pilots and air traffic controllers in real time, controlling aircraft operations, and ensuring safety. The following is a detailed description of an embodiment of the system.
[1281] System configuration
[1282] The system consists of the following main components:
[1283] 1. Device that acquires audio data
[1284] 2. A server that converts the acquired voice data into text data
[1285] 3. Server that translates text data into multiple languages
[1286] 4. Server that performs emotion analysis of text data and voice data
[1287] 5. Server that controls automatic or manual operation based on the results of emotion analysis
[1288] 6. Server that generates advice and suggestions in real time based on the analysis results and the situation
[1289] 7. A server that stores all data and uses it for later analysis
[1290] System Operation
[1291] Acquiring audio data
[1292] Voice information from the user (pilot or air traffic controller) is acquired through a microphone on board the aircraft or through the voice system in the air traffic control room. The terminal captures this voice data and transmits it.
[1293] Voice Recognition
[1294] When the voice data arrives at the server, the voice recognition engine activates and converts it into text data. For example, a voice saying "There is an engine problem, please check!" is converted into the corresponding text.
[1295] Text data translation
[1296] The server then identifies the language of this text data and, if necessary, translates it into the appropriate language using a built-in translation engine. For example, a Japanese sentence like "Engine has a malfunction, please check!" is translated into English as "Engine has a malfunction, please check!"
[1297] Emotion analysis
[1298] The translated text data and the original audio data are analyzed by a server, which uses natural language processing and speech analysis techniques to assess the speaker's emotional state, for example, determining whether they are stressed or calm in an emergency.
[1299] Providing operational control and advice
[1300] The server controls autopilot or manual operation based on the emotion analysis results and text content. If it determines that the risk is high, it sends a signal to the aircraft's autopilot system to perform appropriate operations. It also generates advice and suggestions based on the situation in real time and provides them to the user via the terminal. For example, if an engine abnormality is detected, it immediately provides a voice message advising, "Please perform the engine restart procedure."
[1301] Specific examples
[1302] Emergency Scenarios
[1303] 1. The user (pilot) says, "There's an engine malfunction, please check!"
[1304] 2. The device captures this audio and sends it to the server.
[1305] 3. The server performs speech recognition and converts it into text.
[1306] 4. If necessary, the server translates the text into English.
[1307] 5. The server performs emotion analysis and detects that the pilot is in a high stress state.
[1308] 6. The server signals the autopilot system to execute a safety maneuver.
[1309] 7. The server generates an advice saying, "Autopilot has been strengthened due to an engine abnormality. Please check the engine instruments and perform the engine restart procedure." and notifies the user via the terminal.
[1310] 8. All processed data is stored on the server and used for later analysis.
[1311] This system can significantly improve aircraft safety and minimize communication errors caused by language differences or mental state during an emergency.
[1312] The processing flow will be explained below.
[1313] Step 1:
[1314] The device captures the voice data of the pilot and air traffic controller through microphones on the plane or the audio system in the air traffic control room, and transmits the captured voice data to a server in real time.
[1315] Step 2:
[1316] The server receives the voice data. It converts the voice data into text data using a speech recognition engine. For example, a voice message such as "There is an engine error. Please check!" is converted into text format.
[1317] Step 3:
[1318] The server identifies the language of the converted text data, using a built-in language identification engine to determine in which language the text is written.
[1319] Step 4:
[1320] The server translates the text data into the specified language as needed. For example, the Japanese text "Engine has a malfunction, please check!" is translated into English as "Engine has a malfunction, please check!".
[1321] Step 5:
[1322] The server performs sentiment analysis based on the translated text data and the original audio data, using natural language processing and audio analysis techniques to assess the speaker's emotional state (e.g., stressed, overwhelmed, calm, etc.).
[1323] Step 6:
[1324] The server assesses the risk level based on the results of emotion analysis and the text content. If the risk level is deemed high, the server sends a signal to the autopilot system to temporarily take control of the aircraft.
[1325] Step 7:
[1326] The server generates optimal countermeasures in real time based on the analysis results and the situation. This includes specific procedures and advice. For example, advice such as "Perform the engine restart procedure" is generated.
[1327] Step 8:
[1328] The terminal then notifies the pilot or controller of the generated advice in real time as a voice or text message.
[1329] Step 9:
[1330] The server stores all voice data, text data, translation data, sentiment analysis results, and response actions as logs, which will be used later for further analysis and system improvement.
[1331] Example 1
[1332] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1333] In communication between aircraft pilots and air traffic controllers, errors in information transmission can occur due to language differences or the pilot's mental state during an emergency. Such errors can have a significant impact on aircraft operation and pose a risk of compromising safety. It is also necessary to accurately grasp the pilot's stress level and emotions and provide appropriate responses in real time, but current systems are unable to adequately achieve this.
[1334] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1335] In this invention, the server includes means for acquiring voice data, means for converting the acquired voice data into text data, means for translating the text data into multiple languages, means for performing emotion analysis of the text data and voice data, means for controlling automatic or manual operation based on the emotion analysis results, means for generating advice and suggestions in real time based on the analysis results and the situation, means for storing all data for later analysis, means for transmitting the acquired voice data in digital form, means for converting the voice data into text data using a voice recognition engine, means for passing the voice data and text data to an emotion analysis engine, means for evaluating the emotional state using the emotion analysis engine, means for sending signals to an automatic operation system, and means for generating and transmitting voice messages based on the analysis results. This makes it possible to minimize communication errors due to language differences or mental states during an emergency, thereby significantly improving aircraft safety.
[1336] "Voice data" refers to data in which voice information uttered by a user is recorded in digital format.
[1337] "Text data" is character string information converted from voice data using a voice recognition engine.
[1338] A "translation engine" is a system that converts text data from an identified language to another language.
[1339] An "emotion analysis engine" is a system that analyzes text and audio data to evaluate the speaker's emotional state.
[1340] An "automatic operation system" is a system that performs automatic piloting and other automatic control of an aircraft based on the results of analysis.
[1341] "Digital format" refers to a format in which analog data is converted into numerical data.
[1342] A "voice recognition engine" is software or hardware for converting voice data into text data.
[1343] A "signal" is an electronic signal for transmitting control information.
[1344] "Real-time" refers to data being processed and analyzed as soon as it is acquired.
[1345] "Analysis results" are the analysis results derived by the sentiment analysis engine and other analysis means.
[1346] "Advice and suggestions" are measures and instructions generated based on the analysis results and the situation.
[1347] "Storage" refers to keeping acquired data and analysis results for a certain period of time.
[1348] "Later analysis" refers to re-analysis at a later date using the stored data.
[1349] This invention is a system for analyzing conversations between aircraft pilots and air traffic controllers in real time to control aircraft operations and ensure safety. The system captures voice data, converts it into text data, translates it into multiple languages, performs sentiment analysis, controls automatic or manual operations based on the analysis results, and generates advice and suggestions in real time as needed. Furthermore, all data is stored for later analysis.
[1350] System configuration and operation
[1351] Acquiring audio data
[1352] The user (pilot or air traffic controller) reports the emergency situation by voice, and this voice data is picked up by the terminal in digital format through the microphone on board or the voice system in the air traffic control room, and then transmitted to the server.
[1353] Voice Recognition
[1354] The server passes the received voice data to a voice recognition engine. This engine generally uses high-performance voice recognition technology such as Google Speech-to-Text API or IBM Watson Speech to Text. The voice data is converted into text data. For example, a voice saying "There is an engine problem, please check!" is converted into text "There is an engine problem, please check!"
[1355] Text data translation
[1356] The server then automatically identifies the language of the text data. If the identified language is other than English, the server uses a built-in translation engine (e.g., Microsoft Translator or Google Translate API) to translate the text data into English. For example, the Japanese input "Engine has a malfunction, please check!" is translated into "Engine has a malfunction, please check!"
[1357] Emotion analysis
[1358] The server passes the translated text data and the original audio data to a sentiment analysis engine, which uses natural language processing and speech analysis technologies such as IBM Watson Natural Language Understanding and Microsoft Azure Text Analytics to assess the speaker's emotional state, for example, determining whether they are highly stressed or calm in an emergency.
[1359] Providing operational control and advice
[1360] The server sends a signal to the aircraft's automatic operation system based on the results of the emotion analysis and the content of the text data. If a high level of danger is detected, the server immediately executes safety operations. It also generates advice and suggestions based on the situation in real time and provides them to the user via the terminal. For example, advice such as "Perform engine restart procedures" is immediately provided as a voice message.
[1361] Data storage
[1362] All processed data is stored on a server as a detailed log, which is used for future analysis, system improvement, aircraft accident investigations, and operational optimization.
[1363] Specific examples
[1364] Emergency Scenarios
[1365] 1. The user (pilot) reports audibly, "An engine malfunction has occurred. Please check!"
[1366] 2. The device captures this audio and immediately transmits it digitally to the server.
[1367] 3. The server passes the received voice data to a voice recognition engine and converts it into text data.
[1368] 4. The server identifies the language of the text data and translates it into English using a translation engine if necessary.
[1369] 5. The server analyzes the translated text data and the original audio data using an emotion analysis engine to detect high stress levels.
[1370] 6. The server sends a signal to the automated operation system to immediately execute safety operations, and also creates a situation-specific advice message, such as "Perform engine restart procedure," as a voice message and provides it to the user via the terminal.
[1371] 7. All processing data will be stored on the server and used for future analysis and improvement.
[1372] Prompt Sentence Examples
[1373] Prompt: "Analyze the audio of an aircraft pilot reporting an emergency and provide a detailed description of how the automated flight control system responds and provides advice to the pilot."
[1374] The use of this system can significantly improve aviation safety and minimize communication errors caused by language differences or mental state during an emergency.
[1375] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1376] The flow of this system's program processing
[1377] Step 1:
[1378] The user (pilot or air traffic controller) reports information verbally. For example, in an emergency situation, they might say, "There's an engine malfunction, please check!" This is the input.
[1379] Step 2:
[1380] The terminal captures the user's voice in digital form through an onboard microphone or an audio system in the air traffic control room, and transmits this digital voice data to a server.
[1381] Input: User's voice
[1382] Output: Digital audio data
[1383] Step 3:
[1384] The server passes the received digital voice data to a voice recognition engine (e.g., Google Speech-to-Text API or IBM Watson Speech to Text) and converts it into text data.
[1385] Input: Digital audio data
[1386] Output: Text data: "An engine error has occurred. Please check!"
[1387] Step 4:
[1388] Identifying the language of server-generated text data, for example automatically recognizing that the text is in Japanese.
[1389] Input: Text data
[1390] Output: Language identification result (Japanese)
[1391] Step 5:
[1392] The server uses a built-in translation engine (e.g., Microsoft Translator or Google Translate API) to translate from the identified language to English. For example, "Engine has a malfunction, please check!" is translated to "Engine has a malfunction, please check!"
[1393] Input: Japanese text data
[1394] Output: English text data
[1395] Step 6:
[1396] The server passes the translated English text data and the original audio data to a sentiment analysis engine (e.g., IBM Watson Natural Language Understanding or Microsoft Azure Text Analytics) to evaluate the speaker's emotional state, e.g., to determine whether the speaker has a high stress level.
[1397] Input: English text and audio data
[1398] Output: Emotion analysis results (high stress)
[1399] Step 7:
[1400] Based on the results of the emotion analysis and the content of the text data, the server sends a signal to the aircraft's automatic operation system to perform appropriate safety operations, such as sending a signal to "strengthen autopilot."
[1401] Input: Sentiment analysis results and text data
[1402] Output: Control signal to automatic operation system
[1403] Step 8:
[1404] The server generates advice and suggestions based on the situation in real time and provides them to the user through the terminal, for example, generating a voice message saying "Please perform the engine restart procedure."
[1405] Input: Sentiment analysis results and situation information
[1406] Output: Advice and suggestions (voice message)
[1407] Step 9:
[1408] The server stores all transaction data in a detailed log, which is used for future analysis and system improvement.
[1409] Input: All analysis results and process logs
[1410] Output: Saved processed data
[1411] (Application example 1)
[1412] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1413] In on-site security work, security personnel often experience high levels of stress during emergencies, which can result in delayed responses and mistakes. Furthermore, communication barriers in multilingual environments and difficulty assessing situations in real time can also be a factor, potentially worsening dangerous situations.
[1414] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1415] In this invention, the server includes means for acquiring voice data, means for converting the acquired voice data into text data, means for translating the text data into multiple languages, means for performing emotion analysis of the text data and voice data, means for controlling automatic or manual operation based on the emotion analysis results, means for generating advice and suggestions in real time based on the analysis results and the situation, means for alerting the user when a high-stress state is detected, and means for saving all data for later analysis, thereby enabling quick and appropriate decisions to be made even in emergencies or high-stress situations.
[1416] The "means for acquiring voice data" refers to a device or software that captures the voice uttered by the user in real time and transmits the voice data to the system.
[1417] The "means for converting acquired voice data into text data" refers to a device or software that uses voice recognition technology to convert acquired voice data into a corresponding string of characters.
[1418] A "means for translating text data into multiple languages" is a translation engine or algorithm for converting text data into another language.
[1419] "Means for analyzing emotions in text data and voice data" refers to devices or software that use natural language processing and voice analysis techniques to evaluate the emotions of a speaker from text data or voice data.
[1420] "Means for controlling automatic or manual operations based on the results of emotion analysis" refers to devices or software that appropriately manage or correct automatic operations or manual instructions within the system based on the results of emotion analysis.
[1421] "Means for generating advice and suggestions in real time based on analysis results and the current situation" refers to devices or software that generate appropriate advice and suggestions in a timely manner based on data analysis results and the current situation.
[1422] "Means for alerting the user when a high stress state is detected" refers to devices or software that alert the user by issuing appropriate alerts or notifications when it is determined that the user's stress level is high.
[1423] "Means for storing all data and making it available for later analysis" refers to devices or software that store all data processed by the system and make it available for later analysis, performance evaluation, etc.
[1424] The present invention is a real-time assistant system for supporting security operations. This system has the functions of capturing voice data, converting it into text data, translating it into multiple languages, performing emotion analysis, providing advice in real time, and detecting high-stress situations.
[1425] System configuration
[1426] The system consists of the following main components:
[1427] 1. A means of acquiring voice data: Capture the user's voice in real time, specifically using a smartphone or microphone.
[1428] 2. Means for converting acquired voice data into text data: Convert the voice into text using voice recognition technology (for example, the SpeechRecognition library).
[1429] 3. Means for translating text data into multiple languages: Translate text data into other languages using a translation engine (e.g., Google Translate API).
[1430] 4. Means for sentiment analysis of text and voice data: Sentiment analysis is performed using natural language processing technology and voice analysis technology (e.g., Hugging Face Transformers).
[1431] 5. Means for controlling automatic or manual operations based on the results of sentiment analysis: Based on the results of sentiment analysis, operations within the system are appropriately managed or corrected.
[1432] 6. Means of generating advice and suggestions in real time based on analysis results and the current situation: Generate appropriate advice and suggestions based on the results of data analysis and the current situation.
[1433] 7. A method to alert users when high stress levels are detected: When stress levels are determined to be high, appropriate alerts and notifications will be provided.
[1434] 8. Means of storing all data for later analysis: All processed data will be stored and used for later analysis and performance evaluation.
[1435] System Operation
[1436] The specific operation of the system can be explained as follows:
[1437] Acquiring voice data and converting it to text data
[1438] The voice spoken by the user (security officer) is picked up through the smartphone's microphone. This voice data is converted into text using voice recognition technology. For example, a voice saying "Someone is breaking in! Call for help!" is converted into text data.
[1439] Text data translation
[1440] This text data is translated from Japanese to English using a language recognition function. Using a translation engine, "Someone is trespassing, please call for help immediately!" is converted to "Someone is trespassing, please call for help immediately!"
[1441] Emotion analysis
[1442] The translated text data and the original audio data are analyzed by a sentiment analysis model to determine the user's emergency state or high stress level. If a specific emotional state is detected, appropriate advice is generated.
[1443] Real-time advice and suggestions
[1444] Based on the results of sentiment analysis and the text content, the system generates real-time advice. For example, if a high stress state is detected, the system will provide advice such as "Be careful. Stress is increasing."
[1445] Data storage
[1446] All data is logged and made available for later analysis or performance evaluation.
[1447] Specific examples
[1448] For example, if a security worker says, "Someone is trespassing, please call for help immediately!", the system converts the speech into text and translates it into "Someone is trespassing, please call for help immediately!". If the emotion analysis determines that this is an "emergency situation," the system provides a real-time voice message with advice such as, "Be careful. Stress is building."
[1449] Prompt Sentence Examples
[1450] Voice: "Someone's breaking in, hurry up and call for help!"
[1451] Translation: "Someone is trespassing, please call for help immediately!"
[1452] Sentiment Analysis: "Emergency"
[1453] Advice: "Pay attention. Stress is building."
[1454] This allows security personnel to respond quickly and appropriately.
[1455] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1456] Step 1:
[1457] The device captures the user's voice in real time through a microphone. The input is the user's speech, and the output is digital voice data. This voice data is sent directly to the server.
[1458] Step 2:
[1459] The server converts the acquired voice data into text data using a voice recognition engine (for example, the SpeechRecognition library). The input is digital voice data, and the output is the corresponding text data. This conversion makes it easier to handle voice information as text information.
[1460] Step 3:
[1461] The server sends the generated text data to a translation engine (for example, Google Translate API) to translate it into another language. The input is text data in the original language, and the output is the translated text data. This translation allows for smooth support in multilingual environments.
[1462] Step 4:
[1463] The server performs sentiment analysis using the translated text data and the original voice data. Natural language processing techniques (e.g., Hugging Face Transformers) are used for sentiment analysis. The input is text and voice data, and the output is the analysis result of the emotional state. Based on the analysis result, the user's state of urgency and stress level are evaluated.
[1464] Step 5:
[1465] The server controls automatic or manual operation based on the results of emotion analysis. Specifically, if a high-stress state is detected, it automatically changes the system operation or issues instructions to the user. The input to this step is the emotion analysis result, and the output is appropriate operation control instructions.
[1466] Step 6:
[1467] The server generates advice and suggestions in real time based on the analysis results and the current situation. For example, if a high stress state is detected, advice such as "Be careful. Your stress is increasing" is generated. The input is the analysis results and situation information, and the output is the generated advice.
[1468] Step 7:
[1469] The server sends the generated advice to the terminal and notifies the user. The input is the generated advice, and the output is the notification received by the user.
[1470] Step 8:
[1471] All data is stored by the server. The stored data is used for later analysis and performance evaluation. The input is all the data during and after processing, and the output is the stored data. This storage allows for future optimization and improvement.
[1472] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1473] This invention is a system for analyzing conversations between aircraft pilots and air traffic controllers in real time, controlling aircraft operations, and ensuring safety. The system incorporates an emotion engine for recognizing the user's emotional state, thereby further improving safety. The following is a detailed description of an embodiment of the system.
[1474] System configuration
[1475] The system consists of the following main components:
[1476] 1. Device that acquires audio data
[1477] 2. A server that converts the acquired voice data into text data
[1478] 3. Server that translates text data into multiple languages
[1479] 4. Server that performs emotion analysis of text data and voice data
[1480] 5. Emotion Engine
[1481] 6. Server that controls automatic or manual operation based on the results of emotion analysis
[1482] 7. Server that generates advice and suggestions in real time based on the analysis results and the situation
[1483] 8. A server that stores all data and uses it for later analysis
[1484] System Operation
[1485] Acquiring audio data
[1486] Voice information from the user (pilot or air traffic controller) is acquired through a microphone on board the aircraft or through the voice system in the air traffic control room. The terminal captures this voice data and transmits it.
[1487] Voice Recognition
[1488] When the voice data arrives at the server, the voice recognition engine activates and converts it into text data. For example, a voice message such as "There is an engine error. Please check!" is converted into text format.
[1489] Text data translation
[1490] The server then identifies the language of this text data and, if necessary, translates it into the appropriate language using a built-in translation engine. For example, a Japanese sentence like "Engine has a malfunction, please check!" is translated into English as "Engine has a malfunction, please check!"
[1491] Emotion analysis
[1492] The translated text data and the original audio data are analyzed by a server, which uses natural language processing and speech analysis techniques to assess the speaker's emotional state, for example, determining whether they are highly stressed or calm in an emergency.
[1493] Use of emotion engine
[1494] The server is equipped with an emotion engine that can recognize the user's emotional state in more detail. This emotion engine analyzes both text and voice data and has the ability to highly evaluate the user's emotional state. This allows the system to determine the level of urgency corresponding to a specific emotional state and propose appropriate countermeasures accordingly.
[1495] Providing operational control and advice
[1496] The server controls autopilot or manual operation based on the emotion analysis results and text content. If it determines that the risk is high, it sends a signal to the aircraft's autopilot system to perform appropriate operations. It also generates advice and suggestions based on the situation in real time and provides them to the user via the terminal. For example, if an engine abnormality is detected, it immediately provides a voice message advising, "Please perform the engine restart procedure."
[1497] Data storage
[1498] The server stores all processing data, analysis results, audio data, and response measures as logs, which will be used later for detailed analysis and system improvement.
[1499] Specific examples
[1500] Emergency Scenarios
[1501] 1. The user (pilot) says, "There's an engine malfunction, please check!"
[1502] 2. The device captures this audio and sends it to the server.
[1503] 3. The server performs speech recognition and converts it into text.
[1504] 4. If necessary, the server translates the text into English.
[1505] 5. The server performs emotion analysis and detects that the pilot is in a high stress state.
[1506] 6. The emotion engine analyzes the user's stress level in detail and determines that the level of urgency is high.
[1507] 7. The server signals the autopilot system to execute a safety maneuver.
[1508] 8. The server generates an advice saying, "Autopilot has been strengthened due to an engine abnormality. Please check the engine instruments and perform the engine restart procedure." and notifies the user via the terminal.
[1509] 9. All processed data is stored on the server and used for later analysis.
[1510] The system can significantly improve aviation safety and minimize communication errors by providing countermeasures based on the user's emotional state.
[1511] The processing flow will be explained below.
[1512] Step 1:
[1513] The device captures the voice data of the pilot and air traffic controller through microphones on the plane or the audio system in the air traffic control room, and transmits the captured voice data to a server in real time.
[1514] Step 2:
[1515] The server receives the voice data. It converts the voice data into text data using a speech recognition engine. For example, a voice message such as "There is an engine error. Please check!" is converted into text format.
[1516] Step 3:
[1517] The server identifies the language of the converted text data, using a built-in language identification engine to determine in which language the text is written.
[1518] Step 4:
[1519] The server translates the text data into the specified language as needed. For example, the Japanese text "Engine has a malfunction, please check!" is translated into English as "Engine has a malfunction, please check!".
[1520] Step 5:
[1521] The server performs sentiment analysis based on the translated text data and the original audio data, using natural language processing and audio analysis technologies to assess the speaker's emotional state (e.g., stressed, overwhelmed, calm, etc.).
[1522] Step 6:
[1523] The emotion engine performs detailed analysis of text and voice data to intelligently assess the user's emotional state, for example detecting when a pilot is under high stress.
[1524] Step 7:
[1525] The server assesses the level of danger based on the results of sentiment analysis and the text content. If the level of danger is deemed high, the server sends a signal to the autopilot system to temporarily control the aircraft's operation. For example, if a pilot is under high stress, the autopilot system will automatically maintain the aircraft's altitude and speed.
[1526] Step 8:
[1527] The server generates optimal countermeasures in real time based on the analysis results and the situation. This includes specific procedures and advice. For example, advice such as "Perform the engine restart procedure" is generated.
[1528] Step 9:
[1529] The terminal then notifies the pilot or controller of the generated advice in real time as a voice or text message.
[1530] Step 10:
[1531] The server stores all voice data, text data, translation data, sentiment analysis results, and response actions as logs, which will be used later for further analysis and system improvement.
[1532] Example 2
[1533] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1534] In communications between aircraft pilots and air traffic controllers, differences in emotional state and language during emergencies can make it difficult to understand instructions, delaying appropriate responses. To solve this problem, real-time voice analysis, evaluation of emotional state, and proposal of appropriate countermeasures are required, but previous systems have found it difficult to perform these tasks in an integrated manner.
[1535] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for acquiring voice data, means for converting the acquired voice data into text data, means for translating the text data into multiple languages, means for performing emotion analysis of the text data and voice data, means for controlling automatic operation or manual operation based on the emotion analysis results, means for generating advice and suggestions in real time based on the analysis results and the situation, means for saving all data and using it for later analysis, and means for incorporating an emotion engine that highly evaluates the user's emotional state. This makes it possible to provide prompt measures based on accurate voice recognition and emotion analysis in real time.
[1536] "Voice data" refers to voices uttered by a user captured as digital signals.
[1537] "Means for acquiring" refers to the technology or device for acquiring voice data using a terminal or sensor within the domain.
[1538] "Text data" is voice data that has been analyzed and expressed as text information.
[1539] "Means for converting into text data" refers to technology or devices that convert acquired voice data into text format using natural language processing technology.
[1540] "Translation means" refers to the technology and software used to convert text data into other languages.
[1541] "Sentiment analysis" is the process of analyzing audio and text data to assess the emotional state of a speaker.
[1542] "Means for performing emotion analysis" refers to technologies or devices that use the algorithms or software required for emotion analysis.
[1543] An "emotion engine" is specialized software or algorithms that intelligently assess a user's emotional state.
[1544] "Means for controlling autopilot or manual operation" refers to technology or devices that automatically control aircraft operation based on the results of emotion analysis or in accordance with the pilot's instructions.
[1545] The "means for generating advice or suggestions" refers to technology or software that generates suggestions in real time to help the user take appropriate actions or make appropriate decisions based on the analysis results and the situation.
[1546] "Means for storing data and using it for later analysis" refers to the technology and devices that record all processing data and use it for later data analysis and system improvement.
[1547] This invention is a system for analyzing conversations between aircraft pilots and air traffic controllers in real time to control aircraft operations and ensure safety. The system uses multiple servers, terminals, and software to achieve functions such as voice data acquisition, voice recognition, text data translation, emotion analysis, autopilot or manual operation control, advice and suggestion generation, and data storage.
[1548] System configuration
[1549] The system consists of the following main components:
[1550] 1. Device that acquires audio data
[1551] 2. Server that converts voice data into text data
[1552] 3. Server that translates text data into multiple languages
[1553] 4. Server that performs emotion analysis of text data and voice data
[1554] 5. Emotion Engine
[1555] 6. Server that controls automatic or manual operation based on the results of emotion analysis
[1556] 7. Server that generates advice and suggestions in real time based on the analysis results and the situation
[1557] 8. A server that stores all data and uses it for later analysis
[1558] Hardware and software used
[1559] 1. Terminal: The aircraft's microphone and communication system are used to acquire voice data. Standard data communication protocols are used to transmit the voice data.
[1560] 2. Server: To analyze the voice data, a speech recognition engine such as Google Cloud Speech-to-Text or Amazon Transcribe is used.
[1561] 3. Text data translation: Translate text data into multiple languages using the Google Translate API, etc.
[1562] 4. Sentiment Analysis: Evaluate emotional state using IBM Watson Natural Language Understanding and Amazon's speech analysis technology.
[1563] 5. Emotional Engine: Utilizing specialized engines such as Affectiva to provide a sophisticated assessment of the user's emotional state.
[1564] 6. Autopilot and manual operation control: Works in conjunction with the aircraft's autopilot system (e.g., Honeywell system) to ensure safe operation.
[1565] 7. Advice and suggestion generation: Use dedicated advice generation algorithms to generate context-sensitive advice in real time.
[1566] 8. Data storage: All processed data, analysis results, audio data, and response actions will be stored using cloud storage for later analysis.
[1567] Specific examples
[1568] Emergency scenarios:
[1569] 1. The user (pilot) says, "There's an engine malfunction, please check!"
[1570] 2. The onboard microphone captures the audio and sends the data to the device.
[1571] 3. The device sends the audio data to the server.
[1572] 4. The server receives the audio data and calls the Google Cloud Speech-to-Text service.
[1573] 5. The server converts the voice data into text and generates the text "An engine error has occurred, please check!"
[1574] 6. The server translates the text data into English using the Google Translate API. "Engine has a malfunction, please check!" becomes "Engine has a malfunction, please check!"
[1575] 7. The server sends the text data to IBM Watson Natural Language Understanding for sentiment analysis.
[1576] 8. The server sends the audio data to Amazon Transcribe, which analyzes the emotional state of the audio.
[1577] 9. The emotion engine on the server integrates the text and voice data to perform a detailed emotion assessment.
[1578] 10. The emotion engine determines that the user's stress level is "high."
[1579] 11. The server sends a signal to the autopilot system, generating the message "Perform engine restart procedure."
[1580] 12. The device will announce this message to the user audibly.
[1581] 13. The server records the processed data, analysis results, audio data, and response measures as logs and stores them in a database.
[1582] 14. The stored data can be later analyzed to help improve the system.
[1583] Prompt Sentence Examples
[1584] "Please explain the steps your system takes when a user says, 'There's an engine malfunction, please check!'"
[1585] In this way, the system significantly improves aircraft safety, can provide countermeasures based on the user's emotional state, and minimizes communication errors.
[1586] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1587] Step 1:
[1588] The voice of the user (pilot or air traffic controller) is acquired through an on-board microphone or the voice system in the air traffic control room. This captures the user's real-time voice information as digital voice data. The input is the user's voice, and the output is digital voice data.
[1589] Step 2:
[1590] The device sends the captured audio data to a server. The audio data is transferred to the server via the Internet or a dedicated communication line. The input is the captured digital audio data, and the output is the transferred audio data.
[1591] Step 3:
[1592] The server receives the voice data and calls the Google Cloud Speech-to-Text service to perform speech recognition. The voice data is analyzed and converted into text data. For example, a speech saying "There is an engine problem, please check!" is converted into text "There is an engine problem, please check!" The input is voice data, and the output is text data.
[1593] Step 4:
[1594] The server identifies the language of the generated text data. If necessary, it uses the Google Translate API to translate the text into other languages. For example, text entered in Japanese is translated into English. The input is the text data, and the output is the translated text data.
[1595] Step 5:
[1596] The server sends the text data to IBM Watson Natural Language Understanding for sentiment analysis. The text data is used to evaluate the speaker's emotional state. For example, the text content can be used to analyze a pilot's stress level and urgency. The input is text data, and the output is an evaluation of the emotional state.
[1597] Step 6:
[1598] The server sends the audio data to Amazon Transcribe, which analyzes the tone and pitch of the voice. This allows for a detailed assessment of the speaker's emotional state and urgency from the audio data. The input is the audio data, and the output is the emotional assessment result from the audio analysis.
[1599] Step 7:
[1600] The emotion engine in the server integrates the results of the analysis of the text data and voice data to perform a detailed emotion evaluation. The emotion engine highly evaluates the user's emotional state and determines, for example, that the user is in a high-stress state. The inputs are the emotion analysis results and the voice analysis results, and the output is the integrated emotion evaluation result.
[1601] Step 8:
[1602] The server generates a control signal to an autopilot system or manual operation based on the emotion analysis result, for example, to signal an aircraft's autopilot system to execute emergency measures. The input is the integrated emotion evaluation result, and the output is a control signal.
[1603] Step 9:
[1604] The server generates appropriate advice or suggestions in real time based on the analysis results and the situation, and notifies the user through the terminal. For example, it generates advice such as "Please perform the engine restart procedure." The input is the emotion analysis results and situation information, and the output is the generated advice.
[1605] Step 10:
[1606] The server records and stores all processing data, analysis results, audio data, and response measures as logs. The stored data is used for later detailed analysis and system improvement. The input is the processing data and analysis results, and the output is the stored log data.
[1607] (Application example 2)
[1608] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1609] Conventional call center systems have the problem of being unable to grasp the emotional state of the operator and customer in real time and provide immediate solutions. This has led to problems such as customers' dissatisfaction not being resolved properly and operators feeling high levels of stress. This leads to problems such as lower customer satisfaction and a decline in the work efficiency of operators.
[1610] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1611] In this invention, the server includes means for acquiring voice data, means for converting the acquired voice data into text data, means for translating the text data into multiple languages, means for performing emotion analysis of the text data and voice data, means for controlling automatic or manual operations based on the emotion analysis results, means for generating advice and suggestions in real time based on the analysis results and the situation, means for triage and proposing countermeasures based on the emotional state, and means for saving all data and using it for later analysis. This makes it possible to grasp the emotional states of customers and operators in real time and immediately provide appropriate countermeasures.
[1612] "Audio data" refers to sound information in digital or analog form collected through audio equipment such as a microphone.
[1613] "Text data" is information obtained by converting voice data into a character string format.
[1614] A "multilingual translation method" is a technique or process for converting text data into different languages.
[1615] "Emotion analysis" is a technique that analyzes audio and text data to assess a speaker's emotional state.
[1616] "Autonomous operation" means that the system performs a specific action without human intervention.
[1617] "Manual operation" means direct human involvement in operating a system.
[1618] "Means for generating advice and suggestions in real time" refers to techniques and methods that provide immediate advice to operators or users based on the current situation.
[1619] "Triage" is the prioritization of response measures based on urgency and importance.
[1620] A "solution" is an action or suggestion to be taken in response to a particular situation.
[1621] "Storing data" means recording collected data on a storage medium so that it can be used later.
[1622] "Use for later analysis" means to conduct analysis based on the stored data for future consideration and improvement.
[1623] MODE FOR CARRYING OUT THE INVENTION
[1624] The present invention is a system for supporting appropriate responses by analyzing the emotional states of customers and agents in a call center in real time. An embodiment of this system will be described in detail below.
[1625] System configuration
[1626] The system consists of the following main components:
[1627] 1. Means of capturing voice data: microphones and telephone systems to capture conversations between customers and operators.
[1628] 2. A means of converting the acquired voice data into text data: a speech recognition engine that converts the voice into a string format (e.g., Google Speech-to-Text API).
[1629] 3. A means of translating text data into multiple languages: A system that translates text into the required languages (e.g., Microsoft Azure Cognitive Services).
[1630] 4. Means of sentiment analysis of text and voice data: Software that analyzes the emotional state of customers and operators (e.g., IBM Watson Tone Analyzer).
[1631] 5. Means for controlling automatic or manual operations based on the results of emotion analysis: A control system for performing processing according to emotional states.
[1632] 6. Means of generating advice and suggestions in real time based on analysis results and situations: An engine that generates real time advice and suggestions to operators.
[1633] 7. Triage and suggest responses based on emotional state: A mechanism for prioritizing responses based on urgency and importance.
[1634] 8. A means of storing all data and making it available for later analysis: a database to store the collected and analyzed data.
[1635] System Operation
[1636] Audio Acquisition
[1637] Voice information from users (customers and operators) is acquired through a microphone or telephone system. The terminal captures this voice data and sends it to the server.
[1638] Voice Recognition
[1639] The voice data received by the server is converted into text data by a voice recognition engine (Google Speech-to-Text API).
[1640] Text data translation
[1641] The server then converts this text data into the appropriate language if it needs to be translated using a multi-language translation means.
[1642] Emotion analysis
[1643] The translated text data and the original audio data are then subjected to natural language processing and speech analysis techniques to assess the user's emotional state, using analysis engines such as IBM Watson Tone Analyzer.
[1644] Detailed analysis by Emotion Engine
[1645] The emotion engine performs advanced analysis of text and voice data to determine specific emotional states and their urgency.
[1646] Providing operational control and advice
[1647] Based on the results of the emotion analysis, the server controls automatic or manual operations and generates and notifies advice and suggestions in real time. Appropriate advice is provided to the operator in voice or text format.
[1648] Data storage
[1649] The server stores all audio data, text data, and analysis results in a database for later analysis.
[1650] Specific examples
[1651] Emergency Scenarios
[1652] 1. A customer says, "I recently purchased a broken item and I'm very unhappy with it. I'd like to process the return immediately."
[1653] 2. The device captures this audio and sends it to the server.
[1654] 3. The server performs speech recognition and converts it into text.
[1655] 4. If necessary, the server translates the text.
[1656] 5. The server performs sentiment analysis and detects that the customer is in a high stress state.
[1657] 6. The emotion engine analyzes the customer's stress level in detail and determines that the situation is urgent.
[1658] 7. The server advises the operator, "The customer is unhappy. Please stay calm and explain with specific examples."
[1659] 8. All processed data is stored on the server and used for later analysis.
[1660] This system will significantly improve the quality and efficiency of customer service at call centers, increasing customer satisfaction. It will also reduce stress on operators and contribute to improved work efficiency.
[1661] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1662] Step 1:
[1663] The user (customer and operator) starts a conversation. The terminal captures the voice data through a microphone or telephone system. This captured voice data is recorded in digital format and sent to the server. The input here is the raw voice data, and the output is the audio file sent to the server.
[1664] Step 2:
[1665] The server sends the received voice data to the Google Speech-to-Text API for speech recognition. In this process, the voice data is converted into text data in string format. The input is the voice data, and the output is the converted text data. Specifically, the voice data undergoes signal processing and pattern matching.
[1666] Step 3:
[1667] The server uses Microsoft Azure Cognitive Services to translate the acquired text data into multiple languages as needed. In this step, the language of the text data is identified and converted into the appropriate language. The input is text data, and the output is translated text data. The specific operation utilizes natural language processing algorithms.
[1668] Step 4:
[1669] The server then sends the translated text data and the original audio data to the IBM Watson Tone Analyzer for sentiment analysis. The input is text data and audio data, and the output is the sentiment analysis results. This function evaluates the speaker's emotional state based on the tone and content of the audio.
[1670] Step 5:
[1671] The server further evaluates the emotion analysis results to determine the specific emotional state and its urgency. This is done by the emotion engine, which understands the user's emotional state. The input is the emotion analysis results, and the output is detailed emotional state and urgency information. Specific operations include integrating and evaluating the analysis results.
[1672] Step 6:
[1673] The server generates and notifies the operator in real time with advice and suggestions based on the results of emotion analysis. The input is the detailed emotional state and the customer's comments, and the output is advice and suggestions provided to the operator. Specifically, the system generates appropriate countermeasures and notifies the operator via voice or text.
[1674] Step 7:
[1675] The server stores all data in a database for later analysis. The input is all processed data (audio, text, analysis results, etc.), and the output is the stored data. Specific operations include recording and storing data.
[1676] Prompt Sentence Examples
[1677] Speech input: "I'm very unhappy with the product I recently purchased, which is broken. I'd like to process the return immediately."
[1678] Expected output: "The customer is unhappy. Please stay calm and provide specific examples. We'll expedite the return process."
[1679] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1680] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1681] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1682] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1683] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1684] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1685] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1686] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1687] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1688] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1689] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1690] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1691] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1692] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1693] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1694] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1695] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1696] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1697] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1698] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1699] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1700] The following is further disclosed regarding the above embodiment.
[1701] (Claim 1)
[1702] means for acquiring audio data;
[1703] A means for converting the acquired voice data into text data;
[1704] a means for translating text data into multiple languages;
[1705] means for performing sentiment analysis on text data and audio data;
[1706] A means for controlling automatic or manual operation based on the emotion analysis results;
[1707] A means of generating advice and suggestions in real time based on the analysis and the situation;
[1708] A means to store all data for later analysis;
[1709] A system including:
[1710] (Claim 2)
[1711] 10. The system of claim 1, further comprising means for identifying the language of the text data.
[1712] (Claim 3)
[1713] 10. The system of claim 1, further comprising means for analyzing the tone and pitch of the audio data.
[1714] "Example 1"
[1715] (Claim 1)
[1716] means for acquiring audio data;
[1717] A means for converting the acquired voice data into text data;
[1718] a means for translating text data into multiple languages;
[1719] means for performing sentiment analysis on text data and audio data;
[1720] a means for controlling automatic or manual operation based on the emotion analysis result;
[1721] A means of generating advice and suggestions in real time based on the analysis and the situation;
[1722] A means to store all data for later analysis;
[1723] means for transmitting the captured audio data in digital form;
[1724] means for converting voice data into text data using a voice recognition engine;
[1725] means for passing the voice data and text data to a sentiment analysis engine;
[1726] means for assessing an emotional state using an emotion analysis engine;
[1727] means for transmitting a signal to an automated operating system;
[1728] means for generating and transmitting a voice message based on the analysis results;
[1729] A system including:
[1730] (Claim 2)
[1731] 10. The system of claim 1, further comprising means for identifying the language of the text data and translating it into another language using a translation engine, if necessary.
[1732] (Claim 3)
[1733] 10. The system of claim 1, further comprising means for analyzing the tone and pitch of the speech data using natural language processing and speech analysis techniques to assess the emotional state of the speaker.
[1734] "Application Example 1"
[1735] (Claim 1)
[1736] means for acquiring audio data;
[1737] A means for converting the acquired voice data into text data;
[1738] a means for translating text data into multiple languages;
[1739] means for performing sentiment analysis on text data and audio data;
[1740] a means for controlling automatic or manual operation based on the emotion analysis result;
[1741] A means of generating advice and suggestions in real time based on the analysis and the situation;
[1742] A means for alerting the user when a high stress state is detected;
[1743] A means to store all data for later analysis;
[1744] A system including:
[1745] (Claim 2)
[1746] 10. The system of claim 1, further comprising means for identifying the language of the text data.
[1747] (Claim 3)
[1748] 10. The system of claim 1, further comprising means for analyzing the tone and pitch of the audio data.
[1749] "Example 2: Combining Emotion Engines"
[1750] (Claim 1)
[1751] means for acquiring audio data;
[1752] A means for converting the acquired voice data into text data;
[1753] a means for translating text data into multiple languages;
[1754] means for performing sentiment analysis on text data and audio data;
[1755] A means for controlling automatic or manual operation based on the emotion analysis results;
[1756] A means of generating advice and suggestions in real time based on the analysis and the situation;
[1757] A means to store all data for later analysis;
[1758] means for incorporating an emotion engine that intelligently assesses the user's emotional state;
[1759] A system including:
[1760] (Claim 2)
[1761] 10. The system of claim 1, further comprising means for identifying the language of the text data.
[1762] (Claim 3)
[1763] 10. The system of claim 1, further comprising means for analyzing the tone and pitch of the audio data.
[1764] "Application example 2 when combining emotion engines"
[1765] (Claim 1)
[1766] means for acquiring audio data;
[1767] A means for converting the acquired voice data into text data;
[1768] a means for translating text data into multiple languages;
[1769] means for performing sentiment analysis on text data and audio data;
[1770] a means for controlling automatic or manual operations based on the sentiment analysis results;
[1771] A means of generating advice and suggestions in real time based on the analysis and the situation;
[1772] a means of triage and suggesting responses based on emotional state;
[1773] A means to store all data for later analysis;
[1774] A system including:
[1775] (Claim 2)
[1776] 10. The system of claim 1, further comprising means for identifying the language of the text data.
[1777] (Claim 3)
[1778] 10. The system of claim 1, further comprising means for analyzing the tone and pitch of the audio data. [Explanation of symbols]
[1779] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. means for acquiring audio data; A means for converting the acquired voice data into text data; a means for translating text data into multiple languages; means for performing sentiment analysis on text data and audio data; A means for controlling automatic or manual operation based on the emotion analysis results; A means of generating advice and suggestions in real time based on the analysis and the situation; A means to store all data for later analysis; A system including:
2. 10. The system of claim 1, further comprising means for identifying the language of the text data.
3. 2. The system of claim 1, further comprising means for analyzing the tone and pitch of the audio data.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A