system

A voice-based system converts worker input into text, analyzes comprehension and emotions, and provides adaptive instructions to improve communication and safety in on-site work environments.

JP2026073495APending Publication Date: 2026-05-01SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-18
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Communication between on-site workers and managers is inefficient, leading to misunderstandings and increased risks of accidents due to unclear instructions and ungrasped understanding, particularly when stress levels are high.

Method used

A system that converts voice input from workers into text data, analyzes intentions and comprehension levels, generates tailored instructions, and provides feedback loops to ensure accurate and stress-aware communication.

Benefits of technology

Enhances communication clarity, reduces misunderstandings, and creates a safer work environment by providing adaptive instructions based on real-time understanding and emotional state analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026073495000001_ABST
    Figure 2026073495000001_ABST
Patent Text Reader

Abstract

Provide a system. 【Solution means】 Means for receiving voice input from an operator, Means for converting the received voice input into text data, Means for analyzing the operator's intention from the text data, Means for estimating the operator's level of understanding based on the analysis result, Means for generating additional instructions according to the level of understanding, Means for outputting the generated instructions to the operator as voice, Means for re-analyzing the feedback from the operator and generating instructions again, A system including
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0004] , , , ,

[0005] , , , , ,

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, the method including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance that responds to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In the communication between on-site workers and managers, it is difficult to accurately convey instructions and grasp the degree of understanding, resulting in problems such as misunderstandings and inefficiencies in work. In particular, when work progresses without accurately grasping the decrease in the degree of understanding and the state of tension, there is a problem that the risk of accidents and work mistakes increases. To solve this, a system that provides appropriate instructions according to the situation of the workers is required.

Means for Solving the Problems

[0005] This invention provides a system that receives voice input from a worker and converts it into text data, thereby accurately analyzing the worker's intentions. Based on the analysis results, it estimates the worker's level of understanding, automatically generates additional instructions tailored to that level of understanding, and communicates them to the worker via voice. Furthermore, by re-analyzing the worker's feedback and providing optimized instructions, it realizes a smooth and safe working environment. This system can also estimate the worker's stress level through voice tone analysis, enabling more accurate support.

[0006] "Voice input" refers to the voice signal from the user, which is data that the system receives and processes.

[0007] "Text data" refers to string data obtained by converting voice input using natural language processing technology.

[0008] "Analysis" refers to the process of processing information contained in text data to understand the user's intentions and needs.

[0009] "Comprehension level" is an indicator that shows how accurately a user understands the information and instructions provided.

[0010] "Additional instructions" are supplementary guidance generated and provided by the system based on the user's level of understanding and status.

[0011] "Voice output" refers to the audio signal used to communicate instructions generated by the system to the user.

[0012] "Feedback" refers to the information or reactions that a user returns to a system as a response.

[0013] "Stress level" refers to the degree of psychological state estimated through analysis of the user's voice tone and dialogue. [Brief explanation of the drawing]

[0014] [Figure 1]It is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] It is a conceptual diagram showing an example of the main functions of a data processing device and a smart device according to the first embodiment. [Figure 3] It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which a plurality of emotions are mapped. [Figure 10] It shows an emotion map to which a plurality of emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when an emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when an emotion engine is combined.

MODE FOR CARRYING OUT THE INVENTION

[0015] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0016] First, the terms used in the following description will be explained.

[0017] In the following embodiments, the numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0018] In the following embodiments, the numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor. <00001​​​​​​​​ In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0022] [First Embodiment]

[0023] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0024] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0025] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0026] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0027] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0028] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0029] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0030] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0031] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0032] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0033] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0034] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0035] This invention is implemented as a voice bot system that streamlines communication between field workers and managers. Specifically, a server uses speech recognition technology to convert voice input from workers (users) into text data in real time. The generated text data is analyzed using natural language processing. This analysis makes it possible to accurately understand the user's intentions and work status.

[0036] Next, the server applies a deep learning algorithm to estimate the user's level of understanding based on the analysis results. During this process, it also considers speech characteristics such as tone and speed to evaluate the stress level. This determines how well the user understands the instructions and how confidently they can proceed with the task.

[0037] Based on the user's state, the server generates appropriate additional instructions. These instructions are communicated to the user using speech synthesis technology to facilitate more effective communication. The user provides feedback on the voice-output instructions. This feedback is also analyzed by the server and, if necessary, leads to the generation of new instructions.

[0038] As a concrete example, when a worker inquires about the operating procedure for a specific piece of equipment, the server generates instructions that concisely explain the procedure. If signs of stress are detected in the worker's feedback, the server adds more detailed steps and precautions to the instructions and provides them again. In this way, the system is implemented to aid user understanding and ensure transparency in communication.

[0039] The following describes the processing flow.

[0040] Step 1:

[0041] The server receives voice input from the user and records the received audio in real time. The audio data is stored in a buffer for subsequent processing.

[0042] Step 2:

[0043] The device's speech recognition engine converts received speech data into text data. This conversion process ensures accurate transcription of the speech into text.

[0044] Step 3:

[0045] The server passes the text data to a natural language processing engine, which analyzes the user's speech. Based on the analysis results, it identifies what information or instructions the user is seeking.

[0046] Step 4:

[0047] The server estimates the user's level of understanding based on the analyzed information. Here, it evaluates the level of understanding based on the user's vocal characteristics and word choices, referencing past data and learning models.

[0048] Step 5:

[0049] The server generates appropriate additional instructions, taking into account the estimated level of understanding and stress levels. These instructions are customized to fit the context of the task.

[0050] Step 6:

[0051] The device communicates the generated instructions to the user using speech synthesis technology. This voice output is provided in a clear and easy-to-understand manner.

[0052] Step 7:

[0053] The user provides feedback on the given instructions. The server converts this feedback back into text, analyzes it, and uses it to generate new instructions.

[0054] Step 8:

[0055] The server repeats the above processing cycle as needed, continuing to assist the user in reaching the desired solution.

[0056] (Example 1)

[0057] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0058] There is a need to improve the efficiency of communication between field workers and managers. In particular, a challenge is to provide an environment where workers can accurately understand the instructions they receive and proceed with their work without stress. There is also the problem of difficulty in correcting instructions and incorporating feedback in real time.

[0059] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0060] In this invention, the server includes means for receiving an audio signal, means for converting the received audio signal into text data, and means for analyzing the text data to identify the worker's intent. This allows the worker to instantly transmit their intent to the server through voice input and receive appropriate instructions and feedback according to the level of understanding.

[0061] An "audio signal" is an electrical or digital representation of information transmitted by sound.

[0062] "Text data" refers to text data that represents audio, handwritten text, and other information in a standardized format.

[0063] "Intention" refers to the purpose or wish that the worker is trying to convey through voice input, and understanding that intention leads to more efficient communication.

[0064] "Comprehension level" is an indicator that shows how correctly a worker understands and can execute the instructions and information provided to them.

[0065] A "signal" refers to additional information or instructions given to a worker, and is communicated through voice or visual means.

[0066] "Psychological stress level" refers to the degree of stress and anxiety inferred from a worker's tone of voice and speed, and serves as a criterion for determining whether to provide appropriate support in the work environment.

[0067] "Natural language processing techniques" refer to technologies and approaches for enabling computers to understand and process human language.

[0068] This invention is implemented as a voice bot system to streamline communication between field workers (users) and administrators. The system functions by having a terminal receive the voice spoken by the user at the field site, and a server process that voice.

[0069] First, users perform voice input using a dedicated mobile device or headset. The device is equipped with a high-sensitivity microphone that can eliminate background noise and transmit clear audio to the server. This allows users to perform highly accurate voice input while moving freely.

[0070] The server uses "speech recognition software" to convert speech data to text in real time. This conversion utilizes common APIs, ensuring rapid speech-to-text conversion.

[0071] The server that acquires the text data uses "natural language processing software" to analyze the user's intent. For example, if a user enters the prompt "Tell me how to replace the part right now," the server can accurately interpret that intent and provide an answer by searching for relevant information. This process employs natural language processing technology and involves detailed analysis.

[0072] Furthermore, the server uses a "deep learning model" to assess the user's psychological load from the tone and speed of their voice. This allows it to estimate the user's level of stress and anxiety and generate additional instructions tailored to their understanding. For example, if the server determines that the user is having difficulty operating the system, it will generate cues that include more detailed explanations and priority steps.

[0073] Ultimately, the server uses "speech synthesis software" to provide the generated instructions as audio signals, which are then transmitted to the user. The user continues working based on the received audio and provides feedback as needed. This feedback is also processed by the server, leading to the provision of new information and improvements to the instructions.

[0074] A concrete example is a process where, after a worker on-site reports a machine malfunction, the server asks questions to identify the cause of the malfunction. Examples of prompts used in this process include, "The machine is vibrating abnormally; please diagnose the cause." This allows users to efficiently resolve problems while receiving support.

[0075] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0076] Step 1:

[0077] The user provides voice input through a dedicated terminal. The input consists of the user's questions and instructions, which are received as audio signals by the terminal's microphone. These audio signals are then converted into digital format within the terminal. The output is digitized audio data.

[0078] Step 2:

[0079] The terminal sends digitized audio data to the server. The server receives this data and uses "speech recognition software" to convert it into text data. During the conversion of audio data to text format, phoneme analysis is performed, and the output is generated as text. The output is text data.

[0080] Step 3:

[0081] The server analyzes the text data using "natural language processing software." Contextual analysis and keyword identification are performed to extract intent from the input text data. Data with clearly defined user intent and purpose is then output.

[0082] Step 4:

[0083] Based on the analyzed intent, the server uses a "deep learning model" to evaluate the user's understanding. This process also takes speech properties into consideration. Data regarding the estimated understanding and stress level is output.

[0084] Step 5:

[0085] The server generates additional instructions based on the user's level of understanding. Information processing is performed to format the generated instructions into an easily understandable form. The output is a written instruction provided to the user.

[0086] Step 6:

[0087] The server uses "speech synthesis software" to convert the generated instructions into audio signals. At this stage, the process generates natural-sounding, easily recognizable speech from the text data. The output is audio data transmitted to the user through the terminal.

[0088] Step 7:

[0089] The user performs tasks based on the voice instructions received and provides voice feedback as needed. This feedback is then captured again as digital voice data on the device. The output is the voice data of the feedback.

[0090] Step 8:

[0091] The server receives the audio data of the feedback, analyzes it again, and generates new instructions if necessary. This allows the system to continuously improve communication. The output consists of new instructions and feedback analysis results.

[0092] (Application Example 1)

[0093] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0094] In complex on-site tasks, smooth communication between workers and systems is essential. However, conventional systems based on voice input have difficulty accurately grasping the worker's level of understanding and emotional state. As a result, they are not adequately supporting workers in preventing misunderstandings and ensuring they can work efficiently and with confidence.

[0095] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0096] In this invention, the server includes a medium for acquiring voice data from a worker, means for converting the voice data into text information, and means for recognizing the worker's intentions from the text information. This makes it possible to accurately evaluate the worker's level of understanding and emotional state, and to provide appropriate additional information and instructions in a timely manner.

[0097] A "worker" is a person who operates equipment and facilities and performs tasks in a factory or on-site.

[0098] "Voice data" refers to data of acoustic signals acquired through the voice of a worker.

[0099] "Medium" refers to a physical or virtual device or means used to acquire or provide information.

[0100] "Textual information" refers to the textual representation of language converted from audio data.

[0101] "Means of recognition" refer to methods and technologies for analyzing acquired data and understanding its meaning and purpose.

[0102] "Comprehension level" is an indicator that shows how accurately a worker understands the instructions and information they are given.

[0103] "Means of evaluation" refer to methods and techniques for analyzing data and situations and making judgments about their quality and characteristics.

[0104] "Additional information" refers to detailed information provided to assist workers in understanding the work and to facilitate the smooth progress of the task.

[0105] A "device" is a mechanical or electronic system built to perform a specific function or task.

[0106] A "response" is a reaction or reply that an operator gives to instructions or information from a system.

[0107] "Analysis" is the process of examining data in detail to clarify its structure and meaning.

[0108] "Stress level" refers to the degree of burden or tension that workers feel towards their work and environment.

[0109] "To analyze with high accuracy" means to analyze data and information with great precision and detail.

[0110] This invention will now be described in terms of embodiments for carrying it out. The main element of the invention is an interactive system for smooth information exchange between a worker and a system. This system is realized by acquiring voice data via smart glasses or a headset worn by a worker on site and processing that voice on a server.

[0111] The server converts audio data acquired from the smart glasses' microphone into text in real time using the Google® Speech-to-Text API. Then, it analyzes this text using NLTK (Natural Language Toolkit) to recognize the worker's intent. Based on this analysis, a deep learning model built with PyTorch is used to evaluate the worker's comprehension and stress levels. Based on this evaluation, the server generates additional information tailored to the worker and communicates it to them verbally using Google Text-to-Speech.

[0112] This process continuously provides optimal instructions by constantly acquiring new responses from workers, analyzing them again, and re-evaluating their level of understanding and stress. As a result, workers receive specific instructions tailored to their situation, enabling them to work efficiently and with peace of mind.

[0113] As a concrete example, consider a scenario where an operator troubleshoots a metalworking machine. The operator asks a question by voice, and the question is immediately analyzed to provide specific machine operating procedures and precautions. If the operator shows any signs of uncertainty, more detailed information will be provided later.

[0114] An example of a prompt message is, "Please provide detailed instructions for troubleshooting the metalworking machine." This prompt allows the system to quickly and accurately convey the necessary information, enabling the operator to efficiently resolve the problem accordingly.

[0115] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0116] Step 1:

[0117] The terminal acquires the worker's voice as audio data from the microphone in the smart glasses. The input is the worker's voice, and this audio data serves as the trigger to start the program.

[0118] Step 2:

[0119] The server converts the acquired audio data into text using the Google Speech-to-Text API. The input is audio data, and the output is text data. This process transforms the audio information into parseable text.

[0120] Step 3:

[0121] The server analyzes the converted character information using NLTK to recognize the worker's intent. The input is character information, and the output is data indicating the worker's intent. This analysis process clarifies the information and instructions the worker is seeking.

[0122] Step 4:

[0123] The server evaluates the worker's level of understanding and stress through a deep learning model using PyTorch. Inputs are intent recognition results and emotional data obtained from the worker's voice, while outputs are understanding and stress indicators. This allows the worker's state to be quantified, enabling appropriate responses based on that situation.

[0124] Step 5:

[0125] Based on the evaluation results, the server utilizes a generative AI model to generate optimal additional information and instructions for the worker. The inputs are comprehension levels and stress indicators, while the output is the content of the instructions for the worker. Appropriate content is provided according to the worker's condition.

[0126] Step 6:

[0127] The server uses Google Text-to-Speech to synthesize the generated instructions into speech and sends them to the terminal. The input is the generated instructions, and the output is the audio data. The terminal then transmits this audio to the worker.

[0128] Step 7:

[0129] The user provides feedback to the system again, and this voice feedback triggers a return to step 1 of the next cycle. The input is the worker's new voice feedback, which triggers the process of being analyzed again.

[0130] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0131] This invention is implemented as a voice bot system for improving communication in field work. By incorporating an emotion engine, this system not only provides instructions but also generates adaptive instructions that respond to the user's emotions and stress levels.

[0132] Specifically, the server receives voice input from the user and converts it into text data in real time using speech recognition. The transcribed text is then analyzed by a natural language processing engine to accurately understand the user's intentions and needs.

[0133] Furthermore, the server uses an emotion engine to identify the emotional characteristics of the speech. In this process, various speech features such as tone, speed, and intonation are considered to determine the user's psychological state. Meanwhile, the level of understanding is evaluated based on the analysis results in response to the user's questions and requests.

[0134] By combining emotional information evaluated based on the emotion engine with an assessment of comprehension, the system generates the most appropriate additional instructions. These instructions are customized to take into account the user's psychological state and work context. The terminal communicates these instructions to the user using speech synthesis and subsequently analyzes the received feedback.

[0135] For example, if a user makes a mistake in executing a work procedure, the server detects the user's anxiety and tension through its emotion engine. Based on this emotional information, it generates instructions in a calm tone to alleviate the anxiety and reinstructs the user to follow the correct procedure in a reassuring manner. This system is highly responsive to the user's psychology and state, enabling it to support a safe and secure work environment.

[0136] The following describes the processing flow.

[0137] Step 1:

[0138] The user provides voice input to the device. The device receives this voice input as digital audio data.

[0139] Step 2:

[0140] The server provides the received audio data to the speech recognition engine in real time, where it is converted into text data. The speech recognition engine accurately processes the conversion of speech into text.

[0141] Step 3:

[0142] The server passes the converted text data to a natural language processing engine, which analyzes the user's intent. This analysis helps understand what the user wants.

[0143] Step 4:

[0144] The server uses an emotion engine to analyze the audio data. The emotion engine evaluates the user's emotional state based on factors such as tone, speed, and rhythm of the voice.

[0145] Step 5:

[0146] The server combines the user's intent analysis results and emotion evaluation results to estimate their level of understanding and psychological state. Based on this estimation, it determines what kind of instructions are needed.

[0147] Step 6:

[0148] The server generates additional instructions tailored to the user based on their estimated level of understanding and emotional state. These instructions are designed to maximize the user's safety and efficiency.

[0149] Step 7:

[0150] The terminal uses speech synthesis to communicate additional instructions sent from the server to the user. The user understands the next action through the voice instructions.

[0151] Step 8:

[0152] The user provides feedback on the additional instructions. This feedback is converted back into text on the server and used to analyze the user's understanding and the need for continued support.

[0153] Step 9:

[0154] If necessary, the server optimizes the process, provides more specific instructions, and delivers them again to the user via the terminal. This iterative process continuously supports the user and reduces misunderstandings.

[0155] (Example 2)

[0156] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0157] There is a lack of instructions that take into account the psychological state of workers in on-site work, and it is necessary to reduce inefficiencies and errors caused by misunderstandings and stress. Furthermore, it is necessary to interpret workers' intentions with high accuracy and provide appropriate support accordingly.

[0158] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0159] In this invention, the server includes means for acquiring voice information of the worker, means for identifying emotional characteristics from the voice information, and means for adjusting the content of instructions based on the identified emotional information. This makes it possible to provide instructions that are adaptive and reassuring to the worker, in accordance with their psychological state.

[0160] "Audio information" refers to all audio data emitted by workers, and is material that is analyzed by speech recognition.

[0161] "Textual information" refers to text data converted from audio information and is used to interpret the worker's intent.

[0162] "Interpreting intent" is the process of identifying the purpose or requirements from what the worker says.

[0163] "Understanding level" is an indicator that shows how well a worker understands the information and instructions provided.

[0164] "Additional information" refers to supplementary instructions or explanations generated based on the worker's level of understanding.

[0165] "Provided via voice" means that generated instructions and information are communicated to workers using speech synthesis technology.

[0166] "Reinterpreting responses" is the process of re-analyzing intentions and situations based on feedback received from workers.

[0167] "Identification of emotional characteristics" is a technique that analyzes emotional indicators to estimate a worker's psychological state from their voice information.

[0168] "Adjusting instructions" refers to modifying the generated instructions and information to best suit the worker, based on identified emotional information.

[0169] This invention is implemented as a voice bot system for providing instructions tailored to the psychological state and understanding of workers during on-site work. This system aims to efficiently support workers by combining speech recognition, natural language processing, and sentiment analysis technologies.

[0170] The server processes the voice information obtained from the worker. This voice information is converted into text information using speech recognition software (e.g., speech recognition API). The text information is then analyzed using natural language processing technology (e.g., natural language API) to explore the worker's intentions and requests.

[0171] The server further uses an emotion analysis system to identify the worker's emotional characteristics from the voice information. This allows the worker's psychological state and stress level to be evaluated and used as foundational data to generate optimal instructions.

[0172] The generated instructions are provided to the worker from the terminal using speech synthesis technology (e.g., speech synthesis API). The worker receives these instructions, proceeds with the work, and sends additional voice feedback to the server as needed. The server analyzes this feedback and generates appropriate instructions again based on the newly evaluated information.

[0173] For example, if a worker says anxiously, "I don't know what to do in the next step," the server will identify that anxiety through emotion analysis and generate instructions in a reassuring tone. An example of a prompt message for this purpose would be, "Analyze the worker's voice to understand their emotions and clearly indicate the next work step in a relaxing tone."

[0174] This system enables the provision of adaptive, safe, and efficient work instructions tailored to the psychological state of workers, significantly improving communication in the workplace.

[0175] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0176] Step 1:

[0177] The server receives voice input from the user. The received voice data is sent to the server in real time. The server converts this data into text information using speech recognition technology. Specifically, it performs a speech-to-text conversion process via a speech recognition API. As a result, the voice data is output in text format.

[0178] Step 2:

[0179] The server inputs text information into its natural language processing engine. From this input text, the server performs keyword extraction and contextual analysis to interpret the user's intent. Specifically, it uses a natural language API to identify the meaning and purpose within the text data. As a result, it obtains analysis results regarding the user's intent.

[0180] Step 3:

[0181] The server analyzes the voice characteristics based on the original audio data and performs emotion analysis. The input used is the tone, speed, and intonation of the voice. The server sends this information to the emotion analysis system to evaluate the user's psychological state and stress level. Specifically, it extracts emotional characteristics using a voice analysis algorithm. As a result, emotion evaluation information is output.

[0182] Step 4:

[0183] The server generates the most appropriate additional instructions based on the analyzed user intent and sentiment evaluation. The input provided is the results of intent analysis and sentiment analysis. The server automatically creates situation-appropriate instructions using a generative AI model. Specifically, it operates the generative AI using adaptive prompt sentences to generate customized instructions. The resulting instructions are obtained.

[0184] Step 5:

[0185] The terminal transmits instructions from the server to the user via voice. The input received is the generated instruction content. The terminal uses speech synthesis technology to convert text instructions into natural-sounding speech and provide it to the user. Specifically, it uses a speech synthesis API to create speech from text data and transmits it to the user through a speaker or earphones. As a result, the user receives the instructions and understands the next action.

[0186] Step 6:

[0187] The user acts based on the instructions received and provides feedback as needed. Input includes the results of the user's actions and any questions they may have. The user's feedback is sent back to the server, where its intent and context are interpreted. The specific action involves providing feedback again via voice input, prompting the server to analyze it. As a result, the feedback is re-analyzed by the server, and additional instructions or adjustments are made as necessary.

[0188] (Application Example 2)

[0189] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0190] In on-site work communication, uniform instructions that do not consider the psychological state or emotions of workers can lead to decreased efficiency and safety. In particular, when workers are feeling anxious or stressed, traditional methods may fail to provide appropriate support, leading to misunderstandings and work errors. A system is needed to solve these problems and flexibly respond to the psychological state of workers.

[0191] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0192] In this invention, the server includes means for receiving voice input from a worker, means for converting the received voice input into text data, means for analyzing the worker's intent from the text data, means for estimating the worker's level of understanding based on the analysis results, means for generating additional instructions according to the level of understanding, means for outputting the generated instructions to the worker via voice, means for re-analyzing feedback from the worker and generating instructions again, means for analyzing the worker's emotional state, and means for adjusting instructions based on the emotional state. This enables individualized support that takes into account the worker's psychological state, and makes it possible to provide a safe and effective work environment.

[0193] "Means for receiving voice input" refers to a device or process for capturing voice signals emitted by an operator as digital data.

[0194] "Means for converting to text data" refers to a device or process that uses speech recognition technology to convert a received audio signal into text information.

[0195] "Means for analyzing intent" refers to a device or process that utilizes natural language processing technology to identify the purpose or request of a worker's utterance from text data.

[0196] "Means for estimating comprehension" are methods for evaluating how accurately a worker understands information based on the analyzed intent.

[0197] "Means for generating additional instructions" refers to a device or process for automatically generating instructions to be provided to an operator based on an estimated level of understanding or other situational factors.

[0198] "Means for outputting voice" refers to a device or process for conveying generated instructions to a worker as voice using speech synthesis technology.

[0199] A "means for reanalyzing feedback" refers to a device or process for receiving responses and reactions from workers, reanalyzing that information, and determining further actions.

[0200] "Methods for analyzing emotional states" refer to techniques that analyze the intonation and tone of voice of a worker in order to analyze their emotions and identify their psychological state.

[0201] "Means for adjusting instructions" refers to a device or process for appropriately changing the content and tone of instructions given to a worker based on their analyzed emotional state.

[0202] This invention is implemented as a voicebot system designed to improve communication in the workplace. The system utilizes speech recognition and natural language processing technologies to convert worker voice input into text and analyze intent based on that text. Specifically, a server receives voice spoken by a worker and converts the voice into text in real time via a speech recognition API. For example, the Google Cloud Speech-to-Text API is used.

[0203] The server analyzes the converted text data using a natural language processing engine to accurately understand the worker's intentions and requests. In this process, Hugging Face's Transformers library is utilized, and advanced text analysis is achieved by incorporating a generative AI model.

[0204] Furthermore, the server evaluates the intonation, speed, and tone of the voice through an emotion engine that grasps the emotional state from the worker's voice. For this purpose, open-source libraries for speech emotion recognition, such as "librosa" and "praat," are used.

[0205] The server uses an AI model to customize appropriate instructions based on the user's level of understanding and emotional state. These instructions are adjusted to take emotions into consideration and communicated to the user using speech synthesis technology (e.g., Google Text-to-Speech). For example, if a user is feeling anxious, the system will provide appropriate instructions in a calm tone to help them regain their composure.

[0206] A key feature of this system is its ability to create a feedback loop that reflects the worker's psychological state. The server uses a generative AI model to generate prompts, enabling flexible instructions tailored to individual work scenarios. By handling prompts such as, "The worker is working on a difficult task but is feeling anxious. Please generate instructions to calm them down and explain the precise steps," the server provides effective support that is relevant to the situation on site.

[0207] In this way, the system aims to provide support tailored to the individual psychological needs of workers so that they can perform their duties safely and with peace of mind.

[0208] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0209] Step 1:

[0210] The server receives voice input of the worker's voice collected by the terminal. This voice data is acquired as input and recorded in digital format.

[0211] Step 2:

[0212] The server converts the received audio data into text data using a speech recognition API (e.g., Google Cloud Speech-to-Text API). The input is audio data, and the output is text data, which is the transcription of that audio.

[0213] Step 3:

[0214] The server inputs the converted text data into a natural language processing engine (e.g., Hugging Face Transformers) to analyze the worker's intent. The input is text data, and the output is the extracted intent or command.

[0215] Step 4:

[0216] Based on the analysis results, the server generates prompt sentences using a generative AI model and estimates the operator's level of understanding. The input is the result of intent analysis, and the output is an evaluation of understanding.

[0217] Step 5:

[0218] The server uses a speech emotion recognition library (e.g., librosa) to analyze emotional states from speech data. The input is speech data, and the output is stress levels and emotional characteristics.

[0219] Step 6:

[0220] The server considers both comprehension level and emotional state, and uses a generative AI model to create optimal additional instructions. The input is the result of comprehension and emotional analysis, and the output is the customized instructions.

[0221] Step 7:

[0222] The server outputs the generated instructions as speech using speech synthesis technology (e.g., Google Text-to-Speech). The input is instructions in text format, and the output is instructions that are then transmitted to the worker in speech format.

[0223] Step 8:

[0224] The user resumes work based on the instructions received and inputs subsequent feedback into the terminal. This feedback is sent to the server for further analysis and adjustment of instructions.

[0225] Step 9:

[0226] The server analyzes user feedback and adjusts instructions as needed. The input is feedback data, and the server improves support quality by outputting revised instructions.

[0227] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0228] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0229] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0230] [Second Embodiment]

[0231] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0232] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0233] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0234] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0235] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0236] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0237] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0238] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0239] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0240] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0241] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0242] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0243] This invention is implemented as a voice bot system that streamlines communication between field workers and managers. Specifically, a server uses speech recognition technology to convert voice input from workers (users) into text data in real time. The generated text data is analyzed using natural language processing. This analysis makes it possible to accurately understand the user's intentions and work status.

[0244] Next, the server applies a deep learning algorithm to estimate the user's level of understanding based on the analysis results. During this process, it also considers speech characteristics such as tone and speed to evaluate the stress level. This determines how well the user understands the instructions and how confidently they can proceed with the task.

[0245] Based on the user's state, the server generates appropriate additional instructions. These instructions are communicated to the user using speech synthesis technology to facilitate more effective communication. The user provides feedback on the voice-output instructions. This feedback is also analyzed by the server and, if necessary, leads to the generation of new instructions.

[0246] As a concrete example, when a worker inquires about the operating procedure for a specific piece of equipment, the server generates instructions that concisely explain the procedure. If signs of stress are detected in the worker's feedback, the server adds more detailed steps and precautions to the instructions and provides them again. In this way, the system is implemented to aid user understanding and ensure transparency in communication.

[0247] The following describes the processing flow.

[0248] Step 1:

[0249] The server receives voice input from the user and records the received audio in real time. The audio data is stored in a buffer for subsequent processing.

[0250] Step 2:

[0251] The device's speech recognition engine converts received speech data into text data. This conversion process ensures accurate transcription of the speech into text.

[0252] Step 3:

[0253] The server passes the text data to a natural language processing engine, which analyzes the user's speech. Based on the analysis results, it identifies what information or instructions the user is seeking.

[0254] Step 4:

[0255] The server estimates the user's level of understanding based on the analyzed information. Here, it evaluates the level of understanding based on the user's vocal characteristics and word choices, referencing past data and learning models.

[0256] Step 5:

[0257] The server generates appropriate additional instructions, taking into account the estimated level of understanding and stress levels. These instructions are customized to fit the context of the task.

[0258] Step 6:

[0259] The device communicates the generated instructions to the user using speech synthesis technology. This voice output is provided in a clear and easy-to-understand manner.

[0260] Step 7:

[0261] The user provides feedback on the given instructions. The server converts this feedback back into text, analyzes it, and uses it to generate new instructions.

[0262] Step 8:

[0263] The server repeats the above processing cycle as needed, continuing to assist the user in reaching the desired solution.

[0264] (Example 1)

[0265] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0266] There is a need to improve the efficiency of communication between field workers and managers. In particular, a challenge is to provide an environment where workers can accurately understand the instructions they receive and proceed with their work without stress. There is also the problem of difficulty in correcting instructions and incorporating feedback in real time.

[0267] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0268] In this invention, the server includes means for receiving an audio signal, means for converting the received audio signal into text data, and means for analyzing the text data to identify the worker's intent. This allows the worker to instantly transmit their intent to the server through voice input and receive appropriate instructions and feedback according to the level of understanding.

[0269] An "audio signal" is an electrical or digital representation of information transmitted by sound.

[0270] "Text data" refers to text data that represents audio, handwritten text, and other information in a standardized format.

[0271] "Intention" refers to the purpose or wish that the worker is trying to convey through voice input, and understanding that intention leads to more efficient communication.

[0272] "Comprehension level" is an indicator that shows how correctly a worker understands and can execute the instructions and information provided to them.

[0273] A "signal" refers to additional information or instructions given to a worker, and is communicated through voice or visual means.

[0274] "Psychological stress level" refers to the degree of stress and anxiety inferred from a worker's tone of voice and speed, and serves as a criterion for determining whether to provide appropriate support in the work environment.

[0275] "Natural language processing techniques" refer to technologies and approaches for enabling computers to understand and process human language.

[0276] This invention is implemented as a voice bot system to streamline communication between field workers (users) and administrators. The system functions by having a terminal receive the voice spoken by the user at the field site, and a server process that voice.

[0277] First, users perform voice input using a dedicated mobile device or headset. The device is equipped with a high-sensitivity microphone that can eliminate background noise and transmit clear audio to the server. This allows users to perform highly accurate voice input while moving freely.

[0278] The server uses "speech recognition software" to convert speech data to text in real time. This conversion utilizes common APIs, ensuring rapid speech-to-text conversion.

[0279] The server that acquires the text data uses "natural language processing software" to analyze the user's intent. For example, if a user enters the prompt "Tell me how to replace the part right now," the server can accurately interpret that intent and provide an answer by searching for relevant information. This process employs natural language processing technology and involves detailed analysis.

[0280] Furthermore, the server uses a "deep learning model" to assess the user's psychological load from the tone and speed of their voice. This allows it to estimate the user's level of stress and anxiety and generate additional instructions tailored to their understanding. For example, if the server determines that the user is having difficulty operating the system, it will generate cues that include more detailed explanations and priority steps.

[0281] Finally, the server uses "voice synthesis software" to provide the generated instructions as acoustic signals and transmit them to the user. The user continues the work based on the received voice and provides feedback if necessary. This feedback is also processed by the server, leading to the provision of new information and the improvement of instructions.

[0282] As a specific example, there is a series of processes where, after an operator at the site reports a machine abnormality, the server asks questions to identify the cause of the abnormality. Examples of prompt sentences used in this case include "The vibration of the machine is abnormal. Please diagnose the cause." With these, the user can efficiently solve problems while receiving support.

[0283] The flow of the specific process in Example 1 will be described using FIG. 11.

[0284] Step 1:

[0285] The user performs voice input through a dedicated terminal. What is input is the user's question or instruction, and the microphone of the terminal receives this as an acoustic signal. This acoustic signal is converted into a digital format inside the terminal. The output is digitized voice data.

[0286] Step 2:

[0287] The terminal transmits the digitized voice data to the server. The server receives this data and uses "voice recognition software" to convert it into text data. When converting the voice data into text format, phoneme analysis is performed and output as a sentence. The output is text data.

[0288] Step 3:

[0289] The server analyzes the text data with "natural language processing software". In order to extract the intention from the input text data, context analysis and keyword identification are performed. Data with the user's intention and purpose clarified is output.

[0290] Step 4:

[0291] Based on the analyzed intent, the server uses a "deep learning model" to evaluate the user's understanding. This process also takes speech properties into consideration. Data regarding the estimated understanding and stress level is output.

[0292] Step 5:

[0293] The server generates additional instructions based on the user's level of understanding. Information processing is performed to format the generated instructions into an easily understandable form. The output is a written instruction provided to the user.

[0294] Step 6:

[0295] The server uses "speech synthesis software" to convert the generated instructions into audio signals. At this stage, the process generates natural-sounding, easily recognizable speech from the text data. The output is audio data transmitted to the user through the terminal.

[0296] Step 7:

[0297] The user performs tasks based on the voice instructions received and provides voice feedback as needed. This feedback is then captured again as digital voice data on the device. The output is the voice data of the feedback.

[0298] Step 8:

[0299] The server receives the audio data of the feedback, analyzes it again, and generates new instructions if necessary. This allows the system to continuously improve communication. The output consists of new instructions and feedback analysis results.

[0300] (Application Example 1)

[0301] Next, Application Example 1 will be described. In the following description, the data processing device 12 is referred to as a "server", and the smart glasses 214 are referred to as a "terminal".

[0302] In complex on-site work, it is required to smooth the communication between the worker and the system. However, in a conventional system based on voice input, it is difficult to appropriately grasp the worker's understanding level and emotional state. For this reason, there is a situation where misunderstandings are not prevented and sufficient support for the worker to work with confidence and efficiently is not provided.

[0303] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0304] In this invention, the server includes a medium for acquiring the voice data of the worker, a means for converting the voice data into character information, and a means for recognizing the intention of the worker from the character information. Thereby, it becomes possible to accurately evaluate the worker's understanding level and emotional state and provide appropriate additional information and instructions in a timely manner.

[0305] The "worker" is a person who operates equipment and facilities in a factory or on-site to carry out work.

[0306] The "voice data" is data of an acoustic signal acquired by the worker's vocalization.

[0307] The "medium" is a physical or virtual device or means for acquiring or providing information.

[0308] The "character information" is a text representation of a language converted based on voice data.

[0309] The "means for recognizing" is a method or technology for analyzing the acquired data and understanding its meaning and purpose. =

[0310] The "understanding level" is an index indicating the degree to which the worker accurately grasps the given instructions and information.

[0311] "Means of evaluation" refer to methods and techniques for analyzing data and situations and making judgments about their quality and characteristics.

[0312] "Additional information" refers to detailed information provided to assist workers in understanding the work and to facilitate the smooth progress of the task.

[0313] A "device" is a mechanical or electronic system built to perform a specific function or task.

[0314] A "response" is a reaction or reply that an operator gives to instructions or information from a system.

[0315] "Analysis" is the process of examining data in detail to clarify its structure and meaning.

[0316] "Stress level" refers to the degree of burden or tension that workers feel towards their work and environment.

[0317] "To analyze with high accuracy" means to analyze data and information with great precision and detail.

[0318] This invention will now be described in terms of embodiments for carrying it out. The main element of the invention is an interactive system for smooth information exchange between a worker and a system. This system is realized by acquiring voice data via smart glasses or a headset worn by a worker on site and processing that voice on a server.

[0319] The server converts audio data acquired from the smart glasses' microphone into text in real time using the Google Speech-to-Text API. Then, it analyzes this text using NLTK (Natural Language Toolkit) to recognize the worker's intent. Based on this analysis, a deep learning model built with PyTorch is used to evaluate the worker's understanding and stress level. Based on this evaluation, the server generates additional information tailored to the worker and communicates it to them verbally using Google Text-to-Speech.

[0320] This process continuously provides optimal instructions by constantly acquiring new responses from workers, analyzing them again, and re-evaluating their level of understanding and stress. As a result, workers receive specific instructions tailored to their situation, enabling them to work efficiently and with peace of mind.

[0321] As a concrete example, consider a scenario where an operator troubleshoots a metalworking machine. The operator asks a question by voice, and the question is immediately analyzed to provide specific machine operating procedures and precautions. If the operator shows any signs of uncertainty, more detailed information will be provided later.

[0322] An example of a prompt message is, "Please provide detailed instructions for troubleshooting the metalworking machine." This prompt allows the system to quickly and accurately convey the necessary information, enabling the operator to efficiently resolve the problem accordingly.

[0323] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0324] Step 1:

[0325] The terminal acquires the worker's voice as audio data from the microphone in the smart glasses. The input is the worker's voice, and this audio data serves as the trigger to start the program.

[0326] Step 2:

[0327] The server converts the acquired audio data into text using the Google Speech-to-Text API. The input is audio data, and the output is text data. This process transforms the audio information into parseable text.

[0328] Step 3:

[0329] The server analyzes the converted character information using NLTK to recognize the worker's intent. The input is character information, and the output is data indicating the worker's intent. This analysis process clarifies the information and instructions the worker is seeking.

[0330] Step 4:

[0331] The server evaluates the worker's level of understanding and stress through a deep learning model using PyTorch. Inputs are intent recognition results and emotional data obtained from the worker's voice, while outputs are understanding and stress indicators. This allows the worker's state to be quantified, enabling appropriate responses based on that situation.

[0332] Step 5:

[0333] Based on the evaluation results, the server utilizes a generative AI model to generate optimal additional information and instructions for the worker. The inputs are comprehension levels and stress indicators, while the output is the content of the instructions for the worker. Appropriate content is provided according to the worker's condition.

[0334] Step 6:

[0335] The server uses Google Text-to-Speech to synthesize the generated instructions into speech and sends them to the terminal. The input is the generated instructions, and the output is the audio data. The terminal then transmits this audio to the worker.

[0336] Step 7:

[0337] The user provides feedback to the system again, and this voice feedback triggers a return to step 1 of the next cycle. The input is the worker's new voice feedback, which triggers the process of being analyzed again.

[0338] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0339] This invention is implemented as a voice bot system for improving communication in field work. By incorporating an emotion engine, this system not only provides instructions but also generates adaptive instructions that respond to the user's emotions and stress levels.

[0340] Specifically, the server receives voice input from the user and converts it into text data in real time using speech recognition. The transcribed text is then analyzed by a natural language processing engine to accurately understand the user's intentions and needs.

[0341] Furthermore, the server uses an emotion engine to identify the emotional characteristics of the speech. In this process, various speech features such as tone, speed, and intonation are considered to determine the user's psychological state. Meanwhile, the level of understanding is evaluated based on the analysis results in response to the user's questions and requests.

[0342] By combining emotional information evaluated based on the emotion engine with an assessment of comprehension, the system generates the most appropriate additional instructions. These instructions are customized to take into account the user's psychological state and work context. The terminal communicates these instructions to the user using speech synthesis and subsequently analyzes the received feedback.

[0343] For example, if a user makes a mistake in executing a work procedure, the server detects the user's anxiety and tension through its emotion engine. Based on this emotional information, it generates instructions in a calm tone to alleviate the anxiety and reinstructs the user to follow the correct procedure in a reassuring manner. This system is highly responsive to the user's psychology and state, enabling it to support a safe and secure work environment.

[0344] The following describes the processing flow.

[0345] Step 1:

[0346] The user provides voice input to the device. The device receives this voice input as digital audio data.

[0347] Step 2:

[0348] The server provides the received audio data to the speech recognition engine in real time, where it is converted into text data. The speech recognition engine accurately processes the conversion of speech into text.

[0349] Step 3:

[0350] The server passes the converted text data to a natural language processing engine, which analyzes the user's intent. This analysis helps understand what the user wants.

[0351] Step 4:

[0352] The server uses an emotion engine to analyze the audio data. The emotion engine evaluates the user's emotional state based on factors such as tone, speed, and rhythm of the voice.

[0353] Step 5:

[0354] The server combines the user's intent analysis results and emotion evaluation results to estimate their level of understanding and psychological state. Based on this estimation, it determines what kind of instructions are needed.

[0355] Step 6:

[0356] The server generates additional instructions tailored to the user based on their estimated level of understanding and emotional state. These instructions are designed to maximize the user's safety and efficiency.

[0357] Step 7:

[0358] The terminal uses speech synthesis to communicate additional instructions sent from the server to the user. The user understands the next action through the voice instructions.

[0359] Step 8:

[0360] The user provides feedback on the additional instructions. This feedback is converted back into text on the server and used to analyze the user's understanding and the need for continued support.

[0361] Step 9:

[0362] If necessary, the server optimizes the process, provides more specific instructions, and delivers them again to the user via the terminal. This iterative process continuously supports the user and reduces misunderstandings.

[0363] (Example 2)

[0364] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0365] There is a lack of instructions that take into account the psychological state of workers in on-site work, and it is necessary to reduce inefficiencies and errors caused by misunderstandings and stress. Furthermore, it is necessary to interpret workers' intentions with high accuracy and provide appropriate support accordingly.

[0366] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0367] In this invention, the server includes means for acquiring voice information of the worker, means for identifying emotional characteristics from the voice information, and means for adjusting the content of instructions based on the identified emotional information. This makes it possible to provide instructions that are adaptive and reassuring to the worker, in accordance with their psychological state.

[0368] "Audio information" refers to all audio data emitted by workers, and is material that is analyzed by speech recognition.

[0369] "Textual information" refers to text data converted from audio information and is used to interpret the worker's intent.

[0370] "Interpreting intent" is the process of identifying the purpose or requirements from what the worker says.

[0371] "Understanding level" is an indicator that shows how well a worker understands the information and instructions provided.

[0372] "Additional information" refers to supplementary instructions or explanations generated based on the worker's level of understanding.

[0373] "Provided via voice" means that generated instructions and information are communicated to workers using speech synthesis technology.

[0374] "Reinterpreting responses" is the process of re-analyzing intentions and situations based on feedback received from workers.

[0375] "Identification of emotional characteristics" is a technique that analyzes emotional indicators to estimate a worker's psychological state from their voice information.

[0376] "Adjusting instructions" refers to modifying the generated instructions and information to best suit the worker, based on identified emotional information.

[0377] This invention is implemented as a voice bot system for providing instructions tailored to the psychological state and understanding of workers during on-site work. This system aims to efficiently support workers by combining speech recognition, natural language processing, and sentiment analysis technologies.

[0378] The server processes the voice information obtained from the worker. This voice information is converted into text information using speech recognition software (e.g., speech recognition API). The text information is then analyzed using natural language processing technology (e.g., natural language API) to explore the worker's intentions and requests.

[0379] The server further uses an emotion analysis system to identify the worker's emotional characteristics from the voice information. This allows the worker's psychological state and stress level to be evaluated and used as foundational data to generate optimal instructions.

[0380] The generated instructions are provided to the worker from the terminal using speech synthesis technology (e.g., speech synthesis API). The worker receives these instructions, proceeds with the work, and sends additional voice feedback to the server as needed. The server analyzes this feedback and generates appropriate instructions again based on the newly evaluated information.

[0381] For example, if a worker says anxiously, "I don't know what to do in the next step," the server will identify that anxiety through emotion analysis and generate instructions in a reassuring tone. An example of a prompt message for this purpose would be, "Analyze the worker's voice to understand their emotions and clearly indicate the next work step in a relaxing tone."

[0382] This system enables the provision of adaptive, safe, and efficient work instructions tailored to the psychological state of workers, significantly improving communication in the workplace.

[0383] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0384] Step 1:

[0385] The server receives voice input from the user. The received voice data is sent to the server in real time. The server converts this data into text information using speech recognition technology. Specifically, it performs a speech-to-text conversion process via a speech recognition API. As a result, the voice data is output in text format.

[0386] Step 2:

[0387] The server inputs text information into its natural language processing engine. From this input text, the server performs keyword extraction and contextual analysis to interpret the user's intent. Specifically, it uses a natural language API to identify the meaning and purpose within the text data. As a result, it obtains analysis results regarding the user's intent.

[0388] Step 3:

[0389] The server analyzes the voice characteristics based on the original audio data and performs emotion analysis. The input used is the tone, speed, and intonation of the voice. The server sends this information to the emotion analysis system to evaluate the user's psychological state and stress level. Specifically, it extracts emotional characteristics using a voice analysis algorithm. As a result, emotion evaluation information is output.

[0390] Step 4:

[0391] The server generates the most appropriate additional instructions based on the analyzed user intent and sentiment evaluation. The input provided is the results of intent analysis and sentiment analysis. The server automatically creates situation-appropriate instructions using a generative AI model. Specifically, it operates the generative AI using adaptive prompt sentences to generate customized instructions. The resulting instructions are obtained.

[0392] Step 5:

[0393] The terminal transmits instructions from the server to the user via voice. The input received is the generated instruction content. The terminal uses speech synthesis technology to convert text instructions into natural-sounding speech and provide it to the user. Specifically, it uses a speech synthesis API to create speech from text data and transmits it to the user through a speaker or earphones. As a result, the user receives the instructions and understands the next action.

[0394] Step 6:

[0395] The user acts based on the instructions received and provides feedback as needed. Input includes the results of the user's actions and any questions they may have. The user's feedback is sent back to the server, where its intent and context are interpreted. The specific action involves providing feedback again via voice input, prompting the server to analyze it. As a result, the feedback is re-analyzed by the server, and additional instructions or adjustments are made as necessary.

[0396] (Application Example 2)

[0397] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0398] In on-site work communication, uniform instructions that do not consider the psychological state or emotions of workers can lead to decreased efficiency and safety. In particular, when workers are feeling anxious or stressed, traditional methods may fail to provide appropriate support, leading to misunderstandings and work errors. A system is needed to solve these problems and flexibly respond to the psychological state of workers.

[0399] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0400] In this invention, the server includes means for receiving voice input from a worker, means for converting the received voice input into text data, means for analyzing the worker's intent from the text data, means for estimating the worker's level of understanding based on the analysis results, means for generating additional instructions according to the level of understanding, means for outputting the generated instructions to the worker via voice, means for re-analyzing feedback from the worker and generating instructions again, means for analyzing the worker's emotional state, and means for adjusting instructions based on the emotional state. This enables individualized support that takes into account the worker's psychological state, and makes it possible to provide a safe and effective work environment.

[0401] "Means for receiving voice input" refers to a device or process for capturing voice signals emitted by an operator as digital data.

[0402] "Means for converting to text data" refers to a device or process that uses speech recognition technology to convert a received audio signal into text information.

[0403] "Means for analyzing intent" refers to a device or process that utilizes natural language processing technology to identify the purpose or request of a worker's utterance from text data.

[0404] "Means for estimating comprehension" are methods for evaluating how accurately a worker understands information based on the analyzed intent.

[0405] "Means for generating additional instructions" refers to a device or process for automatically generating instructions to be provided to an operator based on an estimated level of understanding or other situational factors.

[0406] "Means for outputting voice" refers to a device or process for conveying generated instructions to a worker as voice using speech synthesis technology.

[0407] A "means for reanalyzing feedback" refers to a device or process for receiving responses and reactions from workers, reanalyzing that information, and determining further actions.

[0408] "Methods for analyzing emotional states" refer to techniques that analyze the intonation and tone of voice of a worker in order to analyze their emotions and identify their psychological state.

[0409] "Means for adjusting instructions" refers to a device or process for appropriately changing the content and tone of instructions given to a worker based on their analyzed emotional state.

[0410] This invention is implemented as a voicebot system designed to improve communication in the workplace. The system utilizes speech recognition and natural language processing technologies to convert worker voice input into text and analyze intent based on that text. Specifically, a server receives voice spoken by a worker and converts the voice into text in real time via a speech recognition API. For example, the Google Cloud Speech-to-Text API is used.

[0411] The server analyzes the converted text data using a natural language processing engine to accurately understand the worker's intentions and requests. In this process, Hugging Face's Transformers library is utilized, and advanced text analysis is achieved by incorporating a generative AI model.

[0412] Furthermore, the server evaluates the intonation, speed, and tone of the voice through an emotion engine that grasps the emotional state from the worker's voice. For this purpose, open-source libraries for speech emotion recognition, such as "librosa" and "praat," are used.

[0413] The server uses an AI model to customize appropriate instructions based on the user's level of understanding and emotional state. These instructions are adjusted to take emotions into consideration and communicated to the user using speech synthesis technology (e.g., Google Text-to-Speech). For example, if a user is feeling anxious, the system will provide appropriate instructions in a calm tone to help them regain their composure.

[0414] A key feature of this system is its ability to create a feedback loop that reflects the worker's psychological state. The server uses a generative AI model to generate prompts, enabling flexible instructions tailored to individual work scenarios. By handling prompts such as, "The worker is working on a difficult task but is feeling anxious. Please generate instructions to calm them down and explain the precise steps," the server provides effective support that is relevant to the situation on site.

[0415] In this way, the system aims to provide support tailored to the individual psychological needs of workers so that they can perform their duties safely and with peace of mind.

[0416] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0417] Step 1:

[0418] The server receives voice input of the worker's voice collected by the terminal. This voice data is acquired as input and recorded in digital format.

[0419] Step 2:

[0420] The server converts the received audio data into text data using a speech recognition API (e.g., Google Cloud Speech-to-Text API). The input is audio data, and the output is text data, which is the transcription of that audio.

[0421] Step 3:

[0422] The server inputs the converted text data into a natural language processing engine (e.g., Hugging Face Transformers) to analyze the worker's intent. The input is text data, and the output is the extracted intent or command.

[0423] Step 4:

[0424] Based on the analysis results, the server generates prompt sentences using a generative AI model and estimates the operator's level of understanding. The input is the result of intent analysis, and the output is an evaluation of understanding.

[0425] Step 5:

[0426] The server uses a speech emotion recognition library (e.g., librosa) to analyze emotional states from speech data. The input is speech data, and the output is stress levels and emotional characteristics.

[0427] Step 6:

[0428] The server considers both comprehension level and emotional state, and uses a generative AI model to create optimal additional instructions. The input is the result of comprehension and emotional analysis, and the output is the customized instructions.

[0429] Step 7:

[0430] The server outputs the generated instructions as speech using speech synthesis technology (e.g., Google Text-to-Speech). The input is instructions in text format, and the output is instructions that are then transmitted to the worker in speech format.

[0431] Step 8:

[0432] The user resumes work based on the instructions received and inputs subsequent feedback into the terminal. This feedback is sent to the server for further analysis and adjustment of instructions.

[0433] Step 9:

[0434] The server analyzes user feedback and adjusts instructions as needed. The input is feedback data, and the server improves support quality by outputting revised instructions.

[0435] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0436] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0437] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0438] [Third Embodiment]

[0439] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0440] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0441] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0442] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0443] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0444] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0445] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0446] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0447] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0448] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0449] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0450] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0451] This invention is implemented as a voice bot system that streamlines communication between field workers and managers. Specifically, a server uses speech recognition technology to convert voice input from workers (users) into text data in real time. The generated text data is analyzed using natural language processing. This analysis makes it possible to accurately understand the user's intentions and work status.

[0452] Next, the server applies a deep learning algorithm to estimate the user's level of understanding based on the analysis results. During this process, it also considers speech characteristics such as tone and speed to evaluate the stress level. This determines how well the user understands the instructions and how confidently they can proceed with the task.

[0453] Based on the user's state, the server generates appropriate additional instructions. These instructions are communicated to the user using speech synthesis technology to facilitate more effective communication. The user provides feedback on the voice-output instructions. This feedback is also analyzed by the server and, if necessary, leads to the generation of new instructions.

[0454] As a concrete example, when a worker inquires about the operating procedure for a specific piece of equipment, the server generates instructions that concisely explain the procedure. If signs of stress are detected in the worker's feedback, the server adds more detailed steps and precautions to the instructions and provides them again. In this way, the system is implemented to aid user understanding and ensure transparency in communication.

[0455] The following describes the processing flow.

[0456] Step 1:

[0457] The server receives voice input from the user and records the received audio in real time. The audio data is stored in a buffer for subsequent processing.

[0458] Step 2:

[0459] The device's speech recognition engine converts received speech data into text data. This conversion process ensures accurate transcription of the speech into text.

[0460] Step 3:

[0461] The server passes the text data to a natural language processing engine, which analyzes the user's speech. Based on the analysis results, it identifies what information or instructions the user is seeking.

[0462] Step 4:

[0463] The server estimates the user's level of understanding based on the analyzed information. Here, it evaluates the level of understanding based on the user's vocal characteristics and word choices, referencing past data and learning models.

[0464] Step 5:

[0465] The server generates appropriate additional instructions, taking into account the estimated level of understanding and stress levels. These instructions are customized to fit the context of the task.

[0466] Step 6:

[0467] The device communicates the generated instructions to the user using speech synthesis technology. This voice output is provided in a clear and easy-to-understand manner.

[0468] Step 7:

[0469] The user provides feedback on the given instructions. The server converts this feedback back into text, analyzes it, and uses it to generate new instructions.

[0470] Step 8:

[0471] The server repeats the above processing cycle as needed, continuing to assist the user in reaching the desired solution.

[0472] (Example 1)

[0473] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0474] There is a need to improve the efficiency of communication between field workers and managers. In particular, a challenge is to provide an environment where workers can accurately understand the instructions they receive and proceed with their work without stress. There is also the problem of difficulty in correcting instructions and incorporating feedback in real time.

[0475] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0476] In this invention, the server includes means for receiving an audio signal, means for converting the received audio signal into text data, and means for analyzing the text data to identify the worker's intent. This allows the worker to instantly transmit their intent to the server through voice input and receive appropriate instructions and feedback according to the level of understanding.

[0477] An "audio signal" is an electrical or digital representation of information transmitted by sound.

[0478] "Text data" refers to text data that represents audio, handwritten text, and other information in a standardized format.

[0479] "Intention" refers to the purpose or wish that the worker is trying to convey through voice input, and understanding that intention leads to more efficient communication.

[0480] "Comprehension level" is an indicator that shows how correctly a worker understands and can execute the instructions and information provided to them.

[0481] A "signal" refers to additional information or instructions given to a worker, and is communicated through voice or visual means.

[0482] "Psychological stress level" refers to the degree of stress and anxiety inferred from a worker's tone of voice and speed, and serves as a criterion for determining whether to provide appropriate support in the work environment.

[0483] "Natural language processing techniques" refer to technologies and approaches for enabling computers to understand and process human language.

[0484] This invention is implemented as a voice bot system to streamline communication between field workers (users) and administrators. The system functions by having a terminal receive the voice spoken by the user at the field site, and a server process that voice.

[0485] First, users perform voice input using a dedicated mobile device or headset. The device is equipped with a high-sensitivity microphone that can eliminate background noise and transmit clear audio to the server. This allows users to perform highly accurate voice input while moving freely.

[0486] The server uses "speech recognition software" to convert speech data to text in real time. This conversion utilizes common APIs, ensuring rapid speech-to-text conversion.

[0487] The server that acquires the text data uses "natural language processing software" to analyze the user's intent. For example, if a user enters the prompt "Tell me how to replace the part right now," the server can accurately interpret that intent and provide an answer by searching for relevant information. This process employs natural language processing technology and involves detailed analysis.

[0488] Furthermore, the server uses a "deep learning model" to assess the user's psychological load from the tone and speed of their voice. This allows it to estimate the user's level of stress and anxiety and generate additional instructions tailored to their understanding. For example, if the server determines that the user is having difficulty operating the system, it will generate cues that include more detailed explanations and priority steps.

[0489] Ultimately, the server uses "speech synthesis software" to provide the generated instructions as audio signals, which are then transmitted to the user. The user continues working based on the received audio and provides feedback as needed. This feedback is also processed by the server, leading to the provision of new information and improvements to the instructions.

[0490] A concrete example is a process where, after a worker on-site reports a machine malfunction, the server asks questions to identify the cause of the malfunction. Examples of prompts used in this process include, "The machine is vibrating abnormally; please diagnose the cause." This allows users to efficiently resolve problems while receiving support.

[0491] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0492] Step 1:

[0493] The user provides voice input through a dedicated terminal. The input consists of the user's questions and instructions, which are received as audio signals by the terminal's microphone. These audio signals are then converted into digital format within the terminal. The output is digitized audio data.

[0494] Step 2:

[0495] The terminal sends digitized audio data to the server. The server receives this data and uses "speech recognition software" to convert it into text data. During the conversion of audio data to text format, phoneme analysis is performed, and the output is generated as text. The output is text data.

[0496] Step 3:

[0497] The server analyzes the text data using "natural language processing software." Contextual analysis and keyword identification are performed to extract intent from the input text data. Data with clearly defined user intent and purpose is then output.

[0498] Step 4:

[0499] Based on the analyzed intent, the server uses a "deep learning model" to evaluate the user's understanding. This process also takes speech properties into consideration. Data regarding the estimated understanding and stress level is output.

[0500] Step 5:

[0501] The server generates additional instructions based on the user's level of understanding. Information processing is performed to format the generated instructions into an easily understandable form. The output is a written instruction provided to the user.

[0502] Step 6:

[0503] The server uses "speech synthesis software" to convert the generated instructions into audio signals. At this stage, the process generates natural-sounding, easily recognizable speech from the text data. The output is audio data transmitted to the user through the terminal.

[0504] Step 7:

[0505] The user performs tasks based on the voice instructions received and provides voice feedback as needed. This feedback is then captured again as digital voice data on the device. The output is the voice data of the feedback.

[0506] Step 8:

[0507] The server receives the audio data of the feedback, analyzes it again, and generates new instructions if necessary. This allows the system to continuously improve communication. The output consists of new instructions and feedback analysis results.

[0508] (Application Example 1)

[0509] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0510] In complex on-site tasks, smooth communication between workers and systems is essential. However, conventional systems based on voice input have difficulty accurately grasping the worker's level of understanding and emotional state. As a result, they are not adequately supporting workers in preventing misunderstandings and ensuring they can work efficiently and with confidence.

[0511] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0512] In this invention, the server includes a medium for acquiring voice data from a worker, means for converting the voice data into text information, and means for recognizing the worker's intentions from the text information. This makes it possible to accurately evaluate the worker's level of understanding and emotional state, and to provide appropriate additional information and instructions in a timely manner.

[0513] A "worker" is a person who operates equipment and facilities and performs tasks in a factory or on-site.

[0514] "Voice data" refers to data of acoustic signals acquired through the voice of a worker.

[0515] "Medium" refers to a physical or virtual device or means used to acquire or provide information.

[0516] "Textual information" refers to the textual representation of language converted from audio data.

[0517] "Means of recognition" refer to methods and technologies for analyzing acquired data and understanding its meaning and purpose.

[0518] "Comprehension level" is an indicator that shows how accurately a worker understands the instructions and information they are given.

[0519] "Means of evaluation" refer to methods and techniques for analyzing data and situations and making judgments about their quality and characteristics.

[0520] "Additional information" refers to detailed information provided to assist workers in understanding the work and to facilitate the smooth progress of the task.

[0521] A "device" is a mechanical or electronic system built to perform a specific function or task.

[0522] A "response" is a reaction or reply that an operator gives to instructions or information from a system.

[0523] "Analysis" is the process of examining data in detail to clarify its structure and meaning.

[0524] "Stress level" refers to the degree of burden or tension that workers feel towards their work and environment.

[0525] "To analyze with high accuracy" means to analyze data and information with great precision and detail.

[0526] This invention will now be described in terms of embodiments for carrying it out. The main element of the invention is an interactive system for smooth information exchange between a worker and a system. This system is realized by acquiring voice data via smart glasses or a headset worn by a worker on site and processing that voice on a server.

[0527] The server converts audio data acquired from the smart glasses' microphone into text in real time using the Google Speech-to-Text API. Then, it analyzes this text using NLTK (Natural Language Toolkit) to recognize the worker's intent. Based on this analysis, a deep learning model built with PyTorch is used to evaluate the worker's understanding and stress level. Based on this evaluation, the server generates additional information tailored to the worker and communicates it to them verbally using Google Text-to-Speech.

[0528] This process continuously provides optimal instructions by constantly acquiring new responses from workers, analyzing them again, and re-evaluating their level of understanding and stress. As a result, workers receive specific instructions tailored to their situation, enabling them to work efficiently and with peace of mind.

[0529] As a concrete example, consider a scenario where an operator troubleshoots a metalworking machine. The operator asks a question by voice, and the question is immediately analyzed to provide specific machine operating procedures and precautions. If the operator shows any signs of uncertainty, more detailed information will be provided later.

[0530] An example of a prompt message is, "Please provide detailed instructions for troubleshooting the metalworking machine." This prompt allows the system to quickly and accurately convey the necessary information, enabling the operator to efficiently resolve the problem accordingly.

[0531] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0532] Step 1:

[0533] The terminal acquires the worker's voice as audio data from the microphone in the smart glasses. The input is the worker's voice, and this audio data serves as the trigger to start the program.

[0534] Step 2:

[0535] The server converts the acquired audio data into text using the Google Speech-to-Text API. The input is audio data, and the output is text data. This process transforms the audio information into parseable text.

[0536] Step 3:

[0537] The server analyzes the converted character information using NLTK to recognize the worker's intent. The input is character information, and the output is data indicating the worker's intent. This analysis process clarifies the information and instructions the worker is seeking.

[0538] Step 4:

[0539] The server evaluates the worker's level of understanding and stress through a deep learning model using PyTorch. Inputs are intent recognition results and emotional data obtained from the worker's voice, while outputs are understanding and stress indicators. This allows the worker's state to be quantified, enabling appropriate responses based on that situation.

[0540] Step 5:

[0541] Based on the evaluation results, the server utilizes a generative AI model to generate optimal additional information and instructions for the worker. The inputs are comprehension levels and stress indicators, while the output is the content of the instructions for the worker. Appropriate content is provided according to the worker's condition.

[0542] Step 6:

[0543] The server uses Google Text-to-Speech to synthesize the generated instructions into speech and sends them to the terminal. The input is the generated instructions, and the output is the audio data. The terminal then transmits this audio to the worker.

[0544] Step 7:

[0545] The user provides feedback to the system again, and this voice feedback triggers a return to step 1 of the next cycle. The input is the worker's new voice feedback, which triggers the process of being analyzed again.

[0546] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0547] This invention is implemented as a voice bot system for improving communication in field work. By incorporating an emotion engine, this system not only provides instructions but also generates adaptive instructions that respond to the user's emotions and stress levels.

[0548] Specifically, the server receives voice input from the user and converts it into text data in real time using speech recognition. The transcribed text is then analyzed by a natural language processing engine to accurately understand the user's intentions and needs.

[0549] Furthermore, the server uses an emotion engine to identify the emotional characteristics of the speech. In this process, various speech features such as tone, speed, and intonation are considered to determine the user's psychological state. Meanwhile, the level of understanding is evaluated based on the analysis results in response to the user's questions and requests.

[0550] By combining emotional information evaluated based on the emotion engine with an assessment of comprehension, the system generates the most appropriate additional instructions. These instructions are customized to take into account the user's psychological state and work context. The terminal communicates these instructions to the user using speech synthesis and subsequently analyzes the received feedback.

[0551] For example, if a user makes a mistake in executing a work procedure, the server detects the user's anxiety and tension through its emotion engine. Based on this emotional information, it generates instructions in a calm tone to alleviate the anxiety and reinstructs the user to follow the correct procedure in a reassuring manner. This system is highly responsive to the user's psychology and state, enabling it to support a safe and secure work environment.

[0552] The following describes the processing flow.

[0553] Step 1:

[0554] The user provides voice input to the device. The device receives this voice input as digital audio data.

[0555] Step 2:

[0556] The server provides the received audio data to the speech recognition engine in real time, where it is converted into text data. The speech recognition engine accurately processes the conversion of speech into text.

[0557] Step 3:

[0558] The server passes the converted text data to a natural language processing engine, which analyzes the user's intent. This analysis helps understand what the user wants.

[0559] Step 4:

[0560] The server uses an emotion engine to analyze the audio data. The emotion engine evaluates the user's emotional state based on factors such as tone, speed, and rhythm of the voice.

[0561] Step 5:

[0562] The server combines the user's intent analysis results and emotion evaluation results to estimate their level of understanding and psychological state. Based on this estimation, it determines what kind of instructions are needed.

[0563] Step 6:

[0564] The server generates additional instructions tailored to the user based on their estimated level of understanding and emotional state. These instructions are designed to maximize the user's safety and efficiency.

[0565] Step 7:

[0566] The terminal uses speech synthesis to communicate additional instructions sent from the server to the user. The user understands the next action through the voice instructions.

[0567] Step 8:

[0568] The user provides feedback on the additional instructions. This feedback is converted back into text on the server and used to analyze the user's understanding and the need for continued support.

[0569] Step 9:

[0570] If necessary, the server optimizes the process, provides more specific instructions, and delivers them again to the user via the terminal. This iterative process continuously supports the user and reduces misunderstandings.

[0571] (Example 2)

[0572] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0573] There is a lack of instructions that take into account the psychological state of workers in on-site work, and it is necessary to reduce inefficiencies and errors caused by misunderstandings and stress. Furthermore, it is necessary to interpret workers' intentions with high accuracy and provide appropriate support accordingly.

[0574] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0575] In this invention, the server includes means for acquiring voice information of the worker, means for identifying emotional characteristics from the voice information, and means for adjusting the content of instructions based on the identified emotional information. This makes it possible to provide instructions that are adaptive and reassuring to the worker, in accordance with their psychological state.

[0576] "Audio information" refers to all audio data emitted by workers, and is material that is analyzed by speech recognition.

[0577] "Textual information" refers to text data converted from audio information and is used to interpret the worker's intent.

[0578] "Interpreting intent" is the process of identifying the purpose or requirements from what the worker says.

[0579] "Understanding level" is an indicator that shows how well a worker understands the information and instructions provided.

[0580] "Additional information" refers to supplementary instructions or explanations generated based on the worker's level of understanding.

[0581] "Provided via voice" means that generated instructions and information are communicated to workers using speech synthesis technology.

[0582] "Reinterpreting responses" is the process of re-analyzing intentions and situations based on feedback received from workers.

[0583] "Identification of emotional characteristics" is a technique that analyzes emotional indicators to estimate a worker's psychological state from their voice information.

[0584] "Adjusting instructions" refers to modifying the generated instructions and information to best suit the worker, based on identified emotional information.

[0585] This invention is implemented as a voice bot system for providing instructions tailored to the psychological state and understanding of workers during on-site work. This system aims to efficiently support workers by combining speech recognition, natural language processing, and sentiment analysis technologies.

[0586] The server processes the voice information obtained from the worker. This voice information is converted into text information using speech recognition software (e.g., speech recognition API). The text information is then analyzed using natural language processing technology (e.g., natural language API) to explore the worker's intentions and requests.

[0587] The server further uses an emotion analysis system to identify the worker's emotional characteristics from the voice information. This allows the worker's psychological state and stress level to be evaluated and used as foundational data to generate optimal instructions.

[0588] The generated instructions are provided to the worker from the terminal using speech synthesis technology (e.g., speech synthesis API). The worker receives these instructions, proceeds with the work, and sends additional voice feedback to the server as needed. The server analyzes this feedback and generates appropriate instructions again based on the newly evaluated information.

[0589] For example, if a worker says anxiously, "I don't know what to do in the next step," the server will identify that anxiety through emotion analysis and generate instructions in a reassuring tone. An example of a prompt message for this purpose would be, "Analyze the worker's voice to understand their emotions and clearly indicate the next work step in a relaxing tone."

[0590] This system enables the provision of adaptive, safe, and efficient work instructions tailored to the psychological state of workers, significantly improving communication in the workplace.

[0591] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0592] Step 1:

[0593] The server receives voice input from the user. The received voice data is sent to the server in real time. The server converts this data into text information using speech recognition technology. Specifically, it performs a speech-to-text conversion process via a speech recognition API. As a result, the voice data is output in text format.

[0594] Step 2:

[0595] The server inputs text information into its natural language processing engine. From this input text, the server performs keyword extraction and contextual analysis to interpret the user's intent. Specifically, it uses a natural language API to identify the meaning and purpose within the text data. As a result, it obtains analysis results regarding the user's intent.

[0596] Step 3:

[0597] The server analyzes the voice characteristics based on the original audio data and performs emotion analysis. The input used is the tone, speed, and intonation of the voice. The server sends this information to the emotion analysis system to evaluate the user's psychological state and stress level. Specifically, it extracts emotional characteristics using a voice analysis algorithm. As a result, emotion evaluation information is output.

[0598] Step 4:

[0599] The server generates the most appropriate additional instructions based on the analyzed user intent and sentiment evaluation. The input provided is the results of intent analysis and sentiment analysis. The server automatically creates situation-appropriate instructions using a generative AI model. Specifically, it operates the generative AI using adaptive prompt sentences to generate customized instructions. The resulting instructions are obtained.

[0600] Step 5:

[0601] The terminal transmits instructions from the server to the user via voice. The input received is the generated instruction content. The terminal uses speech synthesis technology to convert text instructions into natural-sounding speech and provide it to the user. Specifically, it uses a speech synthesis API to create speech from text data and transmits it to the user through a speaker or earphones. As a result, the user receives the instructions and understands the next action.

[0602] Step 6:

[0603] The user acts based on the instructions received and provides feedback as needed. Input includes the results of the user's actions and any questions they may have. The user's feedback is sent back to the server, where its intent and context are interpreted. The specific action involves providing feedback again via voice input, prompting the server to analyze it. As a result, the feedback is re-analyzed by the server, and additional instructions or adjustments are made as necessary.

[0604] (Application Example 2)

[0605] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0606] In on-site work communication, uniform instructions that do not consider the psychological state or emotions of workers can lead to decreased efficiency and safety. In particular, when workers are feeling anxious or stressed, traditional methods may fail to provide appropriate support, leading to misunderstandings and work errors. A system is needed to solve these problems and flexibly respond to the psychological state of workers.

[0607] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0608] In this invention, the server includes means for receiving voice input from a worker, means for converting the received voice input into text data, means for analyzing the worker's intent from the text data, means for estimating the worker's level of understanding based on the analysis results, means for generating additional instructions according to the level of understanding, means for outputting the generated instructions to the worker via voice, means for re-analyzing feedback from the worker and generating instructions again, means for analyzing the worker's emotional state, and means for adjusting instructions based on the emotional state. This enables individualized support that takes into account the worker's psychological state, and makes it possible to provide a safe and effective work environment.

[0609] "Means for receiving voice input" refers to a device or process for capturing voice signals emitted by an operator as digital data.

[0610] "Means for converting to text data" refers to a device or process that uses speech recognition technology to convert a received audio signal into text information.

[0611] "Means for analyzing intent" refers to a device or process that utilizes natural language processing technology to identify the purpose or request of a worker's utterance from text data.

[0612] "Means for estimating comprehension" are methods for evaluating how accurately a worker understands information based on the analyzed intent.

[0613] "Means for generating additional instructions" refers to a device or process for automatically generating instructions to be provided to an operator based on an estimated level of understanding or other situational factors.

[0614] "Means for outputting voice" refers to a device or process for conveying generated instructions to a worker as voice using speech synthesis technology.

[0615] A "means for reanalyzing feedback" refers to a device or process for receiving responses and reactions from workers, reanalyzing that information, and determining further actions.

[0616] "Methods for analyzing emotional states" refer to techniques that analyze the intonation and tone of voice of a worker in order to analyze their emotions and identify their psychological state.

[0617] "Means for adjusting instructions" refers to a device or process for appropriately changing the content and tone of instructions given to a worker based on their analyzed emotional state.

[0618] This invention is implemented as a voicebot system designed to improve communication in the workplace. The system utilizes speech recognition and natural language processing technologies to convert worker voice input into text and analyze intent based on that text. Specifically, a server receives voice spoken by a worker and converts the voice into text in real time via a speech recognition API. For example, the Google Cloud Speech-to-Text API is used.

[0619] The server analyzes the converted text data using a natural language processing engine to accurately understand the worker's intentions and requests. In this process, Hugging Face's Transformers library is utilized, and advanced text analysis is achieved by incorporating a generative AI model.

[0620] Furthermore, the server evaluates the intonation, speed, and tone of the voice through an emotion engine that grasps the emotional state from the worker's voice. For this purpose, open-source libraries for speech emotion recognition, such as "librosa" and "praat," are used.

[0621] The server uses an AI model to customize appropriate instructions based on the user's level of understanding and emotional state. These instructions are adjusted to take emotions into consideration and communicated to the user using speech synthesis technology (e.g., Google Text-to-Speech). For example, if a user is feeling anxious, the system will provide appropriate instructions in a calm tone to help them regain their composure.

[0622] A key feature of this system is its ability to create a feedback loop that reflects the worker's psychological state. The server uses a generative AI model to generate prompts, enabling flexible instructions tailored to individual work scenarios. By handling prompts such as, "The worker is working on a difficult task but is feeling anxious. Please generate instructions to calm them down and explain the precise steps," the server provides effective support that is relevant to the situation on site.

[0623] In this way, the system aims to provide support tailored to the individual psychological needs of workers so that they can perform their duties safely and with peace of mind.

[0624] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0625] Step 1:

[0626] The server receives voice input of the worker's voice collected by the terminal. This voice data is acquired as input and recorded in digital format.

[0627] Step 2:

[0628] The server converts the received audio data into text data using a speech recognition API (e.g., Google Cloud Speech-to-Text API). The input is audio data, and the output is text data, which is the transcription of that audio.

[0629] Step 3:

[0630] The server inputs the converted text data into a natural language processing engine (e.g., Hugging Face Transformers) to analyze the worker's intent. The input is text data, and the output is the extracted intent or command.

[0631] Step 4:

[0632] Based on the analysis results, the server generates prompt sentences using a generative AI model and estimates the operator's level of understanding. The input is the result of intent analysis, and the output is an evaluation of understanding.

[0633] Step 5:

[0634] The server uses a speech emotion recognition library (e.g., librosa) to analyze emotional states from speech data. The input is speech data, and the output is stress levels and emotional characteristics.

[0635] Step 6:

[0636] The server considers both comprehension level and emotional state, and uses a generative AI model to create optimal additional instructions. The input is the result of comprehension and emotional analysis, and the output is the customized instructions.

[0637] Step 7:

[0638] The server outputs the generated instructions as speech using speech synthesis technology (e.g., Google Text-to-Speech). The input is instructions in text format, and the output is instructions that are then transmitted to the worker in speech format.

[0639] Step 8:

[0640] The user resumes work based on the instructions received and inputs subsequent feedback into the terminal. This feedback is sent to the server for further analysis and adjustment of instructions.

[0641] Step 9:

[0642] The server analyzes user feedback and adjusts instructions as needed. The input is feedback data, and the server improves support quality by outputting revised instructions.

[0643] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0644] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0645] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0646] [Fourth Embodiment]

[0647] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0648] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0649] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0650] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0651] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0652] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0653] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0654] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0655] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0656] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0657] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0658] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0659] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0660] This invention is implemented as a voice bot system that streamlines communication between field workers and managers. Specifically, a server uses speech recognition technology to convert voice input from workers (users) into text data in real time. The generated text data is analyzed using natural language processing. This analysis makes it possible to accurately understand the user's intentions and work status.

[0661] Next, the server applies a deep learning algorithm to estimate the user's level of understanding based on the analysis results. During this process, it also considers speech characteristics such as tone and speed to evaluate the stress level. This determines how well the user understands the instructions and how confidently they can proceed with the task.

[0662] Based on the user's state, the server generates appropriate additional instructions. These instructions are communicated to the user using speech synthesis technology to facilitate more effective communication. The user provides feedback on the voice-output instructions. This feedback is also analyzed by the server and, if necessary, leads to the generation of new instructions.

[0663] As a concrete example, when a worker inquires about the operating procedure for a specific piece of equipment, the server generates instructions that concisely explain the procedure. If signs of stress are detected in the worker's feedback, the server adds more detailed steps and precautions to the instructions and provides them again. In this way, the system is implemented to aid user understanding and ensure transparency in communication.

[0664] The following describes the processing flow.

[0665] Step 1:

[0666] The server receives voice input from the user and records the received audio in real time. The audio data is stored in a buffer for subsequent processing.

[0667] Step 2:

[0668] The device's speech recognition engine converts received speech data into text data. This conversion process ensures accurate transcription of the speech into text.

[0669] Step 3:

[0670] The server passes the text data to a natural language processing engine, which analyzes the user's speech. Based on the analysis results, it identifies what information or instructions the user is seeking.

[0671] Step 4:

[0672] The server estimates the user's level of understanding based on the analyzed information. Here, it evaluates the level of understanding based on the user's vocal characteristics and word choices, referencing past data and learning models.

[0673] Step 5:

[0674] The server generates appropriate additional instructions, taking into account the estimated level of understanding and stress levels. These instructions are customized to fit the context of the task.

[0675] Step 6:

[0676] The device communicates the generated instructions to the user using speech synthesis technology. This voice output is provided in a clear and easy-to-understand manner.

[0677] Step 7:

[0678] The user provides feedback on the given instructions. The server converts this feedback back into text, analyzes it, and uses it to generate new instructions.

[0679] Step 8:

[0680] The server repeats the above processing cycle as needed, continuing to assist the user in reaching the desired solution.

[0681] (Example 1)

[0682] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0683] There is a need to improve the efficiency of communication between field workers and managers. In particular, a challenge is to provide an environment where workers can accurately understand the instructions they receive and proceed with their work without stress. There is also the problem of difficulty in correcting instructions and incorporating feedback in real time.

[0684] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0685] In this invention, the server includes means for receiving an audio signal, means for converting the received audio signal into text data, and means for analyzing the text data to identify the worker's intent. This allows the worker to instantly transmit their intent to the server through voice input and receive appropriate instructions and feedback according to the level of understanding.

[0686] An "audio signal" is an electrical or digital representation of information transmitted by sound.

[0687] "Text data" refers to text data that represents audio, handwritten text, and other information in a standardized format.

[0688] "Intention" refers to the purpose or wish that the worker is trying to convey through voice input, and understanding that intention leads to more efficient communication.

[0689] "Comprehension level" is an indicator that shows how correctly a worker understands and can execute the instructions and information provided to them.

[0690] A "signal" refers to additional information or instructions given to a worker, and is communicated through voice or visual means.

[0691] "Psychological stress level" refers to the degree of stress and anxiety inferred from a worker's tone of voice and speed, and serves as a criterion for determining whether to provide appropriate support in the work environment.

[0692] "Natural language processing techniques" refer to technologies and approaches for enabling computers to understand and process human language.

[0693] This invention is implemented as a voice bot system to streamline communication between field workers (users) and administrators. The system functions by having a terminal receive the voice spoken by the user at the field site, and a server process that voice.

[0694] First, users perform voice input using a dedicated mobile device or headset. The device is equipped with a high-sensitivity microphone that can eliminate background noise and transmit clear audio to the server. This allows users to perform highly accurate voice input while moving freely.

[0695] The server uses "speech recognition software" to convert speech data to text in real time. This conversion utilizes common APIs, ensuring rapid speech-to-text conversion.

[0696] The server that acquires the text data uses "natural language processing software" to analyze the user's intent. For example, if a user enters the prompt "Tell me how to replace the part right now," the server can accurately interpret that intent and provide an answer by searching for relevant information. This process employs natural language processing technology and involves detailed analysis.

[0697] Furthermore, the server uses a "deep learning model" to assess the user's psychological load from the tone and speed of their voice. This allows it to estimate the user's level of stress and anxiety and generate additional instructions tailored to their understanding. For example, if the server determines that the user is having difficulty operating the system, it will generate cues that include more detailed explanations and priority steps.

[0698] Ultimately, the server uses "speech synthesis software" to provide the generated instructions as audio signals, which are then transmitted to the user. The user continues working based on the received audio and provides feedback as needed. This feedback is also processed by the server, leading to the provision of new information and improvements to the instructions.

[0699] A concrete example is a process where, after a worker on-site reports a machine malfunction, the server asks questions to identify the cause of the malfunction. Examples of prompts used in this process include, "The machine is vibrating abnormally; please diagnose the cause." This allows users to efficiently resolve problems while receiving support.

[0700] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0701] Step 1:

[0702] The user provides voice input through a dedicated terminal. The input consists of the user's questions and instructions, which are received as audio signals by the terminal's microphone. These audio signals are then converted into digital format within the terminal. The output is digitized audio data.

[0703] Step 2:

[0704] The terminal sends digitized audio data to the server. The server receives this data and uses "speech recognition software" to convert it into text data. During the conversion of audio data to text format, phoneme analysis is performed, and the output is generated as text. The output is text data.

[0705] Step 3:

[0706] The server analyzes the text data using "natural language processing software." Contextual analysis and keyword identification are performed to extract intent from the input text data. Data with clearly defined user intent and purpose is then output.

[0707] Step 4:

[0708] Based on the analyzed intent, the server uses a "deep learning model" to evaluate the user's understanding. This process also takes speech properties into consideration. Data regarding the estimated understanding and stress level is output.

[0709] Step 5:

[0710] The server generates additional instructions based on the user's level of understanding. Information processing is performed to format the generated instructions into an easily understandable form. The output is a written instruction provided to the user.

[0711] Step 6:

[0712] The server uses "speech synthesis software" to convert the generated instructions into audio signals. At this stage, the process generates natural-sounding, easily recognizable speech from the text data. The output is audio data transmitted to the user through the terminal.

[0713] Step 7:

[0714] The user performs tasks based on the voice instructions received and provides voice feedback as needed. This feedback is then captured again as digital voice data on the device. The output is the voice data of the feedback.

[0715] Step 8:

[0716] The server receives the audio data of the feedback, analyzes it again, and generates new instructions if necessary. This allows the system to continuously improve communication. The output consists of new instructions and feedback analysis results.

[0717] (Application Example 1)

[0718] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0719] In complex on-site tasks, smooth communication between workers and systems is essential. However, conventional systems based on voice input have difficulty accurately grasping the worker's level of understanding and emotional state. As a result, they are not adequately supporting workers in preventing misunderstandings and ensuring they can work efficiently and with confidence.

[0720] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0721] In this invention, the server includes a medium for acquiring voice data from a worker, means for converting the voice data into text information, and means for recognizing the worker's intentions from the text information. This makes it possible to accurately evaluate the worker's level of understanding and emotional state, and to provide appropriate additional information and instructions in a timely manner.

[0722] A "worker" is a person who operates equipment and facilities and performs tasks in a factory or on-site.

[0723] "Voice data" refers to data of acoustic signals acquired through the voice of a worker.

[0724] "Medium" refers to a physical or virtual device or means used to acquire or provide information.

[0725] "Textual information" refers to the textual representation of language converted from audio data.

[0726] "Means of recognition" refer to methods and technologies for analyzing acquired data and understanding its meaning and purpose.

[0727] "Comprehension level" is an indicator that shows how accurately a worker understands the instructions and information they are given.

[0728] "Means of evaluation" refer to methods and techniques for analyzing data and situations and making judgments about their quality and characteristics.

[0729] "Additional information" refers to detailed information provided to assist workers in understanding the work and to facilitate the smooth progress of the task.

[0730] A "device" is a mechanical or electronic system built to perform a specific function or task.

[0731] A "response" is a reaction or reply that an operator gives to instructions or information from a system.

[0732] "Analysis" is the process of examining data in detail to clarify its structure and meaning.

[0733] "Stress level" refers to the degree of burden or tension that workers feel towards their work and environment.

[0734] "To analyze with high accuracy" means to analyze data and information with great precision and detail.

[0735] This invention will now be described in terms of embodiments for carrying it out. The main element of the invention is an interactive system for smooth information exchange between a worker and a system. This system is realized by acquiring voice data via smart glasses or a headset worn by a worker on site and processing that voice on a server.

[0736] The server converts audio data acquired from the smart glasses' microphone into text in real time using the Google Speech-to-Text API. Then, it analyzes this text using NLTK (Natural Language Toolkit) to recognize the worker's intent. Based on this analysis, a deep learning model built with PyTorch is used to evaluate the worker's understanding and stress level. Based on this evaluation, the server generates additional information tailored to the worker and communicates it to them verbally using Google Text-to-Speech.

[0737] This process continuously provides optimal instructions by constantly acquiring new responses from workers, analyzing them again, and re-evaluating their level of understanding and stress. As a result, workers receive specific instructions tailored to their situation, enabling them to work efficiently and with peace of mind.

[0738] As a concrete example, consider a scenario where an operator troubleshoots a metalworking machine. The operator asks a question by voice, and the question is immediately analyzed to provide specific machine operating procedures and precautions. If the operator shows any signs of uncertainty, more detailed information will be provided later.

[0739] An example of a prompt message is, "Please provide detailed instructions for troubleshooting the metalworking machine." This prompt allows the system to quickly and accurately convey the necessary information, enabling the operator to efficiently resolve the problem accordingly.

[0740] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0741] Step 1:

[0742] The terminal acquires the worker's voice as audio data from the microphone in the smart glasses. The input is the worker's voice, and this audio data serves as the trigger to start the program.

[0743] Step 2:

[0744] The server converts the acquired audio data into text using the Google Speech-to-Text API. The input is audio data, and the output is text data. This process transforms the audio information into parseable text.

[0745] Step 3:

[0746] The server analyzes the converted character information using NLTK to recognize the worker's intent. The input is character information, and the output is data indicating the worker's intent. This analysis process clarifies the information and instructions the worker is seeking.

[0747] Step 4:

[0748] The server evaluates the worker's level of understanding and stress through a deep learning model using PyTorch. Inputs are intent recognition results and emotional data obtained from the worker's voice, while outputs are understanding and stress indicators. This allows the worker's state to be quantified, enabling appropriate responses based on that situation.

[0749] Step 5:

[0750] Based on the evaluation results, the server utilizes a generative AI model to generate optimal additional information and instructions for the worker. The inputs are comprehension levels and stress indicators, while the output is the content of the instructions for the worker. Appropriate content is provided according to the worker's condition.

[0751] Step 6:

[0752] The server uses Google Text-to-Speech to synthesize the generated instructions into speech and sends them to the terminal. The input is the generated instructions, and the output is the audio data. The terminal then transmits this audio to the worker.

[0753] Step 7:

[0754] The user provides feedback to the system again, and this voice feedback triggers a return to step 1 of the next cycle. The input is the worker's new voice feedback, which triggers the process of being analyzed again.

[0755] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0756] This invention is implemented as a voice bot system for improving communication in field work. By incorporating an emotion engine, this system not only provides instructions but also generates adaptive instructions that respond to the user's emotions and stress levels.

[0757] Specifically, the server receives voice input from the user and converts it into text data in real time using speech recognition. The transcribed text is then analyzed by a natural language processing engine to accurately understand the user's intentions and needs.

[0758] Furthermore, the server uses an emotion engine to identify the emotional characteristics of the speech. In this process, various speech features such as tone, speed, and intonation are considered to determine the user's psychological state. Meanwhile, the level of understanding is evaluated based on the analysis results in response to the user's questions and requests.

[0759] By combining emotional information evaluated based on the emotion engine with an assessment of comprehension, the system generates the most appropriate additional instructions. These instructions are customized to take into account the user's psychological state and work context. The terminal communicates these instructions to the user using speech synthesis and subsequently analyzes the received feedback.

[0760] For example, if a user makes a mistake in executing a work procedure, the server detects the user's anxiety and tension through its emotion engine. Based on this emotional information, it generates instructions in a calm tone to alleviate the anxiety and reinstructs the user to follow the correct procedure in a reassuring manner. This system is highly responsive to the user's psychology and state, enabling it to support a safe and secure work environment.

[0761] The following describes the processing flow.

[0762] Step 1:

[0763] The user provides voice input to the device. The device receives this voice input as digital audio data.

[0764] Step 2:

[0765] The server provides the received audio data to the speech recognition engine in real time, where it is converted into text data. The speech recognition engine accurately processes the conversion of speech into text.

[0766] Step 3:

[0767] The server passes the converted text data to a natural language processing engine, which analyzes the user's intent. This analysis helps understand what the user wants.

[0768] Step 4:

[0769] The server uses an emotion engine to analyze the audio data. The emotion engine evaluates the user's emotional state based on factors such as tone, speed, and rhythm of the voice.

[0770] Step 5:

[0771] The server combines the user's intent analysis results and emotion evaluation results to estimate their level of understanding and psychological state. Based on this estimation, it determines what kind of instructions are needed.

[0772] Step 6:

[0773] The server generates additional instructions tailored to the user based on their estimated level of understanding and emotional state. These instructions are designed to maximize the user's safety and efficiency.

[0774] Step 7:

[0775] The terminal uses speech synthesis to communicate additional instructions sent from the server to the user. The user understands the next action through the voice instructions.

[0776] Step 8:

[0777] The user provides feedback on the additional instructions. This feedback is converted back into text on the server and used to analyze the user's understanding and the need for continued support.

[0778] Step 9:

[0779] If necessary, the server optimizes the process, provides more specific instructions, and delivers them again to the user via the terminal. This iterative process continuously supports the user and reduces misunderstandings.

[0780] (Example 2)

[0781] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0782] There is a lack of instructions that take into account the psychological state of workers in on-site work, and it is necessary to reduce inefficiencies and errors caused by misunderstandings and stress. Furthermore, it is necessary to interpret workers' intentions with high accuracy and provide appropriate support accordingly.

[0783] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0784] In this invention, the server includes means for acquiring voice information of the worker, means for identifying emotional characteristics from the voice information, and means for adjusting the content of instructions based on the identified emotional information. This makes it possible to provide instructions that are adaptive and reassuring to the worker, in accordance with their psychological state.

[0785] "Audio information" refers to all audio data emitted by workers, and is material that is analyzed by speech recognition.

[0786] "Textual information" refers to text data converted from audio information and is used to interpret the worker's intent.

[0787] "Interpreting intent" is the process of identifying the purpose or requirements from what the worker says.

[0788] "Understanding level" is an indicator that shows how well a worker understands the information and instructions provided.

[0789] "Additional information" refers to supplementary instructions or explanations generated based on the worker's level of understanding.

[0790] "Provided via voice" means that generated instructions and information are communicated to workers using speech synthesis technology.

[0791] "Reinterpreting responses" is the process of re-analyzing intentions and situations based on feedback received from workers.

[0792] "Identification of emotional characteristics" is a technique that analyzes emotional indicators to estimate a worker's psychological state from their voice information.

[0793] "Adjusting instructions" refers to modifying the generated instructions and information to best suit the worker, based on identified emotional information.

[0794] This invention is implemented as a voice bot system for providing instructions tailored to the psychological state and understanding of workers during on-site work. This system aims to efficiently support workers by combining speech recognition, natural language processing, and sentiment analysis technologies.

[0795] The server processes the voice information obtained from the worker. This voice information is converted into text information using speech recognition software (e.g., speech recognition API). The text information is then analyzed using natural language processing technology (e.g., natural language API) to explore the worker's intentions and requests.

[0796] The server further uses an emotion analysis system to identify the worker's emotional characteristics from the voice information. This allows the worker's psychological state and stress level to be evaluated and used as foundational data to generate optimal instructions.

[0797] The generated instructions are provided to the worker from the terminal using speech synthesis technology (e.g., speech synthesis API). The worker receives these instructions, proceeds with the work, and sends additional voice feedback to the server as needed. The server analyzes this feedback and generates appropriate instructions again based on the newly evaluated information.

[0798] For example, if a worker says anxiously, "I don't know what to do in the next step," the server will identify that anxiety through emotion analysis and generate instructions in a reassuring tone. An example of a prompt message for this purpose would be, "Analyze the worker's voice to understand their emotions and clearly indicate the next work step in a relaxing tone."

[0799] This system enables the provision of adaptive, safe, and efficient work instructions tailored to the psychological state of workers, significantly improving communication in the workplace.

[0800] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0801] Step 1:

[0802] The server receives voice input from the user. The received voice data is sent to the server in real time. The server converts this data into text information using speech recognition technology. Specifically, it performs a speech-to-text conversion process via a speech recognition API. As a result, the voice data is output in text format.

[0803] Step 2:

[0804] The server inputs text information into its natural language processing engine. From this input text, the server performs keyword extraction and contextual analysis to interpret the user's intent. Specifically, it uses a natural language API to identify the meaning and purpose within the text data. As a result, it obtains analysis results regarding the user's intent.

[0805] Step 3:

[0806] The server analyzes the voice characteristics based on the original audio data and performs emotion analysis. The input used is the tone, speed, and intonation of the voice. The server sends this information to the emotion analysis system to evaluate the user's psychological state and stress level. Specifically, it extracts emotional characteristics using a voice analysis algorithm. As a result, emotion evaluation information is output.

[0807] Step 4:

[0808] The server generates the most appropriate additional instructions based on the analyzed user intent and sentiment evaluation. The input provided is the results of intent analysis and sentiment analysis. The server automatically creates situation-appropriate instructions using a generative AI model. Specifically, it operates the generative AI using adaptive prompt sentences to generate customized instructions. The resulting instructions are obtained.

[0809] Step 5:

[0810] The terminal transmits instructions from the server to the user via voice. The input received is the generated instruction content. The terminal uses speech synthesis technology to convert text instructions into natural-sounding speech and provide it to the user. Specifically, it uses a speech synthesis API to create speech from text data and transmits it to the user through a speaker or earphones. As a result, the user receives the instructions and understands the next action.

[0811] Step 6:

[0812] The user acts based on the instructions received and provides feedback as needed. Input includes the results of the user's actions and any questions they may have. The user's feedback is sent back to the server, where its intent and context are interpreted. The specific action involves providing feedback again via voice input, prompting the server to analyze it. As a result, the feedback is re-analyzed by the server, and additional instructions or adjustments are made as necessary.

[0813] (Application Example 2)

[0814] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0815] In on-site work communication, uniform instructions that do not consider the psychological state or emotions of workers can lead to decreased efficiency and safety. In particular, when workers are feeling anxious or stressed, traditional methods may fail to provide appropriate support, leading to misunderstandings and work errors. A system is needed to solve these problems and flexibly respond to the psychological state of workers.

[0816] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0817] In this invention, the server includes means for receiving voice input from a worker, means for converting the received voice input into text data, means for analyzing the worker's intent from the text data, means for estimating the worker's level of understanding based on the analysis results, means for generating additional instructions according to the level of understanding, means for outputting the generated instructions to the worker via voice, means for re-analyzing feedback from the worker and generating instructions again, means for analyzing the worker's emotional state, and means for adjusting instructions based on the emotional state. This enables individualized support that takes into account the worker's psychological state, and makes it possible to provide a safe and effective work environment.

[0818] "Means for receiving voice input" refers to a device or process for capturing voice signals emitted by an operator as digital data.

[0819] "Means for converting to text data" refers to a device or process that uses speech recognition technology to convert a received audio signal into text information.

[0820] "Means for analyzing intent" refers to a device or process that utilizes natural language processing technology to identify the purpose or request of a worker's utterance from text data.

[0821] "Means for estimating comprehension" are methods for evaluating how accurately a worker understands information based on the analyzed intent.

[0822] "Means for generating additional instructions" refers to a device or process for automatically generating instructions to be provided to an operator based on an estimated level of understanding or other situational factors.

[0823] "Means for outputting voice" refers to a device or process for conveying generated instructions to a worker as voice using speech synthesis technology.

[0824] A "means for reanalyzing feedback" refers to a device or process for receiving responses and reactions from workers, reanalyzing that information, and determining further actions.

[0825] "Methods for analyzing emotional states" refer to techniques that analyze the intonation and tone of voice of a worker in order to analyze their emotions and identify their psychological state.

[0826] "Means for adjusting instructions" refers to a device or process for appropriately changing the content and tone of instructions given to a worker based on their analyzed emotional state.

[0827] This invention is implemented as a voicebot system designed to improve communication in the workplace. The system utilizes speech recognition and natural language processing technologies to convert worker voice input into text and analyze intent based on that text. Specifically, a server receives voice spoken by a worker and converts the voice into text in real time via a speech recognition API. For example, the Google Cloud Speech-to-Text API is used.

[0828] The server analyzes the converted text data using a natural language processing engine to accurately understand the worker's intentions and requests. In this process, Hugging Face's Transformers library is utilized, and advanced text analysis is achieved by incorporating a generative AI model.

[0829] Furthermore, the server evaluates the intonation, speed, and tone of the voice through an emotion engine that grasps the emotional state from the worker's voice. For this purpose, open-source libraries for speech emotion recognition, such as "librosa" and "praat," are used.

[0830] The server uses an AI model to customize appropriate instructions based on the user's level of understanding and emotional state. These instructions are adjusted to take emotions into consideration and communicated to the user using speech synthesis technology (e.g., Google Text-to-Speech). For example, if a user is feeling anxious, the system will provide appropriate instructions in a calm tone to help them regain their composure.

[0831] A key feature of this system is its ability to create a feedback loop that reflects the worker's psychological state. The server uses a generative AI model to generate prompts, enabling flexible instructions tailored to individual work scenarios. By handling prompts such as, "The worker is working on a difficult task but is feeling anxious. Please generate instructions to calm them down and explain the precise steps," the server provides effective support that is relevant to the situation on site.

[0832] In this way, the system aims to provide support tailored to the individual psychological needs of workers so that they can perform their duties safely and with peace of mind.

[0833] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0834] Step 1:

[0835] The server receives voice input of the worker's voice collected by the terminal. This voice data is acquired as input and recorded in digital format.

[0836] Step 2:

[0837] The server converts the received audio data into text data using a speech recognition API (e.g., Google Cloud Speech-to-Text API). The input is audio data, and the output is text data, which is the transcription of that audio.

[0838] Step 3:

[0839] The server inputs the converted text data into a natural language processing engine (e.g., Hugging Face Transformers) to analyze the worker's intent. The input is text data, and the output is the extracted intent or command.

[0840] Step 4:

[0841] Based on the analysis results, the server generates prompt sentences using a generative AI model and estimates the operator's level of understanding. The input is the result of intent analysis, and the output is an evaluation of understanding.

[0842] Step 5:

[0843] The server uses a speech emotion recognition library (e.g., librosa) to analyze emotional states from speech data. The input is speech data, and the output is stress levels and emotional characteristics.

[0844] Step 6:

[0845] The server considers both comprehension level and emotional state, and uses a generative AI model to create optimal additional instructions. The input is the result of comprehension and emotional analysis, and the output is the customized instructions.

[0846] Step 7:

[0847] The server outputs the generated instructions as speech using speech synthesis technology (e.g., Google Text-to-Speech). The input is instructions in text format, and the output is instructions that are then transmitted to the worker in speech format.

[0848] Step 8:

[0849] The user resumes work based on the instructions received and inputs subsequent feedback into the terminal. This feedback is sent to the server for further analysis and adjustment of instructions.

[0850] Step 9:

[0851] The server analyzes user feedback and adjusts instructions as needed. The input is feedback data, and the server improves support quality by outputting revised instructions.

[0852] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0853] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0854] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0855] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0856] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0857] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0858] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0859] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0860] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0861] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0862] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0863] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0864] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0865] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0866] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0867] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0868] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0869] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0870] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0871] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0872] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[0873] The following is further disclosed regarding the embodiments described above.

[0874] (Claim 1)

[0875] A means of receiving voice input from workers,

[0876] A means of converting received voice input into text data,

[0877] A means of analyzing the worker's intent from text data,

[0878] A means to estimate the worker's level of understanding based on the analysis results,

[0879] A means of generating additional instructions according to the level of understanding,

[0880] A means for outputting the generated instructions to the worker via voice,

[0881] A means of re-analyzing feedback from workers and generating new instructions,

[0882] A system that includes this.

[0883] (Claim 2)

[0884] The system according to claim 1, comprising means for analyzing the tone of a worker's voice to estimate their stress level.

[0885] (Claim 3)

[0886] The system according to claim 1, comprising means for analyzing the worker's intentions with high accuracy using natural language processing technology.

[0887] "Example 1"

[0888] (Claim 1)

[0889] A means for receiving audio signals,

[0890] A means for converting a received audio signal into text data,

[0891] A means of analyzing text data to identify the worker's intent,

[0892] A means of determining the worker's level of understanding based on the analyzed intent,

[0893] A means of generating additional signals according to the level of understanding,

[0894] A means of providing the generated signal to the worker as an audible signal,

[0895] A means for analyzing the worker's response and regenerating the signal,

[0896] A system that includes this.

[0897] (Claim 2)

[0898] The system according to claim 1, comprising means for analyzing the tone of voice of a worker to estimate the level of psychological stress.

[0899] (Claim 3)

[0900] The system according to claim 1, comprising means for identifying the worker's intentions with high accuracy using natural language processing techniques.

[0901] "Application Example 1"

[0902] (Claim 1)

[0903] A medium for acquiring voice data from workers,

[0904] A means of converting acquired audio data into text information,

[0905] A means of recognizing the worker's intent from textual information,

[0906] A means of evaluating the worker's level of understanding based on the recognition results,

[0907] A means of generating additional information according to the level of understanding,

[0908] A device that outputs the generated information to the worker via voice,

[0909] A means of re-recognizing the response from the worker and generating information again,

[0910] A means of analyzing the emotional state of workers and estimating their stress levels,

[0911] A system that includes this.

[0912] (Claim 2)

[0913] The system according to claim 1, comprising a medium that displays instructions in real time via a visual display.

[0914] (Claim 3)

[0915] The system according to claim 1, comprising means for analyzing the worker's intentions with high accuracy using language analysis technology.

[0916] "Example 2 of combining an emotion engine"

[0917] (Claim 1)

[0918] A means of acquiring voice information from workers,

[0919] A means of converting acquired audio information into text information,

[0920] A means of interpreting the worker's intent from textual information,

[0921] A means of evaluating the worker's level of understanding based on the interpretation results,

[0922] A means of generating additional information according to the level of understanding,

[0923] A means of providing the generated information to the worker via voice,

[0924] A means of reinterpreting worker responses and generating new information,

[0925] A means of identifying emotional characteristics from audio information,

[0926] Means for adjusting instructions based on identified emotional information,

[0927] A system that includes this.

[0928] (Claim 2)

[0929] The system according to claim 1, comprising means for analyzing the voice characteristics of a worker to identify their psychological state.

[0930] (Claim 3)

[0931] The system according to claim 1, comprising means for interpreting the worker's intentions with high accuracy using natural language processing technology.

[0932] "Application example 2 when combining with an emotional engine"

[0933] (Claim 1)

[0934] A means of receiving voice input from workers,

[0935] A means of converting received voice input into text data,

[0936] A means of analyzing the worker's intent from text data,

[0937] A means to estimate the worker's level of understanding based on the analysis results,

[0938] A means of generating additional instructions according to the level of understanding,

[0939] A means for outputting the generated instructions to the worker via voice,

[0940] A means of re-analyzing feedback from workers and generating new instructions,

[0941] A means of analyzing the emotional state of workers,

[0942] A means of adjusting instructions based on emotional state,

[0943] A system that includes this.

[0944] (Claim 2)

[0945] The system according to claim 1, comprising means for analyzing the tone of a worker's voice to estimate their stress level, and means for optimizing instructions by combining their emotional state and level of understanding.

[0946] (Claim 3)

[0947] The system according to claim 1, comprising means for analyzing the worker's intent with high accuracy using natural language processing technology, and generating context-appropriate prompt sentences and instructions using a generative AI model. [Explanation of Symbols]

[0948] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. A means of receiving voice input from workers, A means of converting received voice input into text data, A means of analyzing the worker's intent from text data, A means to estimate the worker's level of understanding based on the analysis results, A means of generating additional instructions according to the level of understanding, A means for outputting the generated instructions to the worker via voice, A means of re-analyzing feedback from workers and generating new instructions, A system that includes this.

2. The system according to claim 1, comprising means for analyzing the tone of a worker's voice to estimate their stress level.

3. The system according to claim 1, comprising means for analyzing the worker's intentions with high accuracy using natural language processing technology.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A